TencentOS Kernel Crash Analysis Expert
Paste the following prompt into your AI chat to install this skill:
Please install @user_7ac3e421/tencentos-crash-expert according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem Solved
Analyzing Linux kernel crashes (vmcore) typically relies on manually executing crash commands, which presents two main issues: global states and local call stacks are easily confused, leading to "anchoring bias"; and the lack of a structured validation chain when searching for fix commits and evaluating mitigations often results in unverified, ineffective recommendations. This skill breaks down the analysis into 7 strictly controlled phases, enforcing "evidence-driven" investigation and "causal validation".
Core Capabilities & Key Steps
Powered by the AiCrasher MCP tool for automated crash debugging, the core workflow includes:
* Global Scanning First: In Soft/Hard Lockup scenarios, it mandates a global CPU state check (bt -a) first, preventing the misidentification of the Panic CPU's local snapshot as the root cause.
* Root Cause Routing: After analysis, it forces a "Kernel Defect" or "Non-Kernel Defect" classification. If it is a kernel defect, it proceeds to search for upstream fix commits and distro patch status; if it is a business configuration or user-space issue, it skips patch searching to avoid invalid retrieval.
* Parameter & Unit Cross-Validation: Built-in strict error-prevention mechanisms, such as verifying the RSS unit in ps output (KB, not pages), and requiring a three-step validation (source code, disassembly, and struct offsets) to confirm the actual runtime value of sysctl parameters.
* Standardized Reporting: Finally, it uses the crash_report_generator.py script to render the conclusions and crash command logs (@cmd[] references) into a Chinese HTML report, automatically releasing the MCP session resources.
Boundaries & Caveats
- Environment Dependency: Strongly depends on the registration and session startup of
AiCrasher MCP Server; a session restart is required after initial registration for hot-loading. - Source Code & Disassembly: For
sysctlparameter troubleshooting or patch validation, a matchingkernel-debuginfoor kernel sourcegitrepository is required. Thesymcommand alone cannot locateper-netnsstruct fields. - Long-Context Forgetting Prevention: The skill enforces a "phase-switching rule", requiring the agent to re-read the corresponding
phase*.mdstep file before entering each new phase to prevent early instructions from being diluted in long conversations.
Use Cases
- A production server records a vmcore after a soft lockup, and you need to test whether the panic CPU stack is only local.
- You receive an OOM crash dump and need to decide whether a kernel leak or a bad workload setting made processes unkillable.
- With a kernel source repo available, you need to find the upstream fix commit and identify which distro tag contains it.
- After analyzing a NULL pointer dereference, you need to produce validated sysctl mitigation advice and an HTML report.
Best For
- Ops engineers responsible for postmortems on production Linux server crashes who need vmcore evidence in a deliverable root-cause report.
- Infrastructure engineers handling kernel crash tickets who need to separate kernel defects from workload or user-space triggers.
- Distro kernel patch maintainers who need to verify fix commits, release tags, and disassembly evidence in a local git repo.
- SREs doing security and availability troubleshooting who need vmcore-based validation before recommending mitigations.
Related Skills
A lightweight wrapper and automation tool for Jaeger-related GitHub scenarios.
Vault Wrap is a wrapping skill for Vault and GitHub automation.
Port management, threat-intel audits, drift checks, and multi-node monitoring for self-hosted infrastructure.
Provides health checks for HTTP, TCP, Ping, and log sources with dynamic-threshold alerts, anomaly detection, and scheduled inspection reports.