AI Agent Hub
Back to skills
🔒

Linux System Administration Gotchas

IT Ops & Security Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_f12a44b7/self-dev-linux.

About this skill

The Problem

Linux troubleshooting often fails because the symptom does not match the root cause: chmod 777 masks permission issues, kill sends SIGTERM by default rather than SIGKILL, deleted files do not free disk space until their file descriptors are closed, df / du output appears contradictory, wrong SSH permissions cause silent authentication failures, and Systemd services may look configured but fail to start or lose logs. This skill turns those high-frequency pitfalls into a searchable checklist, helping engineers move beyond “did the command error?” and inspect the actual causes: ACLs, process state, open file handles, inode exhaustion, network persistence, timezone, OOM killer behavior, and container memory limits.

How It Works

It provides scenario-based diagnostic paths rather than a flat list of commands:

  • Permissions: Inspect owner/group and getfacl first, instead of blindly applying chmod 777; use umask 077 for sensitive directories.
  • Processes and resources: Use lsof +L1 to find deleted files still held by processes; use dmesg, vmstat, and /proc/[pid]/status to assess OOM events, swap thrashing, and real memory usage.
  • Disk and services: Check journals, Docker overlay storage, snapshots, and inode pressure; in Systemd, distinguish start, enable, and reload, and configure Storage=persistent and Restart=on-failure.
  • Networking and SSH: Prefer ss over deprecated netstat, ensure firewall rules persist across reboots, check SSH directory and key permissions with 700 / 600, and treat agent forwarding as a security risk.
  • Scheduling: Use absolute paths, explicit TZ, and output redirection in Cron instead of relying on mail delivery or the current shell environment.

Scope and Caveats

It is useful for general Linux system administration, incident diagnosis, and operational security review, especially for quickly checking common pitfalls before deeper debugging. It does not replace distro-specific hardening, container orchestration guidance, or cloud provider snapshot policies; in multi-user, containerized, or complex ACL environments, it should be used alongside concrete permission models, audit practices, and backup workflows.

Use Cases

  • Diagnose unresponsive service ports or silent SSH authentication failures by checking permissions, firewall rules, and host mappings.
  • Investigate disk write failures despite free space, or deleted files still holding space, by inspecting inodes, open handles, and snapshots.
  • Review Cron jobs for pitfalls in PATH, timezone handling, output redirection, and crontab backups.
  • Assess container and host memory pressure using cgroups, OOM logs, and VmRSS to identify real memory consumption.

Best For

  • SREs responsible for production Linux incident diagnosis: need to quickly identify permission, port, process, and disk anomalies.
  • Operations engineers managing SSH jump hosts and server access: need to troubleshoot authentication failures, permissions, and timeouts.
  • Backend engineers deploying scheduled tasks and system services: need to validate Cron, Systemd persistence, and logging behavior.
  • Cloud platform engineers handling container resource issues: need to distinguish cgroups, system OOM, and host memory state.