Clear incident communication (calm updates, timelines, stakeholder management)
Problem-solving and root-cause analysis
Prioritization and risk judgment (what to fix now vs. later)
Linux fundamentals and command-line troubleshooting
Networking basics (DNS, HTTP, load balancing, common failure modes)
Programming/scripting for automation (Python, Go, Java, or similar)
Monitoring and observability (metrics, logs, tracing; meaningful alerts)
Cloud platforms (AWS, Azure, or GCP) and core services (compute, storage, networking)
Containers and orchestration (Docker, Kubernetes)
Infrastructure as Code (Terraform, CloudFormation, Pulumi)
CI/CD and release practices (safe rollouts, feature flags, rollback strategies)
Reliability methods (SLIs/SLOs, error budgets, capacity planning)