๐ Getting started
What is Kubernetes?โ
Kubernetes (often called K8s) runs your applications in containers. It keeps those containers running, moves them between machines, and starts new ones when needed.
That also means there are many moving parts. A container can run out of memory, a deployment can get stuck, or a service can lose its healthy backends.
What is kwatch?โ
kwatch is an open-source Kubernetes incident monitor. It turns failures into clear alerts that explain what broke, why it happened, and what to do next:
- ๐ Watch โ kwatch reads Kubernetes status, events, and recent logs.
- ๐ง Explain โ it connects the clues and finds the likely cause.
- ๐ฃ Alert โ it sends the reason, impact, and next step to your team.
kwatch runs in your own cluster, with no hosted account required. Pair it with Prometheus, Grafana, or Loki for long-term metrics and logs.
๐จ What an alert looks likeโ
Instead of only seeing CrashLoopBackOff, you get a message like:
๐จ OOMKilled โ production / orders-api
Pod: orders-api-7ffc9d4f9-x9p4t
Node: worker-3 ยท severity: high
๐ก Cause: the container exceeded its 512Mi memory limit.
โก๏ธ Next step: increase limits.memory or reduce memory usage.
๐ Recent logs and Kubernetes events are included.
๐ฏ What does kwatch watch?โ
Most monitors are enabled by default:
| Signal | What kwatch explains |
|---|---|
| ๐ฅ Pod crashes | Crash reason, logs, events, and a next step |
| โณ Pending pods | Why the scheduler cannot place a pod |
| ๐ฅ๏ธ Nodes | Readiness and disk or memory pressure |
| ๐ Deployments | Stuck rollouts and unavailable replicas |
| ๐งฉ StatefulSets and DaemonSets | Unavailable or stuck workloads |
| ๐งโ๐ผ Jobs and CronJobs | Failed, suspended, or missed work |
| ๐ HPA | An autoscaler stuck at its replica limit |
| ๐ฃ Cluster autoscaler | Evidence that scaling could not happen |
| ๐พ PVCs | Storage pressure and volume failures |
| ๐ Services and Ingress | Missing or unhealthy backends |
| ๐๏ธ Control plane | API server and platform health signals |
TLS certificate monitoring and heartbeat notifications are opt-in. See the configuration guide or the complete configuration reference for every available key.
๐ Install in three stepsโ
1. Check your toolsโ
You need Bash, kubectl, curl, and cluster install permissions. Confirm
that kubectl can reach your cluster:
kubectl cluster-info
2. Run the managerโ
/bin/bash -c "$(curl -fsSL https://kwatch.dev/kwatch.sh)"
The manager asks where alerts should go, stores credentials safely in a Secret, installs kwatch, and verifies the installation.
3. Check the resultโ
kubectl get pods -n kwatch
By default you should see two kwatch pods with STATUS Running: one active
leader and one standby. Both pods should show their container as READY 1/1;
only the leader is ready for monitoring. A one-replica installation is also
supported, but it has no kwatch self-failover. ๐
๐ ๏ธ What next?โ
- Need the full lifecycle and troubleshooting guide? Read Installation.
- Want Slack, Discord, email, or PagerDuty? Open Channels.
- Want to change thresholds or silence known noise? Read Configuration.
- Want to understand the manager? Read kwatch.sh manager.
- Want to contribute code? Start with Contributing.
If you are unsure where to begin, install with the manager first. You can run it again later to configure, upgrade, check, or uninstall kwatch. โจ