Kafka: Cluster Health Check
A Kafka cluster in KRaft mode doesn’t forgive neglect until the first incident. Health checks need to be regular and quick — no graphs or dashboards, just the terminal. Here’s the command set that covers the typical checklist: processes, quorum, partition leaders, ISR, and a quick status report in one shot.
Checking KRaft Processes
KRaft mode has no separate ZooKeeper — the controller role is either co-located with the broker or isolated on dedicated nodes. First, verify the JVM processes are alive and see which mode each node started in.
For a managed cluster you expect three QuorumControllerMain processes on dedicated controllers and N kafka.Kafka processes on brokers. Mixed nodes (combined mode) run both roles in a single process — normal for smaller installations.
Then check startup logs and configuration:
The process.roles parameter accepts broker, controller, or broker,controller. If empty, the cluster is still in legacy mode with ZooKeeper — adapt the commands below.
Verify all nodes agreed on the quorum:
In the output look for leaderId, votedLeaders, and quorum size 2/3 (for three controllers) or N/N for a fully stabilized cluster.
Broker and Controller Status
Next, determine which brokers are actually responding and which dropped from the registry. Use kafka-broker-api-versions.sh — it returns supported API versions and simultaneously shows whether TCP connectivity reaches the broker.
If a node is unreachable you’ll see a timeout. This is the fastest way to distinguish “broker stuck in JVM” from “network issue.”
Full list of registered brokers and their state:
The command shows a JSON-like listing, but for a tabular report kafka-metadata-quorum.sh is more convenient:
The LEADER column shows the current quorum leader, REPLICAS shows all active nodes. Unresponsive controllers drop from the list.
The lastCaughtUpTime field in describe --status shows how far a controller lags behind the leader. 0 or fresh now — healthy. Lag in minutes — reason to check GC and network latency.
Topics and Partition Leaders
The cluster can be alive but without partition leaders — producers won’t write, consumers won’t read. So the next step is topics and their leaders.
List all topics with partition and replica counts:
The output table contains Leader, Replicas, Isr. If Leader equals -1, the partition has no active leader — this is an emergency, producers will receive NotLeaderForPartitionException.
To get only “bad” partitions in one pipeline:
To view leaders for a specific topic:
ISR and Unavailable Replicas
ISR (in-sync replicas) determines write reliability. Recommended setting is min.insync.replicas >= 2 for critical topics. Check for out-of-sync:
The command returns partitions where Isr is smaller than Replicas. Empty output — good. Any lines in the list — incident.
Full picture across all partitions with problem filtering:
kafka-topics.sh --describe on a topic with thousands of partitions outputs many lines and loads the controller. In production run with --partitions N or filter with awk, otherwise the health check itself becomes a problem.
Additionally — state of a specific replica on the broker side. If you suspect one of the disks is lagging:
In the output check the partition.error field — if non-empty, the replica has issues (offline log dir, disk full, fs in read-only).
Quick Diagnostics in One Command
For daily rounds it’s convenient to bundle all checks into a single script with clear exit codes. Below is a minimal status indicator in bash.
Save as kafka-health.sh, make executable, and wrap in cron or a systemd timer every 60 seconds:
For alerting, replace echo blocks with logger -p local0.err and configure rsyslog to your SIEM. No need for Prometheus node_exporter — events fly into the common channel.
Typical errors on first run and how to interpret them:
| Symptom | Probable Cause | Action |
|---|---|---|
Connection to node -1 could not be established | Broker not registered in cluster but process is alive | Check advertised.listeners, node.id, network ACL |
isLeader: false for all controllers | Quorum lost, controllers < __.min.insync.replicas | Check controller.quorum.voters and disk state on controllers |
--under-replicated-partitions shows entries for >5 min | Broker lagging, slow disk or GC pauses | Take jstack, check iostat -x and LogFlushRate metrics |
kafka-log-dirs.sh returns LogDirOffline | Disk full or failed | Free space, check FS for read-only, restart broker |
Empty --describe output on a running cluster | Wrong bootstrap address passed | Verify advertised.listeners and DNS |
The five command sets above cover 90% of operational “is the cluster alive?” questions. If everything is green — no need to dig deeper; if something is red — kafka-log-dirs.sh and kafka-metadata-quorum.sh describe --status will show which direction to go.