Skip to content

Kafka: Cluster Health Check

A Kafka cluster in KRaft mode doesn’t forgive neglect until the first incident. Health checks need to be regular and quick — no graphs or dashboards, just the terminal. Here’s the command set that covers the typical checklist: processes, quorum, partition leaders, ISR, and a quick status report in one shot.

Checking KRaft Processes

KRaft mode has no separate ZooKeeper — the controller role is either co-located with the broker or isolated on dedicated nodes. First, verify the JVM processes are alive and see which mode each node started in.

ps -ef | grep -E 'kafka.Kafka|QuorumControllerMain' | grep -v grep

For a managed cluster you expect three QuorumControllerMain processes on dedicated controllers and N kafka.Kafka processes on brokers. Mixed nodes (combined mode) run both roles in a single process — normal for smaller installations.

Then check startup logs and configuration:

grep -E 'process.roles|node.id|controller.quorum.voters|listeners=' config/kraft/server.properties
Note

The process.roles parameter accepts broker, controller, or broker,controller. If empty, the cluster is still in legacy mode with ZooKeeper — adapt the commands below.

Verify all nodes agreed on the quorum:

/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server localhost:9092 describe --status

In the output look for leaderId, votedLeaders, and quorum size 2/3 (for three controllers) or N/N for a fully stabilized cluster.

Broker and Controller Status

Next, determine which brokers are actually responding and which dropped from the registry. Use kafka-broker-api-versions.sh — it returns supported API versions and simultaneously shows whether TCP connectivity reaches the broker.

for h in kafka1 kafka3 kafka5; do
  echo "=== $h ==="
  /opt/kafka/bin/kafka-broker-api-versions.sh \
    --bootstrap-server $h:9092 2>&1 | head -n 3
done

If a node is unreachable you’ll see a timeout. This is the fastest way to distinguish “broker stuck in JVM” from “network issue.”

Full list of registered brokers and their state:

/opt/kafka/bin/kafka-broker-api-versions.sh \
  --bootstrap-server kafka1:9092 | head -n 50

The command shows a JSON-like listing, but for a tabular report kafka-metadata-quorum.sh is more convenient:

/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server kafka1:9092 describe --replicas

The LEADER column shows the current quorum leader, REPLICAS shows all active nodes. Unresponsive controllers drop from the list.

Tip

The lastCaughtUpTime field in describe --status shows how far a controller lags behind the leader. 0 or fresh now — healthy. Lag in minutes — reason to check GC and network latency.

Topics and Partition Leaders

The cluster can be alive but without partition leaders — producers won’t write, consumers won’t read. So the next step is topics and their leaders.

List all topics with partition and replica counts:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal

The output table contains Leader, Replicas, Isr. If Leader equals -1, the partition has no active leader — this is an emergency, producers will receive NotLeaderForPartitionException.

To get only “bad” partitions in one pipeline:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '$5 == -1 || $5 == "none" {print}'

To view leaders for a specific topic:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --topic orders.events

ISR and Unavailable Replicas

ISR (in-sync replicas) determines write reliability. Recommended setting is min.insync.replicas >= 2 for critical topics. Check for out-of-sync:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --under-replicated-partitions

The command returns partitions where Isr is smaller than Replicas. Empty output — good. Any lines in the list — incident.

Full picture across all partitions with problem filtering:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '{
    replicas=$5; isr=$7;
    # Replicas and Isr are comma-separated id lists; length by commas
    rep_n = split(replicas, a, ",");
    isr_n = split(isr, b, ",");
    if (rep_n != isr_n) print "UNSYNC:", $0;
    if (a[1] == "-1") print "NO_LEADER:", $0;
  }'
Warning

kafka-topics.sh --describe on a topic with thousands of partitions outputs many lines and loads the controller. In production run with --partitions N or filter with awk, otherwise the health check itself becomes a problem.

Additionally — state of a specific replica on the broker side. If you suspect one of the disks is lagging:

/opt/kafka/bin/kafka-log-dirs.sh \
  --bootstrap-server kafka1:9092 \
  --describe --broker-list 1,3,5 \
  | jq '.[] | .logDirs[] | {broker: .broker, dir: .dir, partitions: (.partitions | length)}'

In the output check the partition.error field — if non-empty, the replica has issues (offline log dir, disk full, fs in read-only).

Quick Diagnostics in One Command

For daily rounds it’s convenient to bundle all checks into a single script with clear exit codes. Below is a minimal status indicator in bash.

#!/usr/bin/env bash
set -u
BOOTSTRAP="${BOOTSTRAP:-kafka1:9092}"
KAFKA_BIN="${KAFKA_BIN:-/opt/kafka/bin}"
fail=0

echo "== Quorum status =="
if ! $KAFKA_BIN/kafka-metadata-quorum.sh --bootstrap-server "$BOOTSTRAP" \
    describe --status 2>&1 | grep -q 'isLeader: true'; then
  echo "WARN: quorum leader not confirmed"; fail=1
fi

echo "== Brokers reachability =="
for h in $(echo "$BOOTSTRAP" | tr ',' ' '); do
  if ! timeout 5 bash -c "echo > /dev/tcp/${h%:*}/${h##*:}"; then
    echo "FAIL: $h unreachable"; fail=1
  fi
done

echo "== Under-replicated partitions =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --under-replicated-partitions 2>/dev/null)
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: 0 under-replicated partitions"
fi

echo "== Partitions without leader =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --exclude-internal 2>/dev/null \
      | awk '$5 == -1 {print}')
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: every partition has a leader"
fi

exit $fail

Save as kafka-health.sh, make executable, and wrap in cron or a systemd timer every 60 seconds:

chmod +x kafka-health.sh
*/1 * * * * /usr/local/bin/kafka-health.sh \
  >> /var/log/kafka-health.log 2>&1
Tip

For alerting, replace echo blocks with logger -p local0.err and configure rsyslog to your SIEM. No need for Prometheus node_exporter — events fly into the common channel.

Typical errors on first run and how to interpret them:

SymptomProbable CauseAction
Connection to node -1 could not be establishedBroker not registered in cluster but process is aliveCheck advertised.listeners, node.id, network ACL
isLeader: false for all controllersQuorum lost, controllers < __.min.insync.replicasCheck controller.quorum.voters and disk state on controllers
--under-replicated-partitions shows entries for >5 minBroker lagging, slow disk or GC pausesTake jstack, check iostat -x and LogFlushRate metrics
kafka-log-dirs.sh returns LogDirOfflineDisk full or failedFree space, check FS for read-only, restart broker
Empty --describe output on a running clusterWrong bootstrap address passedVerify advertised.listeners and DNS

The five command sets above cover 90% of operational “is the cluster alive?” questions. If everything is green — no need to dig deeper; if something is red — kafka-log-dirs.sh and kafka-metadata-quorum.sh describe --status will show which direction to go.