# Kafka: Cluster Health Check

LLMS index: [llms.txt](/en/llms.txt)

---

A Kafka cluster in KRaft mode doesn't forgive neglect until the first incident. Health checks need to be regular and quick — no graphs or dashboards, just the terminal. Here's the command set that covers the typical checklist: processes, quorum, partition leaders, ISR, and a quick status report in one shot.

## Checking KRaft Processes

KRaft mode has no separate ZooKeeper — the `controller` role is either co-located with the broker or isolated on dedicated nodes. First, verify the JVM processes are alive and see which mode each node started in.

```bash
ps -ef | grep -E 'kafka.Kafka|QuorumControllerMain' | grep -v grep
```

For a managed cluster you expect three `QuorumControllerMain` processes on dedicated controllers and N `kafka.Kafka` processes on brokers. Mixed nodes (combined mode) run both roles in a single process — normal for smaller installations.

Then check startup logs and configuration:

```bash
grep -E 'process.roles|node.id|controller.quorum.voters|listeners=' config/kraft/server.properties
```

> [!NOTE]
> The `process.roles` parameter accepts `broker`, `controller`, or `broker,controller`. If empty, the cluster is still in legacy mode with ZooKeeper — adapt the commands below.

Verify all nodes agreed on the quorum:

```bash
/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server localhost:9092 describe --status
```

In the output look for `leaderId`, `votedLeaders`, and quorum size `2/3` (for three controllers) or `N/N` for a fully stabilized cluster.

## Broker and Controller Status

Next, determine which brokers are actually responding and which dropped from the registry. Use `kafka-broker-api-versions.sh` — it returns supported API versions and simultaneously shows whether TCP connectivity reaches the broker.

```bash
for h in kafka1 kafka3 kafka5; do
  echo "=== $h ==="
  /opt/kafka/bin/kafka-broker-api-versions.sh \
    --bootstrap-server $h:9092 2>&1 | head -n 3
done
```

If a node is unreachable you'll see a timeout. This is the fastest way to distinguish "broker stuck in JVM" from "network issue."

Full list of registered brokers and their state:

```bash
/opt/kafka/bin/kafka-broker-api-versions.sh \
  --bootstrap-server kafka1:9092 | head -n 50
```

The command shows a JSON-like listing, but for a tabular report `kafka-metadata-quorum.sh` is more convenient:

```bash
/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server kafka1:9092 describe --replicas
```

The `LEADER` column shows the current quorum leader, `REPLICAS` shows all active nodes. Unresponsive controllers drop from the list.

> [!TIP]
> The `lastCaughtUpTime` field in `describe --status` shows how far a controller lags behind the leader. `0` or fresh `now` — healthy. Lag in minutes — reason to check GC and network latency.

## Topics and Partition Leaders

The cluster can be alive but without partition leaders — producers won't write, consumers won't read. So the next step is topics and their leaders.

List all topics with partition and replica counts:

```bash
/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal
```

The output table contains `Leader`, `Replicas`, `Isr`. If `Leader` equals `-1`, the partition has no active leader — this is an emergency, producers will receive `NotLeaderForPartitionException`.

To get only "bad" partitions in one pipeline:

```bash
/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '$5 == -1 || $5 == "none" {print}'
```

To view leaders for a specific topic:

```bash
/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --topic orders.events
```

## ISR and Unavailable Replicas

ISR (in-sync replicas) determines write reliability. Recommended setting is `min.insync.replicas >= 2` for critical topics. Check for out-of-sync:

```bash
/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --under-replicated-partitions
```

The command returns partitions where `Isr` is smaller than `Replicas`. Empty output — good. Any lines in the list — incident.

Full picture across all partitions with problem filtering:

```bash
/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '{
    replicas=$5; isr=$7;
    # Replicas and Isr are comma-separated id lists; length by commas
    rep_n = split(replicas, a, ",");
    isr_n = split(isr, b, ",");
    if (rep_n != isr_n) print "UNSYNC:", $0;
    if (a[1] == "-1") print "NO_LEADER:", $0;
  }'
```

> [!WARNING]
> `kafka-topics.sh --describe` on a topic with thousands of partitions outputs many lines and loads the controller. In production run with `--partitions N` or filter with awk, otherwise the health check itself becomes a problem.

Additionally — state of a specific replica on the broker side. If you suspect one of the disks is lagging:

```bash
/opt/kafka/bin/kafka-log-dirs.sh \
  --bootstrap-server kafka1:9092 \
  --describe --broker-list 1,3,5 \
  | jq '.[] | .logDirs[] | {broker: .broker, dir: .dir, partitions: (.partitions | length)}'
```

In the output check the `partition.error` field — if non-empty, the replica has issues (offline log dir, disk full, fs in read-only).

## Quick Diagnostics in One Command

For daily rounds it's convenient to bundle all checks into a single script with clear exit codes. Below is a minimal status indicator in bash.

```bash
#!/usr/bin/env bash
set -u
BOOTSTRAP="${BOOTSTRAP:-kafka1:9092}"
KAFKA_BIN="${KAFKA_BIN:-/opt/kafka/bin}"
fail=0

echo "== Quorum status =="
if ! $KAFKA_BIN/kafka-metadata-quorum.sh --bootstrap-server "$BOOTSTRAP" \
    describe --status 2>&1 | grep -q 'isLeader: true'; then
  echo "WARN: quorum leader not confirmed"; fail=1
fi

echo "== Brokers reachability =="
for h in $(echo "$BOOTSTRAP" | tr ',' ' '); do
  if ! timeout 5 bash -c "echo > /dev/tcp/${h%:*}/${h##*:}"; then
    echo "FAIL: $h unreachable"; fail=1
  fi
done

echo "== Under-replicated partitions =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --under-replicated-partitions 2>/dev/null)
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: 0 under-replicated partitions"
fi

echo "== Partitions without leader =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --exclude-internal 2>/dev/null \
      | awk '$5 == -1 {print}')
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: every partition has a leader"
fi

exit $fail
```

Save as `kafka-health.sh`, make executable, and wrap in cron or a systemd timer every 60 seconds:

```bash
chmod +x kafka-health.sh
*/1 * * * * /usr/local/bin/kafka-health.sh \
  >> /var/log/kafka-health.log 2>&1
```

> [!TIP]
> For alerting, replace `echo` blocks with `logger -p local0.err` and configure rsyslog to your SIEM. No need for Prometheus node_exporter — events fly into the common channel.

Typical errors on first run and how to interpret them:

| Symptom | Probable Cause | Action |
|---|---|---|
| `Connection to node -1 could not be established` | Broker not registered in cluster but process is alive | Check `advertised.listeners`, `node.id`, network ACL |
| `isLeader: false` for all controllers | Quorum lost, controllers < `__.min.insync.replicas` | Check `controller.quorum.voters` and disk state on controllers |
| `--under-replicated-partitions` shows entries for >5 min | Broker lagging, slow disk or GC pauses | Take `jstack`, check `iostat -x` and `LogFlushRate` metrics |
| `kafka-log-dirs.sh` returns `LogDirOffline` | Disk full or failed | Free space, check FS for read-only, restart broker |
| Empty `--describe` output on a running cluster | Wrong bootstrap address passed | Verify `advertised.listeners` and DNS |

The five command sets above cover 90% of operational "is the cluster alive?" questions. If everything is green — no need to dig deeper; if something is red — `kafka-log-dirs.sh` and `kafka-metadata-quorum.sh describe --status` will show which direction to go.
