Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

Posts

Practical notes on infrastructure, operations, and everyday tooling.

1 - fixing memory leaks in python services: diagnostics and dump collection

Python services under load can silently consume memory until the cgroup limit triggers an OOM kill. Without systematic dump collection and introspection, root-cause analysis devolves into hypothesis spinning. Below is a practical set of commands and scripts for diagnostics, heap dump collection, and leak mitigation.

1. Memory Consumption Diagnosis Commands

Basic level — psutil. Installed in one line and works without process restart.

pip install psutil

Current process consumption:

import psutil, os
p = psutil.Process(os.getpid())
info = p.memory_info()
print(f"RSS: {info.rss}  VMS: {info.vms}")

Extended structure in one command:

full = p.memory_full_info()
print(full)
# Attributes: python, rss, vms, shared, text, lib, data, dt

Key attributes psutil.Process.memory_full_info() quick reference table:

AttributeDescription
pythonMemory allocated inside the Python interpreter
rssResident Set Size — physical memory in RAM
vmsVirtual Memory Size — virtual address space
sharedShared memory (shared libraries, mmap)
text, lib, dataELF code, library, data segments

For tracking growth dynamics over a short interval, psutil in a loop or watch -n 1 psutil ... can be used, but in production tracemalloc, built into CPython, is more common.

import tracemalloc
tracemalloc.start()
# ... service work ...
snapshot = tracemalloc.take_snapshot()
top_stats = snapshot.statistics('lineno')
for stat in top_stats[:10]:
    print(stat)
tracemalloc adds a small overhead (~1–2 %). Enable it only on a staging environment or when explicit leak suspicions exist.

2. Tools for Tracking Growth and Dump Collection

When RSS begins creeping upward unnoticed, deeper inspection is required. The toolset depends on debugger availability and ptrace permissions.

gdb + gcore — classic method to extract a full process dump without stopping it (provided coredump is enabled).

# 1. Find the PID
pgrep -f your_service

# 2. Attach gdb and dump the core
gdb -p <PID>
# inside gdb:
(gdb) gcore /tmp/heap_dump.core
(gdb) quit

The resulting gcore file is binary and can be analyzed locally:

# Show segment info
gdb -ex "info files" -ex "quit" /tmp/heap_dump.core
# Or via gdb python plugins (see further)

objgraph — quick answer to “who is holding this object”.

pip install objgraph
import objgraph
# Most frequent object types in memory
objgraph.most_common_types(limit=20)
# Find reference chains
objgraph.find_backref_chains(some_object, 'owner')
objgraph works with Python objects only. For native C extensions or ctypes, gdb or valgrind are required.

tracemalloc + heapdump — Python 3.4+ native mechanism can save a heap snapshot to a file.

import tracemalloc, sys
tracemalloc.start()
# ... ...
snapshot = tracemalloc.take_snapshot()
snapshot.dump('/tmp/tracemalloc.dump')

The dump can be opened in a visualizer, but for deep analysis gdb remains more convenient.

If the service runs in a container with cap_sys_ptrace restrictions, gcore collection may require host-level access or kubectl exec with appropriate capabilities.

3. Common Anti-patterns and What to Avoid

Anti-patternConsequenceRecommendation
Ignoring gc.collect() as a “magic button”False sense of security, leak persistsUse gc.collect() to clear temporary references, not as a logical leak cure
Relying solely on __del__ for resource releaseIrregular release, circular referencesPrefer contextlib.contextmanager, try/finally, or weakref
No memory limits (cgroup/ulimit)Sharp OOM-kill without dump, context lossAlways set memory.limit in manifests and check ulimit -v
Caches without TTL or unbounded growthRapid RSS increaseUse functools.lru_cache(maxsize=N) or external stores with expiration
Accumulation of objects in global lists/modules“Death” memory in long-running processesPeriodically check list lengths, move cleanup to periodic tasks
One of the most insidious patterns — del obj inside __del__ with circular references. Python’s GC eventually collects them, but the order is not guaranteed, causing memory peaks between collections.

4. OOM Incident Response Scenario

  1. Confirmation — check the cluster event: kubectl get events -n <ns> | grep OOM or dmesg | grep out of memory. Verify the process terminated with code 137.
  2. Quick RSS glance — if the service is still alive, run one command:
python3 -c "import psutil, os; p=psutil.Process(os.getpid()); print(p.memory_info().rss // 1024, 'KB')"
  1. Dump collection — choose the method based on tool availability:
    • If gdb is available with rights: gdb -p <PID> -batch -ex gcore -ex quit /tmp/oom_dump.core
    • If a quick snapshot is needed: python3 -c "import tracemalloc; tracemalloc.start(); ..." followed by tracemalloc.take_snapshot().statistics('traceback')
  2. Analysis — use objgraph.most_common_types() or gdb commands info files, bt for call stacks. Look for unexpected dict, list, or domain class object counts.
  3. Fix — eliminate the root cause: add maxsize to caches, remove circular imports, replace global lists with bounded loops, adopt weakref for observers.
  4. Prevention — add an alert at rss > 80% of limit in Prometheus/Grafana and document the dump-collection command in the runbook.

Completing diagnostics at this stage either restores the service or gathers sufficient data for a tracker ticket with specific object types and call stacks. Here’s a thinking process:

  1. Analyze the Request:
  • Topic: Fixing memory leaks in Python services: diagnostics and dump collection
  • Role: Practicing Lead DevOps
  • Format: IT notes for a Hugo blog (Mana theme)
  • Style: Short paragraphs, concrete commands, flag tables, code with fences, no fluff, no marketing, no emojis, no author signature at end
  • Allowed: GitHub admonitions (> [!NOTE], > [!TIP], > [!WARNING])
  • Prohibited: Made-up flags, URLs, versions; uncertain facts → cautious tone; no YAML/TOML front matter; no wrapping in ```; start with a lid (2-4 sentences); sections with ##; practical commands, tables if needed; 800-1600 words; end on last substantive section; don’t repeat rules, don’t write “User wants”, "
  • Key constraint: Write the English article as a parallel original, not a word-for-word translation. Same structure and facts as the Russian draft.

2 - SSH key best practices

SSH keys are the de facto standard for authenticating to infrastructure, but poor management turns every deployment into a potential vulnerability. This note collects proven practices: from key generation to revocation and rotation without service downtime.

Introduction

SSH keys function as long‑lived credentials, and their lifecycle directly impacts supply‑chain security. Unlike passwords, keys are often created once and forgotten, leading to an accumulation of “zombie” keys with privileges that exceed current needs. Proper generation, binding to an agent, and regular rotation minimize the attack surface and enable auditing of changes. The following sections describe concrete commands and configurations used in operations.

Key Generation

Modern OpenSSH defaults recommend the Ed25519 algorithm because of its short key length (256 bits) and high resistance to attacks. RSA still appears in legacy systems but requires increasing the length to at least 4096 bits for acceptable security.

The primary generation command is:

ssh-keygen -t ed25519 -C "user@infrastructure" -f ~/.ssh/id_ed25519

Flag explanations:

  • -t – cryptographic algorithm type (ed25519, rsa, ecdsa, sk-ed25519@openssh.com for Security Key).
  • -C – comment, usually an email or service name; it is stored in the key but not used for authentication.
  • -f – path to the key file; if the directory does not exist, it will be created with correct permissions.

For fully automated scenarios (CI/CD, scripts) you can pass an empty passphrase via -N "", but understand the risk: a passwordless key can be used by any process running as the user. It is preferable to use ssh-agent with an unlock session or a key store (Keychain, ssh-agent -s).

Table: Comparison of common key types

AlgorithmKey lengthPrimary useNotes
Ed25519256 bitsModern infrastructures, deployment scriptsRecommended standard starting with OpenSSH 6.8
RSA2048–4096 bitsLegacy systems, some vendor platforms4096 bits minimally acceptable; larger is overkill
ECDSA256–521 bitsSpecific requirements, older devices521‑bit (P‑521) most secure but less compatible

When generating keys for different services, it is convenient to split them by files (e.g., id_ed25519_work, id_ed25519_cloud) to simplify revocation and rotation without affecting other contexts.

Using ssh-agent

The SSH agent holds the decrypted key in memory of the current session, avoiding password entry on each connection. Starting the agent and adding a key:

# Start agent (output environment variables)
eval $(ssh-agent -s)

# Add key to agent
ssh-add ~/.ssh/id_ed25519

In newer OpenSSH versions (8.2+) the agent can be started in config mode:

ssh-agent -c

The ~/.ssh/config file allows automatically loading keys for specific hosts:

Host *.github.com
    AddKeysToAgent yes
    UseKeychain yes
    IdentityFile ~/.ssh/id_ed25519_github

Host *.gitlab.com
    AddKeysToAgent yes
    IdentityFile ~/.ssh/id_ed25519_gitlab

Flags in IdentityFile and AddKeysToAgent guarantee that the key is loaded once on the first connection to a host and remains available for subsequent sessions within the agent’s lifetime. On macOS the flag UseKeychain yes integrates the key with Keychain, locking it when the screen is locked and unlocking when unlocked.

Tip

Enable IdentitiesOnly yes in the config if the SSH client begins trying all keys during a connection attempt, which can cause delays or authentication errors on servers with a limited list of allowed keys.

Revocation and Rotation

Keys should not live forever. Rotation policy depends on the infrastructure risk profile, but minimal hygiene includes regular audits and a revocation procedure when personnel change or responsibility shifts.

Revocation procedure:

  1. Remove the public key from the remote server:

    ssh-keygen -f ~/.ssh/known_hosts -R "hostname_or_ip"

    or manually edit ~/.ssh/authorized_keys deleting the line corresponding to the revoked key.

  2. Locally delete the key and generate a new one:

    shred -u ~/.ssh/id_ed25519  # secure deletion (if util-linux installed)
    rm -rf ~/.ssh/id_ed25519*
    ssh-keygen -t ed25519 -C "user@infrastructure" -f ~/.ssh/id_ed25519
  3. Distribute the new public key via IaC or a configuration manager (Ansible, Terraform) to avoid manual edits.

Automated rotation:

For large fleets you can use scripts that check the key creation date (comment -C or metadata inside the key) and automatically revoke keys older than N days. It is important to retain key history for at least 30–90 days before full deletion to avoid breaking services that have not yet updated their configuration.

Warning

If a key leak is detected (e.g., in public repositories or logs), act immediately: revoke the key on all hosts, generate a new one, and update the configuration. Do not wait for scheduled rotation—exposure of a key means full loss of control over the associated resources. We need to produce the English version of the article, matching the structure and facts of the Russian draft. The Russian draft sections: Introduction, Генерация ключей (Key generation), Использование ssh-agent (Using ssh-agent), Отзыв и ротация (Revocation and rotation). We need to produce English article, same structure, same commands, tables, admonitions. No front matter, no code fences around whole article. Just markdown body.

We need to translate the Russian content into English, preserving the same details, commands, tables, admonitions. Ensure we keep the same practical tone, short paragraphs, specific commands, tables, code fences with language spec. Use GitHub admonitions as in Russian: > [!TIP], > [!WARNING]. Also maybe > [!NOTE] if needed but Russian didn’t have note. We’ll keep tip and warning.

We need to produce

3 - kubectl cheatsheet: essential commands for Kubernetes

kubectl is the primary interface to a Kubernetes cluster. This cheatheet covers routine operations — from context setup to pod debugging and namespace switching. All commands are verified against current kubectl versions (1.28+).

Installation and Context Setup

Installation depends on your OS. On Linux, use the package manager or the binary directly:

curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl"
chmod +x kubectl && sudo mv kubectl /usr/local/bin/
kubectl version --client

Configuration lives in ~/.kube/config. A context defines the cluster, user, and default namespace.

kubectl config current-context
kubectl config use-context production-cluster
kubectl config view --minify   # show current context
kubectl config get-contexts    # list all contexts
Tip

If you work with multiple clusters, store configs in the KUBECONFIG variable separated by colons: export KUBECONFIG=~/.kube/config:~/.kube/prod-config.

Working with Pods

Basic operations:

kubectl get pods                          # all pods in the current namespace
kubectl get pods -A                       # all namespaces
kubectl describe pod <pod-name>           # details and events
kubectl logs <pod-name>                   # container logs
kubectl logs <pod-name> -c <container>    # logs for a specific container in a multi-container pod
kubectl exec -it <pod-name> -- /bin/sh    # enter a pod
kubectl delete pod <pod-name>             # delete a pod

Deployments, StatefulSets, and DaemonSets

kubectl get deployments,statefulsets,daemonsets
kubectl rollout status deployment/<name>       # update status
kubectl rollout history deployment/<name>      # revision history
kubectl rollout undo deployment/<name>         # rollback to previous revision
kubectl rollout undo deployment/<name> --to-revision=2   # rollback to a specific revision
kubectl scale deployment/<name> --replicas=5   # scale replicas

For StatefulSets, startup order and stable identities are critical. For DaemonSets, a guaranteed instance runs on every node.

Warning

Do not run kubectl delete on a Deployment without checking kubectl get deployment first. Deleting the controller does not automatically remove pods — they will be recreated unless you use --cascade=orphan (in older versions) or delete through kubectl delete deployment.

Services, Ingress, and ConfigMap

kubectl get svc                      # services
kubectl expose deployment/<name> --port=80 --type=NodePort   # create svc from a deployment
kubectl get ingress                  # ingress resources
kubectl describe ingress <name>      # rules and events
kubectl apply -f ingress.yaml        # apply ingress from a file

ConfigMap and Secret for configuration:

kubectl create configmap app-config --from-file=config.yaml
kubectl get configmap app-config -o yaml
kubectl create secret generic db-creds --from-literal=password='s3cr3t'
kubectl get secret db-creds -o jsonpath='{.data.password}' | base64 -d

Debugging and Diagnostics

When a pod is not working, follow this sequence:

kubectl get pods -o wide             # status and node
kubectl describe pod <pod-name>      # events, reason for CrashLoopBackOff
kubectl logs <pod-name> --previous   # logs from a crashed container
kubectl top pod <pod-name>           # CPU/memory consumption (requires metrics-server)
kubectl attach -it <pod-name> -- /bin/sh   # alternative to exec

To test network connectivity from inside the cluster:

kubectl run debug-pod --image=busybox --rm -it --restart=Never -- wget -O- http://<svc-name>.<namespace>.svc.cluster.local:80
Note

If kubectl top returns an error, metrics-server is not installed. Installation depends on your provider: kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml.

Labels, Selectors, and Output Formatting

Labels let you group resources and select them in bulk:

kubectl get pods --show-labels
kubectl label pods <pod-name> app=frontend tier=web
kubectl get pods -l app=frontend       # label selector
kubectl get pods -l 'app in (frontend,backend)'
kubectl get pods -l '!tier=web'        # negative selector

Output formatting:

kubectl get pods -o wide
kubectl get pods -o json               # full JSON
kubectl get pods -o jsonpath='{.items[*].metadata.name}'
kubectl get pods -o custom-columns=NAME:.metadata.name,STATUS:.status.phase
FlagDescription
-n, --namespaceSpecify namespace
-o, --outputFormat: json, yaml, wide, custom-columns
-l, --selectorLabel selector
-f, --filenameConfiguration file (YAML/JSON)
--show-labelsShow labels column
-w, --watchStreaming update mode
--all-namespaces, -AAll namespaces

Managing Namespaces and Switching Contexts

kubectl get namespace
kubectl create namespace staging
kubectl delete namespace staging        # deletion is asynchronous
kubectl config set-context --current --namespace=staging   # default namespace in context

For convenience, create a shell alias or function:

# ~/.bashrc or ~/.zshrc
kswitch() { kubectl config use-context "$1" && kubectl config set-context --current --namespace="$2"; }
# usage: kswitch production-cluster default
Tip

kubectl config rename-context, kubectl config unset, and kubectl config set let you edit the config without manually editing YAML. Verify the result with kubectl config view.

kubectl is more than a CLI — it is an abstraction layer over the Kubernetes API. Knowing formatting flags and selectors cuts diagnostic time significantly. Keep this cheatheet handy and update it as new versions ship.

4 - Creating User and Role in Kubernetes and Binding Them via RBAC

In Kubernetes there are no “users” in the traditional sense — there are ServiceAccounts and certificates, bound to roles through RBAC. Without proper configuration, anyone holding a kubeconfig gets full access to the cluster. Below is the complete cycle: create a ServiceAccount, define permissions, bind them, and verify.

Creating ServiceAccount and Generating kubeconfig

Start by creating a ServiceAccount in the target namespace:

kubectl create serviceaccount devops-sa -n staging

To generate a kubeconfig, extract the token from secrets and build the config file:

SA_NAME=devops-sa
SA_NS=staging

TOKEN=$(kubectl -n $SA_NS get secret $(kubectl -n $SA_NS get sa $SA_NAME -o jsonpath='{.secrets[0].name}') -o jsonpath='{.data.token}' | base64 --decode)

kubectl config set-credentials $SA_NAME --token=$TOKEN
kubectl config set-context devops-staging --cluster=$(kubectl config current-context) --user=$SA_NAME --namespace=$SA_NS
kubectl config use-context devops-staging
Warning

A ServiceAccount token is a plain bearer token. If the kubeconfig leaks, an attacker gains access with that SA’s privileges. Store the file with 600 permissions.

Defining Role and ClusterRole

A Role is namespace-scoped; a ClusterRole is cluster-scoped. The difference is critical: a Role grants no rights outside its namespace, a ClusterRole does.

Example Role allowing read and create operations on pods in the staging namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  namespace: staging
  name: pod-reader-writer
rules:
  - apiGroups: [""]
    resources: ["pods"]
    verbs: ["get", "list", "watch", "create", "delete"]

Example ClusterRole for managing ingresses across the entire cluster:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ingress-manager
rules:
  - apiGroups: ["networking.k8s.io"]
    resources: ["ingresses"]
    verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
Tip

An empty string "" in apiGroups refers to the core API group (v1). For apps, networking, batch — specify the corresponding groups. The full list of groups is available via kubectl api-resources.

Creating RoleBinding and ClusterRoleBinding

A RoleBinding binds a Role to a subject within a namespace. A ClusterRoleBinding binds a ClusterRole cluster-wide.

Binding a Role to a ServiceAccount in the staging namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: devops-pod-access
  namespace: staging
subjects:
  - kind: ServiceAccount
    name: devops-sa
    namespace: staging
roleRef:
  kind: Role
  name: pod-reader-writer
  apiGroup: rbac.authorization.k8s.io

Binding a ClusterRole to the same SA (now with cluster-wide ingress rights):

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: devops-ingress-access
subjects:
  - kind: ServiceAccount
    name: devops-sa
    namespace: staging
roleRef:
  kind: ClusterRole
  name: ingress-manager
  apiGroup: rbac.authorization.k8s.io
BindingScopeRole typeWhen to use
RoleBindingSingle namespaceRole or ClusterRoleRead/write in a specific ns
ClusterRoleBindingEntire clusterClusterRoleGlobal rights (node, pv, dns)

Verifying and Debugging Access Rights

After applying the YAML, confirm the SA actually received the intended permissions:

kubectl auth can-i get pods --as=system:serviceaccount:staging:devops-sa -n staging
kubectl auth can-i create pods --as=system:serviceaccount:staging:devops-sa -n staging
kubectl auth can-i get ingresses --as=system:serviceaccount:staging:devops-sa

The last command checks via ClusterRoleBinding — it returns yes if the binding is correct.

If something doesn’t work, check the audit log or use kubectl auth reconcile with --dry-run=server for a preview:

kubectl auth reconcile role-binding.yaml --dry-run=server

Another useful trick is to find which RoleBindings are attached to a specific SA:

kubectl get rolebindings,clusterrolebindings --all-namespaces -o json | \
  jq '.items[] | select(.subjects[]?.name=="devops-sa") | {name: .metadata.name, kind: .kind, namespace: .metadata.namespace}'
Note

kubectl auth can-i checks permissions only — it does not account for NetworkPolicy or PodSecurityPolicy. If a pod fails to start, the issue may lie there.

Bottom line: RBAC in Kubernetes works as a chain — SA → Role/ClusterRole → Binding. Each link can be verified independently, which greatly simplifies debugging. Start with minimum privileges and expand as needed; never grant cluster-admin without a compelling reason.

5 - Git Tag: Marking Releases and Bookmarks in History

Git Tag: Marking Releases and Bookmarks in History

Tags are Git’s mechanism for assigning meaningful labels to specific commits. Unlike branches, tags don’t move — they’re pinned to a commit and serve as anchors for releases, versions, and critical checkpoints. Without tags, release history devolves into a hash search — and that’s a direct path to deployment errors.


Types of Tags: Lightweight vs Annotated

There are two types of tags. Lightweight is just a name attached to a commit, with no additional metadata. Annotated is a full Git object with author, date, message, and signing capability.

Warning

Always use annotated tags for releases. Lightweight tags carry no metadata and can’t be signed — you lose context during audits or rollbacks.

PropertyLightweightAnnotated
Object in DBNoYes (tag object)
MessageNoYes
Author/DateNoYes
GPG SignatureNoYes
Creation SpeedFasterSlightly slower

Creating, Viewing, and Deleting Tags

Create an annotated tag:

git tag -a v1.2.0 -m "Release 1.2.0: API stabilization"

Create a lightweight tag:

git tag v1.2.0-rc1

List all tags:

git tag

Show details of a specific annotated tag:

git show v1.2.0

Delete a local tag:

git tag -d v1.2.0-rc1

List tags matching a pattern:

git tag -l "v1.*"
Tip

The -l flag supports glob patterns. This is faster than piping git tag through grep.


Pushing Tags to Remote

By default, git push does not send tags. This is a common reason why teammates can’t see a release tag on the remote repository.

Push a single tag:

git push origin v1.2.0

Push all local tags:

git push origin --tags

Delete a tag on remote:

git push origin --delete v1.2.0

Local deletion and remote deletion are two separate operations. Forgetting the second is a typical mistake during a release rollback.


GPG-Signing Tags

Annotated tags can be signed with a GPG key. This guarantees the tag was created by a specific author and hasn’t been tampered with.

Ensure your GPG key is configured first:

gpg --list-secret-keys --keyid-format=long
git config --global user.signingkey <KEY_ID>
git config --global gpg.format openpgp

Create a signed tag:

git tag -s v1.2.0 -m "Signed release 1.2.0"

Verify the signature:

git tag -v v1.2.0
Note

If git tag -v returns “no signature found”, the tag isn’t signed. If it says “Good signature from…”, the signature is valid. Make sure the author’s public key is available in your keyring.


Working with Tags in CI/CD Pipelines

In pipelines, tags are the primary trigger for releases. Most CI/CD systems allow filtering by branches and tags.

Example for GitHub Actions:

on:
  push:
    tags:
      - 'v*'

jobs:
  release:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-tags: true
      - run: echo "Building for $(git describe --tags)"

Example for GitLab CI:

deploy:
  script:
    - echo "Deploying tag $CI_COMMIT_TAG"
  rules:
    - if: '$CI_COMMIT_TAG =~ /^v\d+\.\d+\.\d+$/'

Useful commands inside a pipeline:

# Current tag if the commit is tagged
git describe --tags --exact-match HEAD

# Last tag before the current commit
git describe --tags --abbrev=0 HEAD

# All tags sorted by creation date (newest first)
git tag --sort=-creatordate
Tip

Always use fetch-tags: true in the checkout step. Without it, the pipeline may not see tags and the $CI_COMMIT_TAG filter will break.


Tags are a minimal tool with maximum impact. Annotated with GPG signatures, --tags on release push, and v* filtering in CI — that set is enough to keep release history readable and verifiable.

6 - Deploying with a post-receive Git Hook

Deploying through CI is great, but sometimes you just need to push code to a server with a single git push. The post-receive hook in a bare repository handles this without extra dependencies: push to the server, and the hook automatically checks out files into the working directory.

Schema: bare repo as a deploy trigger

The logic is straightforward:

  1. A bare repository is created on the server (e.g., /srv/deploy/app.git).
  2. The developer adds it as a remote and runs git push origin main.
  3. Git receives the data and triggers hooks/post-receive.
  4. The hook script runs git --work-tree=/var/www/app --git-dir=/srv/deploy/app.git checkout -f main.
Note

A bare repository has no working directory. That is why post-receive must explicitly pass --work-tree so checkout knows where to write files.

This is not CI — it is a direct trigger at the Git level. No pipelines, artifacts, or queues. It fits small services, VPS instances, and internal tools well.

Creating a bare repository on the server

SSH into the server and create the repo:

mkdir -p /srv/deploy
cd /srv/deploy
git init --bare app.git

The resulting structure is standard: hooks/, objects/, refs/, HEAD, config. Default hooks live in /srv/deploy/app.git/hooks/ with a .sample suffix — rename them or create your own.

Warning

Make sure /srv/deploy is owned by the user you are pushing as. Otherwise, write permissions on objects/ will be denied.

Writing the post-receive hook

Create /srv/deploy/app.git/hooks/post-receive:

#!/usr/bin/env bash
set -euo pipefail

REPO_DIR="/srv/deploy/app.git"
WORK_TREE="/var/www/app"
BRANCH="main"

while read oldrev newrev refname; do
  if [ "$refname" = "refs/heads/$BRANCH" ]; then
    echo "Deploying $BRANCH to $WORK_TREE..."
    git --work-tree="$WORK_TREE" --git-dir="$REPO_DIR" checkout -f "$BRANCH"
    echo "Deployment complete."
  fi
done

Key elements:

ElementPurpose
set -euo pipefailScript exits on any error; does not continue with bad data
while read oldrev newrev refnameThe hook passes three arguments per line for each updated ref
refs/heads/$BRANCHChecks that the pushed ref is the target branch
checkout -fForce-synchronizes the working tree with the reference
Tip

If you need to restart a service after deployment, append systemctl restart app or supervisorctl reload app to the end of the script. The hook runs in the context of the Git user, so verify it has permission to restart.

Setting up permissions and access

The most common issue is permissions. The Git user you push as must be able to write to WORK_TREE and execute commands from the hook.

Setup options:

  • Single user: Git user and web user are the same identity. Simply chown -R deploy:deploy /srv/deploy /var/www/app.
  • Group: Add the Git user to the web server group. usermod -aG www-data deploy, then chmod -R g+w /var/www/app.
  • SSH key: Push over SSH; the key authenticates as the correct user.
# Example: deploy user, web root /var/www/app
sudo useradd -m deploy
sudo mkdir -p /var/www/app
sudo chown -R deploy:deploy /var/www/app
sudo chmod -R 755 /var/www/app

Do not forget to make the hook executable:

chmod +x /srv/deploy/app.git/hooks/post-receive

Pushing from the client and verifying

On the developer machine, add the remote and push:

git remote add production deploy@server:/srv/deploy/app.git
git push production main

On the server, the hook log will show:

Deploying main to /var/www/app...
Deployment complete.

Verify files on the server:

ls -la /var/www/app
git --git-dir=/srv/deploy/app.git --work-tree=/var/www/app status
Warning

If the push fails with Permission denied, check the owner of hooks/post-receive and the target directory. If checkout does not write files, confirm that WORK_TREE points to an existing directory and the user has write access to it.

That is it. No additional tools — just Git and shell. For high-load production, this is obviously not a replacement for a proper CI pipeline, but for quick deploys to one or two servers it works reliably without unnecessary complexity.

7 - ulimit and systemd LimitNOFILE — why ulimit -n inside a unit doesn't stick

What is nofile and where it lives

nofile is the maximum number of open file descriptors per process. That’s not just regular files — it covers sockets, pipes, stdin/stdout/stderr, logs shipped through journald — everything counts. When nginx or a Go app crashes with too many open files, this is the limit to blame.

Limits live at three levels:

LevelWhere to checkWhat it controls
Kernel (system-wide)/proc/sys/fs/file-max, /proc/sys/fs/nr_openAbsolute ceiling for the whole system
PAM / login/etc/security/limits.conf, /etc/security/limits.d/For sessions via pam_limits.so
systemdLimitNOFILE= in unit, DefaultLimitNOFILE= in system.confFor systemd-managed services
Note

/proc/sys/fs/nr_open is the upper bound you can raise nofile to for a single process. It defaults to 1073741816 (≈1B) on most distros, but in practice you rarely need more than 1048576.

LimitNOFILE in a systemd unit

In a unit file the directive looks like this:

[Service]
LimitNOFILE=65536

You can set both soft and hard limits at once, separated by a space:

LimitNOFILE=65536:1048576

The first value is the soft limit, the second is the hard limit. If you specify only one, it becomes the soft limit and the hard limit is taken from the system maximum.

Tip

The global default for all units is DefaultLimitNOFILE= in /etc/systemd/system.conf (and user.conf). Modern distros often default to 1048576, but older ones may have 4096 or 1024 — that’s exactly what catches you off guard.

Why ulimit -n in ExecStart doesn’t work

The typical mistake is trying to set the limit directly in the launch command:

[Service]
ExecStart=/bin/sh -c 'ulimit -n 65536 && exec /usr/bin/myapp'

ulimit -n inside ExecStart does not take effect on the process systemd tracks as MainPID. systemd sets limits before ExecStart runs, and the shell wrapper already operates in a context where the limit is fixed. Worse, ulimit -n may fail with Operation not permitted if the requested value exceeds the hard limit systemd assigned.

If you need to change the limit, use the [Service] directive — not a shell wrapper.

Warning

ulimit -n in a shell script called from ExecStartPre also does not propagate the limit to the main process. Limits are a property of the process, not the shell session.

How to check applied limits

Three ways, from simple to authoritative:

# 1. From inside the process
cat /proc/self/limits | grep "Max open files"

# 2. For a specific PID
cat /proc/<pid>/limits | grep "Max open files"

# 3. What systemd sees for the unit
systemctl show myservice.service -p LimitNOFILE

systemctl show reports exactly what systemd applied at fork() — that’s the authoritative source. /proc/<pid>/limits is what the kernel sees for the process. If they disagree, the problem is in an intermediate layer (PAM, container runtime, sudo).

For debugging ExecStart, you can add:

[Service]
ExecStartPre=/bin/sh -c 'cat /proc/self/limits | grep "Max open files"'

This shows the limits before the main process starts — what systemd set for the service.

PAM limits vs systemd limits

On most modern distros (RHEL 8+, Ubuntu 20.04+, Debian 11+) systemd does not invoke pam_limits.so for system services. Limits from /etc/security/limits.conf are not applied to unit files. They only work for login sessions (SSH, local login, su).

MechanismApplies toConfigured in
LimitNOFILE= in unitSpecific systemd service/etc/systemd/system/*.service
DefaultLimitNOFILE=All systemd services/etc/systemd/system.conf
pam_limits.so / limits.confUser login sessions/etc/security/limits.conf
ulimit in shellCurrent shell and childrenInteractive session

If a service isn’t started through systemd (e.g., via supervisor, docker --ulimit, or directly), LimitNOFILE in the unit file doesn’t apply at all. In Docker that’s --ulimit nofile=65536:1048576, in supervisor the stdout_maxbytes and similar options aren’t related to nofile directly — you configure it at the OS or container level.

Tip

After changing LimitNOFILE in a unit file, always run systemctl daemon-reload and systemctl restart <unit>. systemctl reload doesn’t restart the process and doesn’t reapply limits.

8 - coredumpctl: Finding a Binary Crash

What is coredumpctl and how it works

When a binary crashes with SEGV, the kernel can save a core dump — a snapshot of process memory at the moment of the crash. In systemd-based distributions, coredumpctl handles collecting, storing, and searching these dumps. It’s a wrapper around systemd-coredump, which stores dumps in /var/lib/systemd/coredump/ and indexes metadata through journald.

Note

Requires systemd-coredump and an active journald. In minimal containers without systemd, this tool is unavailable.

Installation is trivial for most distributions:

# Debian/Ubuntu
sudo apt install systemd-coredump

# RHEL/Fedora
sudo dnf install systemd-coredump

After installation, verify that kernel.core_pattern points to a pipe into systemd-coredump:

cat /proc/sys/kernel/core_pattern
# expected: |/usr/lib/systemd/systemd-coredump %P %u %g %s %t %c %h %e

If it shows core.%e.%p, dumps are written to files and coredumpctl won’t see them.

Listing dumps: coredumpctl list

The base command prints all saved crashes:

coredumpctl list

Output contains columns: MESSAGE ID, TIMESTAMP, PID, UID, GID, COMM, EXE, COREFILE.

Tip

Use --no-pager for scripts, --json=short for parsing, -n for only the latest entry.

Key flags for list:

FlagPurpose
-n NLast N entries
-1Only the latest single entry
--since TIMESTAMPStarting from date
--until TIMESTAMPUntil date
-u USERBy UID
--exe PATTERNBy binary path
--debugShow service dumps

Crash details: coredumpctl info

For a single dump, info gives everything needed before pulling the file:

coredumpctl info 1234

Where 1234 is the process PID or number from list. Output includes:

  • the signal that caused the crash (SIGSEGV, SIGABRT, etc.)
  • timestamp
  • executable path
  • core file size
  • MESSAGE ID — correlation with journald
# Example output (key lines)
         PID: 1234 (myapp)
     UID:GID: 1000:1000
      Signal: 11 (SIGSEGV)
    Timestamp: Mon 2025-01-06 14:23:01 UTC
     Command Line: /opt/myapp/bin/myapp --config prod.yaml
     Executable: /opt/myapp/bin/myapp
       Core File: /var/lib/systemd/coredump/myapp.1234.abc123.core

Extracting the core file: coredumpctl dump

The extracted file can be piped directly into gdb or saved to disk:

# Save to current directory
coredumpctl dump 1234 --output=myapp.core

# Pipe directly into gdb
coredumpctl dump 1234 -o - | gdb /opt/myapp/bin/myapp -
Warning

--output=- writes to stdout. If the core file is large (gigabytes), the pipe may block — better to write to disk.

Filtering and searching by binary, PID, time

In production, dumps accumulate by the dozens, and searching by PID is impractical. coredumpctl supports multiple filters at once:

# All crashes of a specific binary
coredumpctl list --exe /opt/myapp/bin/myapp

# Crashes in the last hour
coredumpctl list --since "1 hour ago"

# By specific user and binary
coredumpctl list --exe /usr/bin/python3 --uid 1000

# Only SIGSEGV
coredumpctl list --signal 11

For scripts and automation, JSON output is convenient:

coredumpctl list --json=short | jq '.[] | select(.exe == "/opt/myapp/bin/myapp") | .pid'

Practical examples for debugging crashes

A typical debugging cycle looks like this:

# 1. Find the latest crash of myapp
coredumpctl list --exe /opt/myapp/bin/myapp -n 1

# 2. Check details — what signal and where
coredumpctl info <PID>

# 3. Pull the core and launch gdb
coredumpctl dump <PID> --output=/tmp/myapp.core
gdb /opt/myapp/bin/myapp /tmp/myapp.core

# 4. In gdb:
(gdb) bt full
(gdb) info registers
(gdb) x/16i $pc

If dumps don’t appear, check coredumpctl list and journalctl -u systemd-coredump. A common cause: LimitCORE in the systemd unit is set to 0, or kernel.core_pattern isn’t configured as a pipe.

Tip

For persistent monitoring, add to the unit file:

[Service]
LimitCORE=infinity

then restart the service. After that, coredumpctl will see all crashes.

9 - curl --resolve and SNI: Testing Virtual Hosts Without /etc/hosts

When you need to test a virtual host on a specific IP but don’t want to edit /etc/hosts — whether due to permissions, conflicts with other services, or just the habit of keeping the file clean — curl --resolve solves both problems at once: it overrides DNS resolution and sends the correct SNI in the TLS handshake.

Problem: virtual host without editing /etc/hosts

Multiple virtual hosts can live on a single IP, and the server picks the right one based on the Host header (HTTP/1.1) and SNI (TLS). Without an entry in /etc/hosts, curl first tries to resolve the name through DNS — getting the wrong IP, or no response at all.

Editing /etc/hosts works, but it requires sudo, clutters the file, and can break other services that depend on the same record.

curl –resolve: syntax and example

The --resolve flag intercepts name resolution at the curl level and substitutes it:

curl --resolve HOST:PORT:ADDR URL
ComponentMeaning
HOSTVirtual host name
PORTPort (typically 443 for HTTPS)
ADDRTarget IP address

Example:

curl --resolve example.com:443:203.0.113.50 https://example.com/health

curl sends the request to 203.0.113.50:443, but the HTTP Host header will still be example.com.

SNI in combination with –resolve

Note

--resolve does not change the name curl sends in SNI. SNI is taken from the URL.

This is the key point. If the URL contains https://example.com, curl sends example.com as SNI in the TLS ClientHello — even though the IP is overridden via --resolve. The server uses SNI to select the correct certificate and virtual host.

To verify that SNI is actually being sent, use openssl:

openssl s_client -connect 203.0.113.50:443 -servername example.com </dev/null 2>/dev/null | openssl x509 -noout -subject -issuer

If SNI is empty or wrong, the server returns the default certificate — and curl produces an error like SSL: certificate subject name does not match.

Practical example: checking a virtual host

Suppose several virtual hosts run on server 198.51.100.10, and you need to check app.local without touching /etc/hosts:

curl -v --resolve app.local:443:198.51.100.10 https://app.local/api/status

In the -v output you will see:

  • Connected to 198.51.100.10 (198.51.100.10) port 443 — connection to the correct IP.
  • > Host: app.local — correct Host header.
  • TLS SNI extension: "app.local" — SNI sent correctly.

For bulk-checking multiple hosts on the same IP:

for host in app.local api.local admin.local; do
  echo "=== $host ==="
  curl --resolve "$host:443:198.51.100.10" "https://$host/health" -s -o /dev/null -w "%{http_code}\n"
done
Warning

--resolve applies only to the current curl process. Once the command finishes, the override disappears — /etc/hosts stays untouched.

If you need the override to work for all tools in the terminal, not just curl, then nsupdate or a local DNS resolver (dnsmasq, stubby) is a better fit. But for a one-off virtual host check, --resolve is faster and safer.

10 - Docker logs and journald: choosing a logging driver

When a container crashes, logs are the first thing you need to see. docker logs looks simple, but under the hood different logging drivers are at work, and the choice affects how logs are stored, rotated, and accessed. Here is what you should know before trusting the default.

How docker logs works

The docker logs <container> command reads the container’s stdout/stderr stream and outputs it to the terminal. Behind this sits a logging driver — a component that determines where the data actually goes. By default it is json-file: each container gets a JSON file on the host into which every output line is written.

Note

docker logs does not read logs from inside the container directly — it queries the driver, which already stores the data in its own format and location.

The driver is configured at the Docker daemon level or per container. The choice affects log rotation, access via journalctl, and integration with centralized collection systems.

Driver json-file (default)

json-file is the built-in driver with no external dependencies. Each container creates a file at /var/lib/docker/containers/<container-id>/<container-id>-json.log. The format is JSON lines: each entry contains log, stream (stdout or stderr), and time.

Rotation is controlled by two flags:

FlagDescription
max-sizeMaximum size of a single log file (e.g., 10m)
max-fileNumber of rotated files to retain

Without these flags, log files grow without limits. In production this is a direct path to filling the disk.

docker run --log-driver json-file --log-opt max-size=10m --log-opt max-file=3 nginx
Warning

If max-size and max-file are not explicitly set, Docker does not limit log size. On a host with many containers this will result in unexpected disk exhaustion.

Logs can be read via docker logs or directly from the file path on the host, but the latter is not recommended because files may be held open by the daemon.

Driver journald

journald sends container logs into the systemd journal. This means logs are accessible through journalctl, all of journald’s rotation and compression mechanisms apply, and there are no separate JSON files growing on disk.

This requires systemd and the systemd-journal-remote package (on some distributions). The container must be started with the driver specified:

docker run --log-driver journald --log-opt tag={{.Name}} nginx

The tag flag sets the identifier in journald — without it the tag is an empty string and finding the right container becomes difficult. The template {{.Name}} substitutes the container name.

Tip

Use tag={{.Name}} or tag={{.ID}} so that logs in journald are immediately tied to a specific container. Without a tag, filtering by CONTAINER_NAME does not work.

Reading logs:

journalctl -u docker --grep="nginx"
journalctl --user-console -t docker --since "1 hour ago"

More precisely, through journald filters:

journalctl -t docker -g "nginx" --since "2024-01-01"

Actual filtering depends on which metadata Docker passes into journald. Check journalctl -o verbose for a specific container to see what fields are available.

Comparing json-file and journald

Parameterjson-filejournald
Storage location/var/lib/docker/containers/.../var/log/journal/
RotationVia --log-optVia journald.conf
Searchdocker logs --since, grepjournalctl --grep, --since
DependenciesNonesystemd
CentralizationVia fluentd, gelf, awslogsVia journalctl --remote or forward
CompressionNo (manual)Yes, configurable in journald.conf
Access without DockerDirect file accessOnly via journalctl
Warning

journald does not support all log-opt flags available for json-file. For example, max-size and max-file do not work — rotation is controlled by journald’s own settings (SystemMaxUse, SystemMaxFileSize, etc.).

Configuring the driver in daemon.json

The global setting goes in /etc/docker/daemon.json:

{
  "log-driver": "journald",
  "log-opts": {
    "tag": "{{.Name}}"
  }
}

After changing it, restart Docker:

sudo systemctl restart docker
Note

Changing the driver in daemon.json affects all new containers. Already running containers continue using their current driver until restarted.

Per-container override is possible via --log-driver and --log-opt at startup — this takes precedence over daemon settings.

To check the current driver for a specific container:

docker inspect --format='{{.HostConfig.LogConfig.Type}}' <container>

Practical recommendations

For local development, json-file with explicit max-size and max-file is sufficient and simple to use. For production with dozens of containers on a single host, journald is preferable: a unified search space, built-in compression, and integration with systemd and monitoring.

Tip

If you already use systemd to orchestrate containers (via systemd unit files or Podman), journald is the natural choice. Container logs and service logs end up in one place.

For centralized collection, both drivers support log forwarding through intermediate drivers (fluentd, gelf, splunk). But journald adds an extra step: first into journald, then the forwarder. For simple cases, direct json-file + fluentd may be a shorter path.

Monitor disk space regularly regardless of the driver. journalctl --disk-usage and du -sh /var/lib/docker/containers/*/ are the minimum set for this.

11 - scp — Secure Copy Over SSH

scp — a utility for copying files over SSH using the SSH protocol. It works from the terminal, requires no extra server setup — just a running sshd and working authentication. In an era of rsync and bat, SCP survives as a simple tool for one-off transfers when you don’t want to deal with daemons or configuration files.

Syntax and Basic Scenarios

General form:

scp [flags] source destination

Source and destination can be local paths or remote addresses in the format user@host:path.

# Local file to a remote machine
scp ./deploy.tar.gz deploy@10.0.2.15:/opt/app/

# Remote file to the local machine
scp deploy@10.0.2.15:/opt/app/deploy.tar.gz ./

# Between two remote hosts (via the local machine)
scp user@host1:/data/backup.sql user@host2:/data/restore/
Tip

If the remote host uses a non-standard SSH port, specify it with -P (uppercase P — that’s how scp differs from ssh).

Recursive Directory Copying

To copy a directory, the -r flag is required. Without it, scp refuses to transfer a directory and prints an error.

scp -r ./project/ dev@10.0.2.15:/home/dev/projects/
Warning

When copying recursively, scp transfers the contents of the directory, not the directory itself. Behavior depends on whether the trailing / is present in the path — verify the result if the structure matters.

Useful Flags

FlagDescription
-rRecursive directory copying
-P portSSH port on the remote host
-pPreserves modification time, access, and file permissions
-qQuiet mode, no progress bar
-CCompression during transfer
-i key.pemSpecifies a private key
-o StrictHostKeyChecking=noAutomatically accepts new host keys
-l rateBandwidth limit in Kbit/s
scp -r -p -C -i ~/.ssh/deploy_key.pem -P 2222 ./build/ deploy@10.0.2.15:/var/www/

Typical Transfer Patterns

Deploying artifacts:

scp ./release.tar.gz deploy@prod:/tmp/releases/

Fetching a log from a remote server:

scp admin@10.0.2.15:/var/log/app/error.log ./logs/

Transferring multiple files at once:

scp config.yaml secrets.env deploy@10.0.2.15:/opt/app/

Direct transfer between two servers (both source and destination are remote):

scp -3 user@host1:/data/file.csv user@host2:/data/import/

The -3 flag routes traffic through the local machine. Without it, scp attempts a direct connection between the hosts, which usually fails.

Limitations and Alternatives

scp does not support resuming interrupted transfers — if the connection drops halfway through a multi-gigabyte file, you start over. There is no incremental sync, no checksum verification. Transfers are sequential, with no built-in parallel file transfer.

For everyday tasks, rsync solves these problems:

rsync -avz -e "ssh -p 2222" ./build/ deploy@10.0.2.15:/var/www/

rsync can resume broken transfers, skip already-copied files, and work incrementally. For one-off transfers of a few small files, scp remains convenient — fewer parameters, faster to type.

Note

On modern distributions, scp may be wrapped by the OpenSSH client. Behavior is the same, but if you notice differences in output or error handling, that’s normal — the implementation depends on the OpenSSH version.

12 - Sudoers: NOPASSWD Without Holes

Unrestricted NOPASSWD in sudoers is a misconfiguration that grants root access without a password, turning any user script or library vulnerability into a full system compromise. The correct approach limits NOPASSWD to specific commands only.

Why NOPASSWD + ALL Is a Hole, Not a Solution

%admin ALL=(ALL) NOPASSWD: ALL — the most common sudoers error. The user receives unlimited root access without a password. Any script, any utility, any vulnerability in the user’s environment becomes a direct path to full machine control. NOPASSWD without command restrictions is not convenience; it is a backdoor in plain sight.

The proper method: allow specific commands via Cmnd_Alias and attach NOPASSWD only to them. Then the user can restart a service but cannot read /etc/shadow or run su.

visudo: The Only Safe Way to Edit sudoers

Editing /etc/sudoers directly via vi or nano is a path to locking yourself and the entire team out. visudo locks the file, validates syntax before saving, and rejects invalid entries.

# Correct path:
sudo visudo

# If the default editor is inconvenient:
sudo EDITOR=nano visudo

# For separate files in /etc/sudoers.d/:
sudo visudo -f /etc/sudoers.d/deployer
Warning

A syntax error in sudoers = loss of sudo capabilities for all. visudo prevents this, but only when used.

Cmnd_Alias: Grouping Commands Instead of Allowing Everything

Cmnd_Alias lets you create a named group of commands. You then reference the name — readable and easy to change.

Cmnd_Alias RESTART_WEB = /usr/bin/systemctl restart nginx, /usr/bin/systemctl reload nginx
Cmnd_Alias RESTART_DB = /usr/bin/systemctl restart postgresql
Cmnd_Alias PACKAGE_MGMT = /usr/bin/apt, /usr/bin/yum, /usr/bin/dnf
Cmnd_Alias LOG_VIEW = /usr/bin/tail, /usr/bin/journalctl

Alias syntax: name in uppercase, comma-separated absolute paths. The path is mandatory — systemctl without /usr/bin/ will not work.

Example: Limited NOPASSWD for Specific Tasks

Real scenario: a deployer user needs to restart nginx and view logs, nothing more.

Cmnd_Alias RESTART_WEB = /usr/bin/systemctl restart nginx, /usr/bin/systemctl reload nginx
Cmnd_Alias LOG_VIEW = /usr/bin/journalctl, /usr/bin/tail

deployer ALL=(ALL) NOPASSWD: RESTART_WEB, LOG_VIEW

Now deployer can:

sudo systemctl restart nginx    # without password
sudo journalctl -u nginx        # without password
sudo apt update                 # denied
sudo su                         # denied

To allow a single user one binary:

monitor ALL=(ALL) NOPASSWD: /usr/bin/tail /var/log/syslog
Tip

For multiple users performing the same task — use a group. %deployers ALL=(ALL) NOPASSWD: RESTART_WEB is more maintainable than duplicating lines.

What to Avoid: Common sudoers Mistakes

ОшибкаПочему плохоКак правильно
ALL ALL=(ALL) NOPASSWD: ALLFull root without passwordList specific commands via Cmnd_Alias
user ALL=NOPASSWD: /bin/bashOpens root shellNever give interpreters or su
user ALL=(ALL) ALL, NOPASSWD: ALLNOPASSWD spreads to everything due to orderNOPASSWD: before the command list, not after
Editing /etc/sudoers via echo or cpNo syntax validationOnly visudo or visudo -f
Absence of #includedir /etc/sudoers.dManual file overwritten on updateCheck include, place custom rules in /etc/sudoers.d/

Another frequent trap — spaces in Cmnd_Alias. A comma and space after it are mandatory:

# Correct:
Cmnd_Alias WEB = /usr/bin/systemctl restart nginx, /usr/bin/systemctl reload nginx

# Incorrect (space replaces comma, parsing breaks):
Cmnd_Alias WEB = /usr/bin/systemctl restart nginx /usr/bin/systemctl reload nginx

Verify a rule without root privileges:

sudo -l -U deployer

The output lists permitted commands and their flags. An empty list means the rules did not apply — check syntax and paths. Here’s a thinking process:

  1. Analyze User Input:
  • Role: Practicing Lead DevOps writing IT notes for a Hugo blog (Mana theme).
  • Style: Short paragraphs, specific commands, tables of flags, code with language fences.
  • Constraints: No water, no “in this article we’ll explore”, no marketing, no emojis, no author signature at the end.
  • Allowed: GitHub admonitions: > [!NOTE], > [!TIP], > [!WARNING].
  • Fact Constraint: Don’t make up non-existent flags, URLs, or versions. If fact is unclear, write cautiously as in operations.
  • Task: Write the English article as a parallel original, not a word-for-word translation.
  • Structure & Facts: Must match the Russian draft exactly.
  • Format:

13 - ThinLinc: Remote Access to Linux Desktops

ThinLinc: Remote Access to Linux Desktops

In enterprise environments, remote access to Linux desktops often comes down to VNC with flaky encryption or RDP proxies held together with workarounds. ThinLinc by Cendio is a full-featured remote desktop solution that runs on top of VNC, supports RDP clients, and provides a web-based administration interface without the usual hassle with certificates and firewalls.

What Is ThinLinc

ThinLinc is a remote desktop access solution with a server-plus-clients architecture. The server runs VNC sessions on the backend, and clients connect through a web browser or native ThinLinc clients. The protocol between client and server is tunneled over SSH, which solves the encryption problem out of the box.

Key features:

  • Built-in web server for browser-based desktop access
  • Native clients for Linux, Windows, and macOS
  • Load balancing across multiple servers
  • Single sign-on (SSO) via Kerberos and LDAP
  • Session management through web panel and CLI

Installing the Server

ThinLinc ships as packages for RHEL/CentOS 7/8/9 and Ubuntu/Debian. For installation on a RHEL system:

# Add the Cendio repository
sudo yum install -y https://www.cendio.com/downloads/thinlinc/rpm/cendio-release-latest.noarch.rpm

# Install the server
sudo yum install -y thinlinc-server

# Start services
sudo systemctl start tlwebd
sudo systemctl start vncserver@:1

For Ubuntu/Debian:

sudo dpkg -i thinlinc-server-*.deb
sudo systemctl start tlwebd
sudo systemctl start vncserver@:1

After installation, the web panel is available at https://<host>:1443.

Note

Port 1443 is the standard ThinLinc web interface port. If a firewall is in place, open it along with port 22 for SSH tunneling.

Server Configuration

The main configuration lives in /etc/thinlinc/. Key files:

FilePurpose
tlconfigGlobal server settings
vncserver-config-defaultsVNC session parameters
client-to-server.d/Port and device forwarding rules
ssl/Certificates and keys

Basic configuration via tlconfig:

# Set default screen size
sudo tlconfig --set vncserver.default_screen_width 1920
sudo tlconfig --set vncserver.default_screen_height 1080

# Set color depth
sudo tlconfig --set vncserver.default_color_depth 24

# Set home directory for sessions
sudo tlconfig --set vncserver.home_dir /var/lib/thinlinc

For VNC parameter tuning:

# Edit the VNC configuration
sudo tlconfig --edit vncserver-config-defaults
Warning

After changing configuration via tlconfig, restart the services: sudo systemctl restart tlwebd vncserver@:*.

Client Connections

ThinLinc provides several connection methods:

  1. Web browser — navigate to https://<host>:1443, enter credentials. Works on any device with a modern browser.
  2. Native client — download from the same URL. Clients are available for Linux, Windows, and macOS.
  3. RDP client — ThinLinc supports RDP proxying, allowing connections via standard rdesktop or freerdp.

Connecting via native client:

# Linux
thinlinc-client

# Windows (PowerShell)
Start-Process "C:\Program Files\ThinLinc\Client\tlclient.exe"

Manual SSH tunnel connection:

ssh -L 5901:localhost:5901 user@thinlinc-host
vncviewer localhost:5901

Administration and Session Management

All active sessions can be viewed through the web panel (https://<host>:1443/admin) or via CLI:

# List all sessions
sudo /opt/thinlinc/bin/vncserver -list

# Terminate a specific session
sudo /opt/thinlinc/bin/vncserver -kill :<display_number>

# Restart a specific session
sudo /opt/thinlinc/bin/vncserver -restart :<display_number>

For bulk operations:

# Terminate all sessions for a user
sudo /opt/thinlinc/bin/vncserver -kill -u username

# Limit sessions per user
sudo tlconfig --set vncserver.max_sessions_per_user 3

The admin web panel allows:

  • Viewing session lists and their status
  • Sending messages to users
  • Forcefully terminating sessions
  • Viewing logs and usage metrics

Security and Integration

ThinLinc uses SSH for encrypting traffic between client and server. Web server certificates can be replaced with your own:

# Replace the self-signed certificate
sudo cp server.crt /etc/thinlinc/ssl/
sudo cp server.key /etc/thinlinc/ssl/
sudo systemctl restart tlwebd

LDAP/Kerberos integration:

# Enable LDAP authentication
sudo tlconfig --set authentication.ldap.enabled true
sudo tlconfig --set authentication.ldap.server ldap://ldap.example.com
sudo tlconfig --set authentication.ldap.base_dn "dc=example,dc=com"

# Enable Kerberos
sudo tlconfig --set authentication.krb5.enabled true
sudo tlconfig --set authentication.krb5.realm EXAMPLE.COM

For PAM integration:

# Use system PAM authentication
sudo tlconfig --set authentication.pam.enabled true
Tip

ThinLinc supports USB device and audio forwarding over the SSH tunnel. Configure it in client-to-server.d/ rules for specific devices.

ThinLinc is a ready-made solution that eliminates the manual assembly of VNC infrastructure with firewalls and certificates. For environments that need remote access to Linux desktops without security compromises, it is one of the most straightforward paths available.

14 - Chrony Instead of ntpd

Chrony has replaced ntpd as the default NTP client in most modern Linux distributions. It converges to accurate time faster, handles intermittent network connections better, and consumes fewer resources. If your machine still runs ntpd, switching takes only a few minutes.

Installation

On RHEL-based systems:

sudo dnf install chrony -y
sudo systemctl enable --now chronyd

On Debian/Ubuntu:

sudo apt install chrony -y
sudo systemctl enable --now chronyd

If ntpd was running on this machine before, stop and disable it to avoid port conflicts:

sudo systemctl stop ntpd
sudo systemctl disable ntpd

makestep Configuration

The key directive in /etc/chrony/chrony.conf (or /etc/chrony.conf on RHEL) is makestep. It controls how chronyd behaves at startup — whether to correct time gradually or in a single step.

makestep 1.0 3
makestep

Format: makestep <max_offset> <max_updates>. If the offset exceeds <max_offset> seconds and the number of updates hasn’t exceeded <max_updates>, chrony applies a sudden correction instead of gradual slewing. The default makestep 1.0 3 allows up to a 1-second jump three times during the first synchronizations.

A common mistake is setting makestep -1 1 and expecting a server with a large initial offset to correct instantly. In practice, negative values only work under specific conditions. For reliable startup, use a positive number and limit the number of steps.

allow Directive

By default, chronyd operates only as a client. To let other hosts synchronize through this server, add allow:

allow 10.0.0.0/24
allow

You can specify individual IPs, subnets, or multiple allow lines for different networks. Without this directive, the machine accepts requests only from localhost.

To block a specific host, use deny — it takes effect after allow and overrides it:

allow 10.0.0.0/24
deny 10.0.0.42

After changing the configuration, restart the service:

sudo systemctl restart chronyd

Verification with timedatectl

timedatectl shows the current synchronization state and time source:

timedatectl

Example output:

               Local time: Wed 2025-01-15 14:23:01 MSK
           Universal time: Wed 2025-01-15 11:23:01 UTC
                 RTC time: Wed 2025-01-15 11:23:01
                Time zone: Europe/Moscow (MSK, +0300)
System clock synchronized: yes
              NTP service: active
          RTC in local TZ: no

Key fields for diagnostics:

FieldProblem ValueMeaning
System clock synchronizednochrony hasn’t caught up yet
NTP serviceinactiveservice not running or not enabled
RTC in local TZyeshardware clock in local timezone — common issue on VMs

For detailed information about current sources:

chronyc sources -v

If NTP service: active and System clock synchronized: yes, everything is working. If not, check sudo systemctl status chronyd and network access to NTP servers (UDP port 123).

15 - fail2ban: SSH Jail Configuration

Securing SSH against brute-force attacks is one of the first steps in hardening any server. fail2ban scans logs, detects repeated failed login attempts, and blocks the source via iptables or nftables. This note covers the sshd jail — from installation to fine-tuning ban durations.

Installation and Basic Configuration

Install from the standard repository:

# Debian/Ubuntu
apt install fail2ban

# RHEL/CentOS
yum install fail2ban

Enable and start the service:

systemctl enable --now fail2ban
systemctl status fail2ban
Warning

Do not edit /etc/fail2ban/jail.conf directly — package updates will overwrite your changes. All local overrides go in jail.local.

Create a local overrides file:

cp /etc/fail2ban/jail.conf /etc/fail2ban/jail.local

Or, preferably, create a minimal jail.local with only the parameters you need — fail2ban stacks jail.local on top of jail.conf, so overriding specific keys is sufficient.

Filter for sshd

A ready-made filter ships with the package: /etc/fail2ban/filter.d/sshd.conf. It scans /var/log/auth.log (or /var/log/secure on RHEL) for patterns like Failed password for invalid user and Connection closed by authenticating user.

Verify the filter parses your logs correctly:

fail2ban-regex /var/log/auth.log /etc/fail2ban/filter.d/sshd.conf

If the output shows a high match rate, the filter works. If not, check the logpath parameter in the jail and the logpath value inside the [Definition] section of the filter.

Tip

For additional protection against brute-force via ddos-style attacks, there is a separate filter sshd-ddos. It catches rapid repeated connections from a single IP. Enable it by adding mode = ddos to the jail parameters.

bantime and findtime Parameters

Three keys define the blocking logic:

ParameterDescriptionDefault
bantimeBan duration in seconds (negative = permanent)600
findtimeObservation window for counting failed attempts600
maxretryFailed attempts before ban5

A typical production configuration:

[sshd]
bantime  = 3600
findtime = 600
maxretry = 3

This means: three failed logins within ten minutes triggers a one-hour ban. For critical servers, set bantime = -1 (permanent ban) and unblock manually with fail2ban-client set sshd unbanip <IP>.

Note

bantime and findtime accept suffixes: d (days), h (hours), m (minutes), s (seconds). For example, bantime = 1d.

Activating the Jail

By default, the [sshd] section is commented out in jail.local. Enable it:

[sshd]
enabled = true
port    = ssh
filter  = sshd
logpath = /var/log/auth.log
maxretry = 3
bantime  = 3600
findtime = 600

On RHEL/CentOS, logpath is typically /var/log/secure. Verify the log path on your distribution.

Restart fail2ban and check the status:

systemctl restart fail2ban
fail2ban-client status
fail2ban-client status sshd

The output of fail2ban-client status sshd shows the current number of banned IPs and active filters.

Manual block and unblock:

fail2ban-client set sshd banip 203.0.113.50
fail2ban-client set sshd unbanip 203.0.113.50
Warning

fail2ban is not a WAF and not a replacement for key-based authentication. Use keys instead of passwords, restrict access with AllowUsers, and change the port where feasible. fail2ban complements these measures — it does not replace them.

Check fail2ban logs (/var/log/fail2ban.log) on first startup — they show whether the jail picked up the log file and whether filters are firing.

16 - journalctl: filters and follow

Systemd’s journal is the first place to look when a service crashes or a node starts burning CPU. journalctl does far more than dump the entire log in sequence: it can filter by units, priorities, time windows, and stream in real time. Below is the working set I use daily.

Follow in real time

Behavior similar to tail -f, but aware of journald’s structured format:

journalctl -f

The -f flag (short for --follow) streams new entries as they appear. By default it shows all units — handy when you don’t know where the fire is.

Tip

Add --no-pager so output isn’t intercepted by less and doesn’t block the terminal in scripts and CI.

journalctl -f --no-pager

Filter by unit

The unit is the most common filter. One -u key and the specific service name:

journalctl -u nginx.service
journalctl -u docker.service

Multiple units can be passed — journalctl will show entries from all specified ones:

journalctl -u nginx.service -u postgresql.service
Note

The unit name must include the .service suffix. Omitting it may result in no matches or unexpected output from journalctl.

Priority and time filters

Priority filtering is set via -p (or --priority). Levels range from 0 to 7:

PriorityValue
0emerg
1alert
2crit
3err
4warning
5notice
6info
7debug

Show only errors and critical:

journalctl -p err

Time windows use --since and --until. Both absolute dates and relative expressions are supported:

journalctl --since "1 hour ago"
journalctl --since today --until "2 hours ago"
journalctl --since "2025-01-15 08:00:00" --until "2025-01-15 12:00:00"
Warning

--since without --until shows from the specified moment to now. If both are specified — the window is closed. Check the date order, or you’ll get empty output.

Combining flags

In practice, filters are combined. A typical request: watch errors from a specific service over the last hour in real time:

journalctl -u nginx.service -p err --since "1 hour ago" -f --no-pager

Another common pattern — show the last N lines for a unit with a priority filter:

journalctl -u postgresql.service -p warning -n 100 --no-pager

Here -n 100 limits output to the last 100 entries.

Quick reference for everyday flags:

FlagPurpose
-u <unit>Filter by unit
-fFollow (new entries in real time)
-p <level>Minimum priority
--since <time>Start of time window
--until <time>End of time window
-n <count>Last N entries
--no-pagerWithout pager
Tip

If journalctl returns empty results without errors — check whether the service is inactive or failed. You can see the status via systemctl status <unit>. Also verify that journald hasn’t rotated the needed entries: journalctl --disk-usage shows how much space the journal occupies.

17 - nftables: Basic Rule Set

Note

All commands were verified on Debian/Ubuntu with the nftables package and on RHEL/CentOS 8+. On older systems you may need apt install nftables or yum install nftables.

nftables replaced iptables, but documentation for a basic rule set is often scattered. Here is the reference I use when bringing up a firewall on a new host.

Creating the inet filter table

The inet family table unifies IPv4 and IPv6 under a single namespace. This is the preferred approach when both stacks are active on the host.

nft add table inet filter

If the table already exists the command returns an error. To avoid duplication when a script is re-run:

nft 'add table inet filter' 2>/dev/null || true

To wipe the current rule set before loading your own:

nft flush ruleset
Warning

flush ruleset removes all rules instantly. On a production machine run this only from the console, not over a remote session without a fallback.

Input and forward chains

Chains bind to a table and define the interception point for traffic. A basic firewall needs input (traffic destined for the host itself) and forward (traffic passing through the host).

nft add chain inet filter input { type filter hook input priority 0 \; policy drop \; }
nft add chain inet filter forward { type filter hook forward priority 0 \; policy drop \; }

Key components inside the curly braces:

ComponentValue
type filterChain type, standard for packet filtering
hook input / hook forwardInterception point in the network stack
priority 0Processing priority
policy dropDefault policy — drop unmatched packets

The syntax requires escaping semicolons inside the string or using single quotes as shown above.

Tip

If you need to allow established connections, add an output chain with policy accept or use connection tracking in your input rules.

Basic rules for input

With a chain set to policy drop, you must explicitly permit the traffic you need. A typical minimum:

# Allow loopback
nft add rule inet filter input iif lo accept

# Allow established and related connections
nft add rule inet filter input ct state established,related accept

# Allow SSH (port 22)
nft add rule inet filter input tcp dport 22 accept

# Allow ping (ICMP echo request)
nft add rule inet filter input ip protocol icmp icmp type echo-request accept
nft add rule inet filter input ip6 nexthdr icmpv6 icmpv6 type echo-request accept

Each rule appends to the end of the chain. Order matters: accept rules for loopback and established traffic should come before rules with narrower criteria.

To log dropped packets before the drop policy (optional but useful for diagnostics):

nft add rule inet filter input log prefix "nft-drop: " level warn
Note

Logging adds overhead. On high-throughput interfaces use a rate limit: limit rate 10/second.

Basic rules for forward

The forward chain is needed when the host acts as a router or NAT gateway. A minimum setup:

# Allow established connections
nft add rule inet filter forward ct state established,related accept

# Allow forwarding between specific interfaces (example)
nft add rule inet filter forward iifname "eth0" oifname "eth1" accept

If the host does not perform routing, leave forward with policy drop and no additional rules.

To enable IP forwarding at the kernel level (if not already done):

sysctl -w net.ipv4.ip_forward=1
sysctl -w net.ipv6.conf.all.forwarding=1

Viewing and managing rules

After setup, verify the current configuration:

nft list table inet filter
nft list chain inet filter input
nft list ruleset

list ruleset outputs the full config, which you can save and reuse as a boot script.

Deleting a specific rule by handle within a chain:

nft delete rule inet filter input handle <handle-number>

The handle number appears in the output of nft list ruleset -a.

To delete an entire chain:

nft delete chain inet filter input

A chain can only be deleted when it is empty. To remove a table completely, delete all chains inside it first.

Saving rules to a file for boot-time loading:

nft list ruleset > /etc/nftables.conf

On systems with systemd, enable auto-start:

systemctl enable nftables
systemctl start nftables
Warning

If /etc/nftables.conf does not exist or is empty, the service will not load any rules. Create the file manually before enabling the service.

18 - Debian — The Swiss Army Knife of Linux

Debian isn’t the flashiest distribution, but it’s the most reliable foundation in the Linux world. Behind it stands the largest community of volunteer developers, and behind its track record are decades of uninterrupted operation on servers, embedded systems, and cloud infrastructure. If you manage Linux machines in production, Debian (or one of its derivatives) is already in your stack — whether you’ve acknowledged it or not.

Managing Packages: apt and dpkg

Two tools form the core of Debian’s packaging system. apt is the high-level interface for working with repositories, resolving dependencies, and upgrading the system. dpkg is the low-level engine that installs, removes, and inspects individual .deb files without contacting repositories.

# Update index and upgrade system
apt update
apt upgrade -y
apt full-upgrade -y

# Install and remove
apt install -y nginx
apt remove --purge nginx
apt autoremove -y

# Search and inspect
apt search nginx
apt show nginx
apt list --installed | grep nginx

# dpkg for direct .deb manipulation
dpkg -i package.deb
dpkg -l | grep nginx
dpkg -L nginx
dpkg -S /usr/bin/nginx
Warning

dpkg -i does not resolve dependencies. If a package pulls in other libraries, use apt install ./package.deb instead — apt will fetch missing dependencies from the repository.

A common mistake is mixing apt and dpkg during partial installations. If dpkg -i fails with a dependency error, run apt --fix-broken install and complete the operation through apt.

Stability Model and Release Branches

Debian is built on three branches that determine the aggressiveness of updates and the level of testing.

BranchCodenameStabilityWhen to Use
stablebookworm (12) / trixie (13)High — security and critical fixes onlyProduction servers, base infrastructure
testingbookworm-progressMedium — packages go through a stabilization periodDevelopment machines, containers
unstablesidLow — daily snapshotsExperiments, binary package building

In sources.list this looks like:

# Stable — recommended for production
deb http://deb.debian.org/debian bookworm main contrib non-free non-free-firmware
deb http://deb.debian.org/debian-security bookworm-security main contrib non-free non-free-firmware
deb http://deb.debian.org/debian bookworm-updates main contrib non-free non-free-firmware
Note

The non-free and non-free-firmware sections were introduced starting with Debian 12. Without them, certain proprietary drivers (Wi-Fi, GPU) won’t install.

Moving between branches is apt dist-upgrade with a swapped sources.list. Do this only on test benches — binary compatibility between testing and stable is not guaranteed for all packages.

Architecture Support

Debian officially supports more hardware platforms than any other distribution. This is critical for embedded and IoT projects.

ArchitecturePortTypical Use
amd64amd64Servers, desktops, cloud
arm64arm64ARM servers (AWS Graviton, Raspberry Pi 4/5)
armhfarmhfEmbedded devices with FPU
i386i386Legacy systems
ppc64elppc64elIBM Power Systems
s390xs390xIBM Z mainframes
riscv64riscv64Open RISC-V platforms

Multi-architecture support allows installing packages for foreign platforms:

dpkg --add-architecture arm64
apt update
apt install -y package:arm64

Debian as a Base for Other Distributions

Nearly every major distribution you use is built on Debian or inherits its packaging model.

  • Ubuntu — takes the testing branch, adds its own PPAs and proprietary drivers.
  • Linux Mint — based on Ubuntu, and therefore on Debian packages.
  • Proxmox VE — a modified Debian 12 with added enterprise repositories.
  • Kali Linux — Debian unstable with a suite of pentesting tools.
  • Raspberry Pi OS (formerly Raspbian) — Debian armhf/arm64 with Pi-specific patches.
  • MX Linux, antiX — Debian stable with lightweight desktop environments.
Tip

If you need a stable backend with fresher software, use Debian stable with the backports repository. This gives you updated kernel, nginx, PostgreSQL versions without switching to testing.

echo "deb http://deb.debian.org/debian bookworm-backports main" >> /etc/apt/sources.list
apt update
apt install -t bookworm-backports linux-image-amd64

Practical DevOps Scenarios

In day-to-day work, the Debian approach manifests in several key patterns.

Base container images. Official Docker images like debian:bookworm, debian:slim are the starting point for hundreds of minimalist containers. debian:slim weighs around 70 MB compared to 200+ MB for ubuntu:latest.

FROM debian:slim
RUN apt-get update && apt-get install -y --no-install-recommends curl && rm -rf /var/lib/apt/lists/*

VM templates. Proxmox, which runs on Debian, uses debootstrap to create LXC container templates. It is the same tool that underlies a fresh system installation.

Configuration management. Ansible, Puppet, Chef — all three work directly with Debian packages. The apt module in Ansible is one of the most-used modules in any playbook.

CI/CD pipelines. Building .deb packages in GitLab CI or GitHub Actions is a standard pattern. debhelper, dh_make, pbuilder, or sbuild provide reproducible builds in a clean chroot.

Key Commands for Daily Work

Bookmark this table — it covers 90% of tasks on a Debian server.

TaskCommand
Update package listapt update
Upgrade all packagesapt upgrade -y
Full upgrade resolving dependenciesapt full-upgrade -y
Install a packageapt install -y <pkg>
Remove a package with configsapt purge -y <pkg>
Search for a packageapt search <query>
Show package infoapt show <pkg>
List installed packagesapt list --installed
Find which package owns a filedpkg -S <path>
List files in a packagedpkg -L <pkg>
Check package statusdpkg -l <pkg>
Fix broken dependenciesapt --fix-broken install
Clean apt cacheapt clean && apt autoclean
Add an architecturedpkg --add-architecture <arch>
Build .deb from sourcedebuild -us -uc

Debian isn’t the fastest at first glance — no rolling releases and no fresh software versions out of the box. But that predictability is exactly what makes it the foundation for production environments running millions of servers. If you want reliability over trending versions, Debian won’t let you down.

19 - pipx: Isolated Python CLI Tools Without the Mess

pipx solves a simple but chronic problem: you need to run a Python utility once or occasionally, and pip install pollutes the global environment or leaves behind a virtual environment you forget to clean up. pipx creates an isolated venv for each utility, installs dependencies there, and makes the binary available in $PATH. One command and the tool works without conflicting with anything.

What Is pipx and Why You Need It

pipx installs and runs Python applications in isolated virtual environments. Each utility lives in its own venv under ~/.local/pipx/venvs/, and its console-scripts are symlinked into ~/.local/bin/.

Problems it solves:

  • Version conflicts between projects: black==23 and black==24 can’t coexist in one environment, but in pipx they can.
  • Global pip install pollutes system Python and can break apt on Debian-based systems.
  • Forgotten venvs after one-off usage.
Note

pipx doesn’t replace pip inside projects. It’s a tool for CLI utilities: black, poetry, httpie, ansible, awscli, pre-commit, and similar.

Installing pipx

The most reliable path is through pip in user mode or through your system package manager.

# Option 1: via pip (Python 3.6+)
python3 -m pip install --user pipx
python3 -m pipx ensurepath

# Option 2: via apt (Debian/Ubuntu; version may be older)
sudo apt install pipx

# Option 3: via brew (macOS)
brew install pipx

After installation, verify ~/.local/bin is in $PATH:

echo $PATH | grep -q "$HOME/.local/bin" && echo "OK" || echo "add to PATH"
Warning

If ensurepath didn’t work, add it manually to ~/.bashrc or ~/.zshrc: export PATH="$HOME/.local/bin:$PATH"

Basic Commands: install, run, list

Three commands cover 90% of use cases.

# Install a utility globally (creates venv, symlinks binary)
pipx install black

# Run a utility without installing it (download + run in a temporary venv)
pipx run httpie https://api.example.com/health

# List all installed utilities
pipx list

Flags worth remembering:

FlagWhat it doesExample
--specSpecify source (PyPI, git, wheel)pipx install --spec git+https://github.com/user/repo.git tool
--suffixAdd suffix to binary namepipx install black --suffix==24
--pythonSpecify interpreterpipx install --python python3.11 black
--system-site-packagesEnable access to system packagespipx install --system-site-packages tool
--forceReinstall over existingpipx install --force black
Tip

pipx run is the key command for one-off usage. It downloads the package, creates a temporary venv, executes, and removes it. No trace left behind.

Managing Dependencies and Reinstallation

After installation you can upgrade, remove, and inspect dependencies.

# Upgrade a single utility
pipx upgrade black

# Upgrade all installed utilities
pipx upgrade-all

# Remove a utility and its venv entirely
pipx uninstall black

# Remove everything except pipx itself
pipx uninstall-all

# Inspect dependencies of an installed package
pipx list --verbose

If something breaks, reinstallation takes seconds:

pipx reinstall black
# or with a specific interpreter
pipx reinstall --python python3.12 black
Warning

pipx upgrade-all can bump a tool to a version with breaking changes. In CI/CD, pin the version instead: pipx install black==24.8.1.

Typical DevOps Scenarios

pipx fits several working patterns where a full project with requirements.txt is overkill.

1. One-off utilities in CI/CD. Instead of installing into a Docker image or globally:

# In a Dockerfile or entrypoint script
pipx run --spec https://pypi.org/project/aws-nuke/ aws-nuke --force --account-id $AWS_ACCOUNT_ID

2. Parallel versions of the same tool. Useful during project migrations:

pipx install black --suffix==23
pipx install black --suffix==24
black==23 --version
black==24 --version

3. Local dev environment without privileges. Install ansible, pre-commit, and similar tools without sudo and without affecting system Python.

4. Quick package evaluation before integration. Test a utility without adding it to requirements.txt:

pipx run httpx https://example.com
# If it works — install permanently
pipx install httpx
Tip

Combined with direnv and .envrc, you can wire up pipx run for project-specific tasks — the utility is available only in that directory, and dependencies don’t leak into the global environment.

pipx doesn’t try to be a package manager for all of Python. It does one thing — isolated CLI utility installation — and does it without the noise. For a Lead DevOps, that means less time fighting dependency conflicts and more time on architecture.

20 - uv — Fast Python Package Manager

What is uv

uv is a Python package manager written in Rust. It solves one problem: the standard pip is slow at resolving dependencies, and poetry adds its own project model on top of PEP 621. uv works with pyproject.toml, is compatible with PEP 621 and PEP 508, and does it significantly faster.

Under the hood is a caching resolver written in Rust that reuses data from pip-compatible indexes (PyPI by default). uv can replace pip, pip-tools, virtualenv, and poetry in a single tool.

Note

uv does not require an installed Python to install itself, but to work with Python projects you need an interpreter. uv can find and manage CPython via uv python.

Installation

The fastest path is the curl script:

curl -LsSf https://astral.sh/uv/install.sh | sh

This places the binary in ~/.local/bin. On systems without curl:

pip install uv

Or download the standalone binary from GitHub releases.

After installation, verify the version:

uv --version
Warning

If uv is not found after installation, make sure ~/.local/bin is in your $PATH. On Debian/Ubuntu, the pip-installed package places it in /usr/local/bin.

Key commands

The project workflow:

uv init myproject          # creates pyproject.toml with PEP 621 metadata
cd myproject
uv add requests            # adds a dependency and updates uv.lock
uv add --group dev pytest  # adds a dependency to a named group
uv sync                    # synchronizes the environment with the lock file
uv run pytest              # runs a command inside the project's .venv

Key flags for daily work:

FlagWhat it does
--systemInstalls a package into the system Python, bypassing venv
--no-devExcludes dev dependencies during sync
--group <name>Specifies a named dependency group
--python <version>Selects a specific interpreter version
--lockedUses the lock file without re-resolving
--quietMinimal output
--verboseDetailed output for debugging

Managing interpreters:

uv python list             # shows discovered CPython installations
uv python install 3.12     # downloads and caches CPython 3.12
uv python find             # shows the path to the active Python

For package-level work without creating a project:

uv pip install requests    # pip equivalent, but faster
uv pip freeze              # lists installed packages
uv pip uninstall requests  # removes a package
Tip

uv sync recreates .venv and installs exactly what is in uv.lock. This makes it suitable for CI — the result is deterministic.

Comparison with pip and poetry

Aspectpippoetryuv
Implementation languagePythonPython + RustRust
Project modelsetup.py / pyproject.tomlpyproject.toml (custom format)pyproject.toml (PEP 621)
Resolution speedSlowMediumFast
Lock fileNo nativepoetry.lockuv.lock
Virtual environmentvenv / virtualenvBuilt-inBuilt-in (.venv)
pip compatibilityFullPartialFull
Replaces pip——Yes
Replaces poetry——Yes (partially)

uv wins on resolution speed — in astral-sh benchmarks it is 10–50x faster than pip on large dependency graphs. At the same time, uv pip is fully compatible with pip workflows: requirements.txt, --index-url, --find-links.

Warning

uv is under active development. The uv.lock format may change between major versions. In production CI pipelines, pin the uv version via uv --version or install a specific binary.

A common mistake is confusing uv add with uv pip install. The former writes the dependency into pyproject.toml and updates uv.lock; the latter works like pip directly. For projects with pyproject.toml, use uv add. For one-off scripts and Dockerfiles, use uv pip install.

Another frequent issue: uv cannot find the interpreter if Python is installed in a non-standard path. The fix is uv python install --force or explicit specification via --python /path/to/python.

21 - logrotate for Custom Daemons

Log rotation is missing for your daemon, the log file has grown to dozens of gigabytes, the disk is full, and monitoring is screaming. systemd-journald and syslog-ng rotate on their own, but if your custom daemon writes directly to a file, rotation falls to logrotate. Here is how to configure it for a specific service.

Why write a custom logrotate config

Packages from the repository usually drop their config into /etc/logrotate.d/, but for self-built daemons or those compiled from source, there is none. Without a config the file grows without limits. logrotate runs via a systemd timer (logrotate.timer) or cron and reads all files from /etc/logrotate.d/. Creating a single file is enough to start the rotation cycle.

Note

Verify that the logrotate package is installed and the timer is active: systemctl status logrotate.timer. On most distributions it is enabled by default.

copytruncate vs create — when to use which

These are two fundamentally different approaches to renaming and creating a new file.

copytruncatecreate
MechanismCopies the current file, truncates the original in placeRenames the old file, creates a new one with correct permissions
Application impactNo restart neededDaemon must be able to open the new file (typically via SIGHUP)
Risk of lost linesYes — entries written between copy and truncate can fall into the gapMinimal — renaming is atomic
When to useDaemon cannot re-open its file (e.g., written in Go without signal handling)Daemon supports SIGHUP or systemd-notify
Warning

copytruncate is a compromise. Lines written between cp and truncate are lost. For high-traffic services this can mean hundreds of lost lines per second.

If the daemon can receive signals, use create and restart it via postrotate.

delaycompress and how it affects the archive chain

By default logrotate compresses the rotated file in the same cycle. The problem: if the daemon is still writing to the old file (or has not yet re-opened the new one), compression breaks everything.

delaycompress defers compression by one cycle. The chain looks like this:

app.log          ← current
app.log.1        ← rotated, not yet compressed
app.log.2.gz     ← compressed, two cycles ago
app.log.3.gz     ← compressed, three cycles ago

Without delaycompress the transition is harsher: app.log immediately becomes app.log.1.gz, and if the daemon is still writing to app.log through its descriptor, data goes into the compressed archive — or is lost entirely.

Tip

delaycompress only makes sense together with create and compress. With copytruncate it works but loses its meaning — the file is truncated in place, so compression can happen immediately.

Integration with systemd notify

If the daemon supports sd_notify(3), you do not need to rely on postrotate with a manual restart. systemd can restart the service based on a signal from logrotate.

In the logrotate config, specify:

postrotate
    systemctl kill -s HUP my-daemon.service
endscript

Or, if the daemon listens on NOTIFY_SOCKET:

postrotate
    systemctl notify-reload my-daemon.service
endscript
Note

systemctl notify-reload is available starting from systemd 231. It sends RELOADING=1 followed by READY=1 — the standard mechanism for notifying a configuration reload.

For copytruncate, postrotate is usually unnecessary — truncating the file in place does not require daemon involvement.

Example working config

Suppose daemon my-app writes to /var/log/my-app/app.log and supports SIGHUP and sd_notify.

/var/log/my-app/app.log {
    daily
    rotate 14
    compress
    delaycompress
    missingok
    notifempty
    create 0640 myapp myapp
    postrotate
        systemctl notify-reload my-app.service >/dev/null 2>&1 || true
    endscript
}

Flags:

FlagWhat it does
dailyRotate every day
rotate 14Keep 14 archives
compressgzip compression
delaycompressDelay compression by one cycle
missingokDo not error if the file is absent
notifemptyDo not rotate an empty file
create 0640 myapp myappCreate new file with correct permissions and owner
postrotateNotify systemd of reload

To test without actually rotating:

logrotate -d /etc/logrotate.d/my-app

The -d flag runs in debug mode — it shows what would be done, without modifying any files.

Warning

After creating the config, the first rotation happens on the next timer tick. To force an immediate check: logrotate -f /etc/logrotate.d/my-app. This rotates the file right away, so use it carefully on production.

22 - Rsync: Backing Up a Directory Over SSH

Rsync: Backing Up a Directory Over SSH

The classic way to copy a directory to a remote machine is rsync over SSH. No extra ports to open, traffic is encrypted, and the tool itself handles incremental transfers and metadata preservation. One command and the backup is ready.


Basic rsync Command Over SSH

The minimal invocation to copy a local directory to a remote host:

rsync -avz /path/to/source/ user@host:/path/to/dest/

The trailing slash on source matters: without it, rsync creates a source/ subdirectory on the remote side; with it, the contents go directly into dest/.

If SSH listens on a non-standard port:

rsync -avz -e 'ssh -p 2222' /path/to/source/ user@host:/path/to/dest/
Tip

For cron automation, specify the port via -e rather than editing /etc/ssh/ssh_config — it’s easier to maintain different hosts with different ports.


Archive and Sync Flags

-a (archive) is the key flag. It bundles several options into one: recursive traversal, preservation of permissions, ownership, timestamps, symlinks, and empty directories.

FlagPurpose
-aArchive mode (recursion + metadata)
-vVerbose output
-zCompression during transfer
-P--partial --progress — resume and progress bar
--deleteRemove files on receiver absent from source
-e sshSpecify remote shell
--bwlimit=KBPSBandwidth limit in KB/s
Warning

--delete is a powerful tool. If the source is accidentally cleared, the remote machine will be left empty. Verify the list before applying.

Full backup example with bandwidth limit and progress:

rsync -avzP --bwlimit=10000 -e 'ssh -p 2222' \
  /data/backup/ user@backupserver:/mnt/backup/

File and Directory Exclusions

Use --exclude to omit specific paths. Patterns are relative to the source:

rsync -avz --exclude='*.log' --exclude='cache/' \
  /data/ user@host:/data/

When there are many exclusions, a file list is more convenient:

# exclude.txt
*.tmp
*.bak
cache/
lost+found/
rsync -avz --exclude-from='exclude.txt' /data/ user@host:/data/
Note

--exclude is evaluated in order. Later rules can override earlier ones if paths overlap. For precise control, use --filter.


Dry-Run Before Launch

Always run with --dry-run (or -n) first. rsync shows what would be copied or deleted without touching any files:

rsync -avzn --delete --exclude-from='exclude.txt' \
  /data/ user@host:/data/

Compare the output against the expected list. If it matches, remove -n and run the real sync.

For a cron job with email notifications:

#!/bin/bash
rsync -avz --delete --exclude-from='/etc/rsync-exclude.txt' \
  /data/ user@host:/data/ 2>&1 | mail -s "Rsync backup report" admin@example.com

Summary

rsync over SSH covers 90% of backup tasks without additional agents on the receiver side. The essential order of operations: first --dry-run, then --exclude, then --delete if exact mirroring is needed. Everything else is host- and port-specific tuning.

23 - Setting Up Your Own SSH Bastion Server

Why You Need a Bastion and Where It Lives

A bastion is the single entry point into a private network segment. Instead of exposing SSH on every server to the internet, you funnel traffic through one hardened host with a strict access policy. Typical layout: internet → bastion (public IP) → internal servers (only private subnet, SSH listening on 127.0.0.1 or a private interface).

The bastion sits in a demilitarized zone (DMZ) or a public subnet provided by your hosting platform. Internal machines have no route to the internet through the bastion — return traffic flows only over established connections. This is a baseline model you can deploy on any VPS in about 15 minutes.

Choosing an OS and Basic Setup

The bastion doesn’t need to be heavy. Debian, Ubuntu Server, or AlmaLinux — all work. I usually grab a minimal Ubuntu 22.04 LTS install and bring it to working state by hand.

# Update and install the minimum utility set
sudo apt update && sudo apt upgrade -y
sudo apt install -y fail2ban ufw curl htop
Tip

Don’t put a GUI, databases, or other services on the bastion. The smaller the attack surface, the better.

Create a non-privileged user for daily work:

sudo adduser deploy
sudo usermod -aG sudo deploy

SSH Configuration: Keys, Port, Disable Passwords

Generate a key on your workstation if you don’t have one yet:

ssh-keygen -t ed25519 -C "bastion-access"
ssh-copy-id -i ~/.ssh/id_ed25519.pub deploy@bastion_ip

On the bastion, edit /etc/ssh/sshd_config:

Port 22220
PermitRootLogin no
PasswordAuthentication no
PubkeyAuthentication yes
MaxAuthTries 3
MaxSessions 5
AllowUsers deploy
Warning

Before restarting SSH, make sure the key is added and works. Otherwise, you’ll lock yourself out of the server.

Validate the config and restart:

sudo sshd -t
sudo systemctl restart sshd

sshd_config keys and what they do:

ParameterValuePurpose
Port22220Non-standard port, reduces log noise
PermitRootLoginnoBlocks direct root login
PasswordAuthenticationnoKeys only
MaxAuthTries3Limits authentication attempts
AllowUsersdeployWhitelist of allowed users

Firewall and Access Restrictions

UFW is the simplest way to close everything unnecessary:

sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow 22220/tcp comment "SSH bastion"
sudo ufw enable

If the bastion is only for your IP, restrict it further:

sudo ufw allow from YOUR_IP to any port 22220 proto tcp
Note

If your IP is dynamic, use a VPN instead of exposing the port. Opening SSH to the internet without source restrictions is a bad practice.

For internal traffic, add a rule on the bastion to allow packet forwarding:

sudo sysctl -w net.ipv4.ip_forward=1

Add net.ipv4.ip_forward = 1 to /etc/sysctl.conf so the setting survives a reboot.

Fail2ban and Brute-Force Protection

Install and configure Fail2ban to protect against password guessing:

sudo apt install fail2ban -y
sudo cp /etc/fail2ban/jail.conf /etc/fail2ban/jail.local

In /etc/fail2ban/jail.local:

[sshd]
enabled = true
port = 22220
filter = sshd
logpath = /var/log/auth.log
maxretry = 3
bantime = 3600
findtime = 600

Restart:

sudo systemctl restart fail2ban
sudo systemctl enable fail2ban
sudo fail2ban-client status sshd

Logging and Auditing

The bastion must log everything. On Debian/Ubuntu, SSH logs go to /var/log/auth.log. For centralized collection, configure rsyslog to ship to a separate SIEM or at least a second server:

# Add to /etc/rsyslog.conf:
*.* @logserver_ip:514

For user action auditing, attach auditd:

sudo apt install auditd -y
sudo auditctl -w /etc/ssh/sshd_config -p wa -k ssh_config
sudo auditctl -w /home -p r -k user_files

View events:

sudo ausearch -k ssh_config
sudo ausearch -k user_files
Warning

Logs on the bastion are an attacker’s first target. Set up remote shipping as early as possible — otherwise, on compromise, you lose the incident history.

Port Forwarding and Tunnels Through the Bastion

The bastion’s main job is to give access to internal machines without exposing their ports. Three ways to do it:

1. Port forwarding via SSH tunnel:

ssh -L 5432:10.0.1.5:5432 deploy@bastion_ip -p 22220

Now local port 5432 on your machine proxies to PostgreSQL on 10.0.1.5.

2. SOCKS proxy for access to all internal hosts:

ssh -D 1080 deploy@bastion_ip -p 22220

Point your browser or proxychains at 127.0.0.1:1080.

3. Reverse tunnel for accessing your local machine from the bastion network:

ssh -R 9090:localhost:8080 deploy@bastion_ip -p 22220

This lets someone reach a local service on port 9090 via the bastion.

For persistent tunnels, use autossh or configure ~/.ssh/config:

Host bastion
    HostName bastion_ip
    Port 22220
    User deploy
    IdentityFile ~/.ssh/id_ed25519
    ServerAliveInterval 60
    ServerAliveCountMax 3

Now ssh bastion is all you need to connect, and tunnels are built with a single command.

Tip

For team operations, store keys in 1Password or HashiCorp Vault, and grant bastion access through ephemeral sessions with expiring tokens. This reduces the risk of key compromise.

24 - tmux on Prod After Screen

Why We Switched from Screen to tmux

Screen was our primary tool for about five years. After migrating the cluster to new servers it became obvious: screen drops sessions on SSH disconnect when hardstatus isn’t configured, and screen -r recovery sometimes hits a race condition when multiple admins connect simultaneously. tmux solves both problems out of the box — sessions live in server memory, are bound to a socket, and reconnection doesn’t depend on the TCP connection state.

The migration took half a day: we set up a shared config, distributed keybindings, and validated on staging. We went to production a week later — after tmux survived two incidents where screen would have lost context.

Key Differences from GNU Screen

Note

We’re not retelling Screen’s history. The focus is on what tmux gives differently.

The main difference is architectural. tmux uses a client-server model with a separate server process per session. Screen is also client-server, but its session is bound to the terminal less reliably and is lost more often on disconnect.

AspectScreentmux
Session recoveryscreen -r (may fail)tmux attach -t <name> (stable)
256-color supportLimitedFull, default-terminal "screen-256color"
Input synchronizationmultiuser + acladdset -g allow-rename off + shared sessions
Configuration~/.screenrc~/.tmux.conf
State after disconnectOften lostSession stays alive, socket remains
Scriptingscreen -Xtmux send-keys, tmux split-window

Another thing — tmux has a proper copy-mode. In Screen you had to fight with text selection through escape sequences. In tmux, Ctrl+B [ drops you into scrollback with search.

When to Use tmux, When Not To

Warning

tmux is not a panacea. There are scenarios where it’s excessive or even harmful.

tmux makes sense when:

  • multiple admins work on the same server simultaneously;
  • sessions are long-running (monitoring, deploys, debugging);
  • you need reliable copy-mode and scrollback;
  • automation through the tmux CLI (CI/CD scripts that send commands into a session).

tmux is not needed when:

  • one admin per server, short sessions;
  • memory-constrained systems — a tmux server consumes more RAM than screen (though on modern machines this is negligible);
  • you’re in a containerized environment with an ephemeral filesystem — the tmux config won’t survive container recreation without a volume.
Tip

Inside Docker, prefer docker exec -it <container> bash over running tmux inside the container. tmux in a container only makes sense for stateful services that require persistent shell access.

Production Commands

The default prefix is Ctrl+B. All commands follow it.

Creating and attaching:

tmux new -s prod-db
tmux attach -t prod-db
tmux list-sessions

Managing windows and panes:

Ctrl+B c          # new window
Ctrl+B &          # kill window
Ctrl+B %          # split vertically
Ctrl+B "          # split horizontally
Ctrl+B o          # switch between panes
Ctrl+B z          # toggle fullscreen for current pane

Working with copy:

Ctrl+B [          # enter copy-mode
Space             # start selection
Enter             # copy
Ctrl+B ]          # paste clipboard buffer

Scripted invocation (for automation):

tmux send-keys -t prod-db "systemctl restart api" C-m
tmux split-window -t prod-db -h "tail -f /var/log/api.log"

Session persistence across reboot:

Warning

tmux won’t survive a reboot. If you need a live session after restart, use the tmux-resurrect plugin or a cron script that recreates sessions from a state file.

Minimal ~/.tmux.conf for production:

set -g default-terminal "screen-256color"
set -g base-index 1
set -g pane-base-index 1
set -g allow-rename off
set -g set-titles on
set -g set-titles-string "#S:#I.#P #W"
set -g status-keys vi
set -g mouse on
bind -n M-1 select-pane -t 0
bind -n M-2 select-pane -t 1
bind -n M-3 select-pane -t 2

After editing the config: tmux source-file ~/.tmux.conf.

Bottom Line

tmux replaced Screen on production not because it’s “newer,” but because its sessions don’t get lost on disconnect, copy works without workarounds, and CLI scripting gives predictability for automation. The trade-off is slightly higher memory usage and the need to maintain a single config across all servers. For our stack of 12 production servers with three admins each, it paid off in the first week. The task is to write an English IT blog post for “Lead DevOps” blog, parallel to the Russian draft provided. The English version should be the original article (not a word-for-word translation), same structure and facts, practical tone, same commands and tables.

Key constraints:

25 - bpftrace: one-liners that replace strace in production

strace halts a process on every syscall. On a live server at 2000 RPS, that means timeouts and alerts. bpftrace runs through eBPF in the kernel — tracing happens in parallel, without stopping anything. The overhead difference is orders of magnitude.

Installation

# Debian / Ubuntu
sudo apt install bpftrace

# RHEL / CentOS / Fedora
sudo dnf install bpftrace

# Arch
sudo pacman -S bpftrace

# Verify
sudo bpftrace -V

Full probe coverage requires debug symbols:

# Debian
sudo apt install linux-image-$(uname -r)-dbg
sudo apt install systemtap-sdt-dev

Check available probes:

sudo bpftrace -l | grep sched_process
# sched:sched_process_exec
# sched:sched_process_fork
# sched:sched_process_exit

bpftrace Syntax in 60 Seconds

One-liner format:

sudo bpftrace -e 'probe { action }'

Structure: what (probe) and what to do (action). Probe types:

TypeExampleDescription
kprobekprobe:do_sys_openat2kernel function entry
kretprobekretprobe:do_sys_openat2kernel function return
tracepointsyscalls:sys_enter_openatstable kernel tracepoint
usdtusdt:/bin/python3:probeuser-level static trace
profileprofile:hz:99timer-based sampling

Built-in variables available in actions:

pid          # process ID
tid          # thread ID
comm         # process name
nsecs        # nanoseconds timestamp
curtask      # current task_struct
args         # probe arguments (if available)

Example — all execve calls:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve { join(args->argv); }'

exec: Who Is Launching Processes

Want to understand which process triggers fork/exec across the system:

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_execve {
        time("%H:%M:%S ");
        printf("%s (PID %d) exec: %s\n", comm, pid, args->argv[0]);
    }
'

Output during 10 seconds of monitoring:

19:42:15 bash (PID 12441) exec: /usr/bin/ls
19:42:15 bash (PID 12441) exec: /usr/bin/cat
19:42:17 systemd (PID 1) exec: /usr/sbin/CROND
19:42:17 CROND (PID 8921) exec: /bin/sh
19:42:17 CROND (PID 8921) exec: /usr/sbin/sendmail

Track a specific process and its children:

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_execve /pid == 1234/ {
        printf("child exec: %s\n", args->argv[0]);
    }
'

The filter /pid == 1234/ uses standard syntax. Without it, bpftrace catches everything.

open: Which Files a Process Opens

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_open,
    tracepoint:syscalls:sys_enter_openat {
        printf("%s (PID %d) -> %s\n", comm, pid, str(args->filename));
    }
'

Filter by process name:

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_openat /comm == "nginx"/ {
        @[str(args->filename)] = count();
    }
'

This is aggregation — counting how many times each file was opened. @ is the built-in variable for maps. Output after Ctrl+C shows a sorted table.

Monitor open failures (ENOENT, EACCES):

sudo bpftrace -e '
    tracepoint:syscalls:sys_exit_openat {
        if (args->ret < 0) {
            printf("%s error %d on %s\n", comm, args->ret, str(args->filename));
        }
    }
'

Network: Connections and Dropped Packets

Monitor outbound connections:

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_connect {
        printf("%s (PID %d) connect to port %d\n", comm, pid, args->uservaddr->sin_port >> 8);
    }
'

Dropped iptables packets:

sudo bpftrace -e '
    kprobe:nf_hook_slow {
        @drops[comm] = count();
    }
'
Note

Not all kprobes are available on every kernel. Verify with sudo bpftrace -l | grep nf_hook.

Aggregation by port — a common task:

sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_connect {
        @port = count();
    }
' 2>/dev/null | sort -rn | head -20

Errors: getpid Does Not Exist

Familiar functions may be absent in bpftrace. This is not bash — different rules apply.

Familiar Functionbpftrace Equivalent
getpid()pid
strace -p PIDbpftrace -e '... /pid == N/ {...}'
readlink /proc/PID/fd/Nnsecs, curtask

Calling getpid() inside bpftrace produces a compilation error — the BPF program has no access to libc.

Warning

bpftrace cannot trace a process already running under active strace. They conflict at the ptrace level.

Error output — use strerror():

sudo bpftrace -e '
    tracepoint:syscalls:sys_exit_openat {
        if (args->ret < 0) {
            printf("%s: %s\n", str(args->filename), strerror(-args->ret));
        }
    }
'

bpftrace vs strace: Overhead Comparison

strace uses ptrace(PTRACE_SYSCALL). On every syscall, the kernel stops the process, copies data to userspace, then resumes. This is synchronous.

bpftrace compiles a BPF program and loads it into the kernel. Tracing happens in kernel context without stopping the process. Data accumulates in a ring buffer and is read asynchronously.

Comparison on nginx, 5000 RPS:

Methodp99 LatencyCPU OverheadObservability
No tracing12ms——
strace -p PID340ms18%syscalls
bpftrace one-liner14ms0.3%syscalls + aggregation
Tip

For a quick check: strace -c -p PID gives a summary table of syscalls. bpftrace does the same via count() and hist().

When bpftrace Is Enough vs When You Need strace

bpftrace — for a system-wide view. Monitoring all processes, aggregation, heat maps, catching anomalies without affecting production.

strace — for deep-diving a specific request. Detailed log of every syscall with arguments and returns to reproduce a problem.

# bpftrace: aggregation — who opens the most files
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { @[comm] = count(); }'

# strace: detailed log of a single request
strace -f -e openat -s 200 curl localhost/api/endpoint

Three rules:

  1. Don’t know the process — use bpftrace.
  2. Know the PID and need detailed logs — use strace -p PID.
  3. On production under load — bpftrace only.

One-liners as aliases:

echo 'alias bt="sudo bpftrace"' >> ~/.bashrc
alias bt-who-exec='sudo bpftrace -e "tracepoint:syscalls:sys_enter_execve { printf(\"%s %s\\n\", comm, str(args->argv[0])); }"'
alias bt-files='sudo bpftrace -e "tracepoint:syscalls:sys_enter_openat { @[str(args->filename)] = count(); }"'

bpftrace goes further — kernel memory, CPU profiling, allocator debugging. For a basic start, these four one-liners cover most of what you need.

26 - ip: Network Setup and Diagnostics in CLI

When ifconfig returns nothing and configuring a route requires a separate command, that’s not a system bug. It’s iproute2 — the package that replaced net-tools in modern Linux distributions. The ip utility from iproute2 is the standard interface for managing the Linux network stack. It covers interfaces, addresses, routes, ARP cache, routing policies, and namespace isolation.

Why iproute2 replaced net-tools

net-tools (ifconfig, route, arp, netstat, nameif) originated in BSD and migrated to Linux in the 1990s. By the 2000s it became clear: they cannot handle VLAN, IPsec, QoS, multicast routing, or Policy Routing. Each task required a separate command with unrelated syntax.

iproute2 consolidated everything into one ip utility with subcommands. The Linux kernel communicates with the network subsystem via netlink sockets — ip talks to them directly, while ifconfig parses /proc/net/. In systemd-based distributions (RHEL 7+, Ubuntu 16.04+, Debian 9+) ip is installed by default. net-tools remains in repositories for compatibility, but kernel developers haven’t added new functionality since 2001.

Install if needed: apt install iproute2 or yum install iproute.

ip link operates at L2 — listing and controlling interface state.

ip link show
ip link show eth0
ip link show type bridge

Output shows index, name, MAC address, MTU, state (UP/DOWN), and error/packet counters.

Bring an interface up or down:

ip link set eth0 up
ip link set eth0 down
Warning

Taking down an interface severs connectivity. When working remotely, wrap in a script with a timeout and auto-recovery.

Set MTU, change MAC, or rename:

ip link set eth0 mtu 9000
ip link set eth0 address 02:42:ac:11:00:02
ip link set eth0 name enp0s3

Create virtual interfaces (VETH pair for namespace or bridge):

ip link add veth0 type veth peer name veth1
ip link add br0 type bridge
ip link set veth0 master br0

Delete an interface:

ip link del veth0
FlagPurpose
showdisplay interfaces (shorthand: ip l)
setmodify interface parameters
add / delcreate or delete virtual interface
masterattach interface to a bridge

ip addr: address binding and diagnostics

ip addr manages IP addresses (L3).

ip addr show
ip addr show eth0

Add an address:

ip addr add 192.168.1.10/24 dev eth0

Add a secondary address (alias) on the same interface:

ip addr add 192.168.1.11/24 dev eth0
Note

Secondary addresses in Linux are not aliases in the ifconfig sense — they are part of a single address entity. The command ifconfig eth0:0 created a pseudodevice with a separate name; ip works differently.

Remove an address:

ip addr del 192.168.1.10/24 dev eth0

Flush all addresses from an interface:

ip addr flush dev eth0

Useful during reconfiguration: clear old addresses and assign new ones without restarting the service.

Specify scope and label:

ip addr add 10.0.0.5/8 dev eth0 scope host label eth0:internal

scope host — address only for local sockets; scope global — routable.

SubcommandAction
addassign address
delremove address
showdisplay addresses
flushclear interface addresses

ip route: default and static routes

ip route works with the routing table.

ip route show

Add a default route (gateway):

ip route add default via 192.168.1.1 dev eth0

Add a specific route:

ip route add 10.20.0.0/16 via 192.168.1.254 dev eth0

Route to a host via direct ARP (no route, L2 only):

ip route add 192.168.1.50/32 dev eth0

Remove a route:

ip route del default via 192.168.1.1

Replace a route (updates if exists, creates if not):

ip route replace default via 10.0.0.1 dev eth0

Get the route the kernel will choose for an address:

ip route get 8.8.8.8

Add a route to a different table (default is table 254):

ip route add default via 10.0.0.1 dev eth0 table 100
SubcommandPurpose
show / listdisplay routing table
addadd a route
delremove a route
replacemodify or create a route
getshow route to an address
flushclear route cache

ip neigh: ARP/NDP cache

ip neigh manages the neighbour table — ARP for IPv4, NDP for IPv6.

ip neigh show
ip neigh show dev eth0

Add a static ARP entry:

ip neigh add 192.168.1.1 lladdr 00:11:22:33:44:55 dev eth0 nud permanent

nud (Neighbour Unreachability Detection) defines the state:

  • permanent — entry never expires
  • noarp — managed by protocol but not removed
  • reachable / stale / delay / probe — automatic states

Remove an entry:

ip neigh del 192.168.1.1 dev eth0

Flush all neighbours on an interface:

ip neigh flush dev eth0
Tip

After changing a gateway MAC address, flushing the ARP cache speeds up connectivity recovery: ip neigh flush dev eth0.

ip rule: routing policies

ip rule determines which routing table is used for a packet.

ip rule show

Standard output:

0:      from all lookup local
32766:  from all lookup main
32767:  from all lookup default

Add a rule for source IP:

ip rule add from 10.0.0.5 table 100

Rule for incoming interface:

ip rule add iif eth0 table 100

Remove a rule:

ip rule del from 10.0.0.5 table 100
Note

Rules are checked in order (priority). Low number means high priority. Add rules with priority between existing ones if ordering matters.

ActionPurpose
fromsource IP or CIDR
todestination IP or CIDR
iifincoming interface
lookuprouting table
prionumeric priority

ip maddr: multicast addresses

ip maddr displays and manages multicast groups on an interface.

ip maddr show eth0

Add an interface to a multicast group:

ip maddr add 239.0.0.1 dev eth0

Remove:

ip maddr del 239.0.0.1 dev eth0

Unlike unicast, multicast addressing is used in broadcast domains, routing protocols (OSPF, RIP), discovery services, and streaming. In most tasks this command is unnecessary, but when configuring clusters or monitoring via specific protocols — it will be required.

ip netns: network stack isolation

ip netns creates isolated network namespaces. Each namespace has its own interfaces, addresses, routes, ARP table, and rules.

Create a namespace:

ip netns add testns

Run a command inside a namespace:

ip netns exec testns ip link show

Bring up an interface in a namespace:

ip netns exec testns ip link set lo up
ip netns exec testns ip addr add 127.0.0.1/8 dev lo

Move a VETH interface into a namespace:

ip link set veth1 netns testns

Delete a namespace:

ip netns del testns

List namespaces:

ip netns list
Tip

Containers (docker, podman, LXC) use netns underneath. If a container has no network — check the host namespace: ip netns exec <container_pid> ip addr.

ifconfig, arp, route — legacy compatibility

net-tools is formally available in all major distribution repositories. The source code is unmaintained, but packages persist for compatibility.

Legacy commandEquivalent ipStatus
ifconfigip addr, ip linkdeprecated
route -nip routedeprecated
arp -aip neighdeprecated
netstat -tulpnss -tulpndeprecated
nameifip link namedeprecated

ss from iproute2 replaces netstat — faster and provides more socket information.

Warning

Network initialization scripts in old distributions may rely on ifconfig. On modern systems, systemd-networkd, NetworkManager, and cloud-init use ip directly or through their own abstractions.

If ifconfig is missing in a fresh distribution — that’s expected. Configuration via ip covers all current scenarios: from address assignment to complex routing policies and service isolation via namespaces.

27 - nslookup and drill: DNS resolution in terminal

The server won’t resolve a domain, but pings fly through. No familiar dig at hand — the BIOS is already loading a minimal busybox. Or on a host without bind-tools. nslookup and drill fill this gap: the first one is built into almost everything, the second gives more context when debugging.

nslookup: interactive and one-liner modes

nslookup ships with bind-utils and isc-dhcp-client. It works in two modes.

One-liner query:

nslookup example.com
nslookup example.com 8.8.8.8

Interactive mode starts with no arguments. Typical session:

$ nslookup
> server 1.1.1.1
Default server: 1.1.1.1
Address: 1.1.1.1#53
> set type=MX
> example.com
Server:         1.1.1.1
Address:        1.1.1.1#53

example.com     mail exchanger = 10 mx1.example.com.
> exit

Switching servers inside a session only changes the resolver for that query. If you need a permanent resolver — edit /etc/resolv.conf.

DNS record types in queries

By default nslookup queries A records. For other types use set type=:

Record typePurposeExample output
AIPv4 address93.184.216.34
AAAAIPv6 address2606:2800:220:1::
MXMail exchanger10 mail.example.com
TXTText records, SPFv=spf1 include:_spf.example.com ~all
NSAuthoritative serversa.iana-servers.net
SOAStart of Authorityserial 2005080901
CNAMECanonical nameexample.com canonical name = www.example.com
PTRReverse resolution34.216.184.93.in-addr.arpa name = example.com

One-liner equivalent — -type= flag:

nslookup -type=ANY example.com
nslookup -type=MX github.com
Warning

ANY queries are often blocked at the resolver level. The recursor returns SERVFAIL or an empty response. Do not rely on ANY when troubleshooting.

drill: output with Resource Record type

drill is part of ldns. It returns results in classic DNS format with ANSWER, AUTHORITY, ADDITIONAL sections:

drill A example.com
drill MX github.com @1.1.1.1

drill output is more readable when tracing a CNAME chain:

$ drill CNAME www.cloudflare.com
;; ->>HEADER<<- opcode: QUERY, rcode: NOERROR
;; flags: qr rd ra; QUERY: 1, ANSWER: 1, AUTHORITY: 0, ADDRIBUTE: 0

;; QUESTION SECTION:
;www.cloudflare.com.   IN   CNAME

;; ANSWER SECTION:
www.cloudflare.com.  300  IN  CNAME  cloudflare.com.

;; AUTHORITY SECTION:
;; ADDITIONAL SECTION:

Without the @server flag, drill reads the resolver from /etc/resolv.conf.

DNSSEC validation

drill checks the DNSSEC trust chain:

drill -S _dmarc.example.com TXT @1.1.1.1

The -S flag requests the DS record higher in the chain and validates the signature. On an invalid chain:

drill: RRSIG validation failed: Signature has expired

nslookup does not validate DNSSEC — it only sends queries with the DO flag (include RRSIG in the response). For full validation you need drill or delv.

NXDOMAIN and SERVFAIL: reading response codes

First step on any error — look at the response code.

NXDOMAIN (code 3) — the domain does not exist. Source: authoritative server for the zone. If dig +short returns nothing, and nslookup says ** server can't find example.invalid, that’s NXDOMAIN. Causes: typo in the domain, stale CNAME, deleted zone.

SERVFAIL (code 2) — the resolver couldn’t answer. Causes: broken DNSSEC validation, exceeded timeout, circular reference in NS records, overloaded authoritative server. nslookup shows ** server can't find example.com: Server failed.

REFUSED (code 5) — the recursor refused to answer. Usually ACL on the DNS server or rate limiting.

nslookup example.com 10.0.0.1
# Server:  10.0.0.1
# Address: 10.0.0.1#53
# ** server can't find example.com: Server failed

Key flags for nslookup and drill

nslookup

FlagEffect
-type=RRRecord type (A, MX, TXT, ANY)
hostRedirect to specified server
-port=53Non-standard port (e.g. 5353 for mDNS)
-timeout=5Timeout in seconds
-retry=3Number of retries
-vcTCP instead of UDP
nslookup -type=TXT -port=5353 _http._tcp.local 224.0.0.251

drill

FlagEffect
@serverServer to query
-QQuiet mode, answer only
-TShow response time
-SDNSSEC validation
-DForce DNSSEC (query with DO flag)
-p portNon-standard port
-t timeoutTimeout in seconds
drill -TD -S TXT dkim._domainkey.example.com @8.8.8.8

-T is useful for comparing latency between resolvers:

drill A google.com @1.1.1.1
# Query timeout: 2
# Answer received in 45ms

Installation

# Debian / Ubuntu
apt install dnsutils ldnsutils

# RHEL / CentOS / Fedora
dnf install bind-utils ldns

# Alpine
apk add bind-tools ldns

In minimal busybox images you already have a simplified nslookup. The full feature set is available after installing dnsutils.

28 - journalctl: Filtering and Formatting systemd Logs

Logs disappeared. Server rebooted, and the familiar less /var/log/syslog returns nothing. On modern distros with systemd, logs are collected by journald and read with journalctl. Without knowing its filters, system debugging turns into guesswork.

Why Logs Disappear After Reboot

By default, journal stores data in /run/log/journal/ — a tmpfs that wipes on reboot. To make logs survive reboots, create the directory:

sudo mkdir -p /var/log/journal
sudo systemd-tmpfiles --create --prefix /var/log/journal

Then restart systemd-journald:

sudo systemctl restart systemd-journald

Check current location and size:

journalctl --disk-usage
Note

Fresh CentOS/RHEL 8+ and Fedora create /var/log/journal automatically. Debian and Ubuntu typically do not.

Filtering by Unit and Time Range

The most common case — logs for a specific service:

journalctl -u nginx.service
journalctl -u postgresql@main.service

Combine multiple units by repeating the flag:

journalctl -u nginx.service -u php-fpm.service

Time filters are for incident debugging:

# Last hour
journalctl --since "1 hour ago"

# Specific day
journalctl --since "2025-01-15" --until "2025-01-15 23:59:59"

# Last 24 hours
journalctl --since "yesterday"

# From 08:00 to 09:00
journalctl --since "today 08:00" --until "today 09:00"

Night crash? Look at logs from that period, not the entire buffer.

Warning

If a time filter returns empty output, check the timezone. journalctl stores timestamps in UTC, but --since interprets local time.

Filtering by Priority

Log levels match syslog:

LevelNumberDescription
emerg0System unusable
alert1Immediate action required
crit2Critical condition
err3Error
warning4Warning
notice5Normal but significant
info6Informational
debug7Debug-level messages
# Errors and critical only
journalctl -p err -l

# Warnings through errors
journalctl -p warning..err

# Everything from notice upward
journalctl -p notice

The -l flag shows full hostnames instead of truncated ones.

Kernel and Boot Logs

The kernel sends its messages separately. The -k flag replaces dmesg:

# Kernel messages for current boot
journalctl -k

# Kernel messages for previous boot
journalctl -k -b -1

List all boots:

journalctl --list-boots

Output:

-2 5d3c1a9... Mon 2025-01-13 08:00:00 — Mon 2025-01-13 18:00:00
-1 a7b2d8f... Mon 2025-01-13 18:05:00 — Tue 2025-01-14 08:00:00
 0 c9e1f3a... Tue 2025-01-14 08:05:00 — currently running

Select a specific boot:

journalctl -b 5d3c1a9...

For boot analysis, use systemd-analyze:

systemd-analyze blame | head -20
systemd-analyze critical-chain nginx.service

Piping journalctl output to grep loses metadata. Use -g (–grep) instead:

# Search for DB connection failures
journalctl -g "connection.*failed" -u myapp.service

# Authentication errors
journalctl -g "auth.*fail" -p err

The -g flag supports basic regex. For complex conditions, combine with --since:

journalctl -u nginx.service --since "1 hour ago" | grep -E "(timeout|502|503)"

This keeps the unit and time selection intact, then filters by pattern.

Follow Mode (-f)

Analogous to tail -f for journald. Unlike watching a log file, follow works with any filter:

journalctl -u nginx.service -f
journalctl -f -p err
journalctl -u nginx.service -p err -f

Ctrl+C stops follow in a terminal. From scripts, wrap with timeout or send a signal.

Tip

Run -f in a separate tmux/screen pane. If the pane closes, logs keep going to journald — you won’t lose data.

Flags combine with AND: -u nginx -p err shows errors from nginx only. For OR across units, use journal fields:

journalctl --no-pager _SYSTEMD_UNIT=nginx.service OR _SYSTEMD_UNIT=php-fpm.service -p err

Other useful fields:

# By UID
journalctl --no-pager _UID=1000

# By executable
journalctl --no-pager _EXE=/usr/sbin/nginx

# List values seen for a field
journalctl --no-pager -F _SYSTEMD_UNIT

Output Formats

By default, journalctl paginates output. For scripts and piping to jq, you need machine-readable format:

# JSON Lines (jq-friendly)
journalctl -u nginx -n 50 -o json

# JSON with pretty structure
journalctl -u nginx -n 50 -o json-pretty
FlagDescriptionUse Case
-o shortClassic syslogDefault
-o short-isoISO 8601 timestampsSIEM logging
-o short-preciseMillisecond precisionPrecise timing
-o verboseAll fieldsMaximum detail
-o jsonJSON Linesjq, Splunk, ELK
-o catMESSAGE field onlyMinimal output
# Messages only, no metadata — equivalent to tail -f /var/log/app.log
journalctl -u myapp -f -o cat
Tip

-n 100 limits output to the last 100 lines. --no-pager disables pagination for scripts.

Cleanup and Size Management

journald rotates logs by size and time. Configure in /etc/systemd/journald.conf:

[Journal]
SystemMaxUse=500M
SystemMaxFileSize=50M
MaxRetentionSec=30day

Apply without restart:

sudo systemd-tmpfiles --create /etc/tmpfiles.d/journald.conf
sudo killall -USR1 systemd-journald

Free up space manually:

# Show disk usage
journalctl --disk-usage

# Delete logs older than N days
sudo journalctl --vacuum-time=7days

# Delete logs, keeping last N megabytes
sudo journalctl --vacuum-size=200M

# Delete old journal files (not current)
sudo journalctl --vacuum-files=5
Warning

--vacuum-* only removes files exceeding the limit. To free space reliably, increase SystemMaxUse and restart journald.

Common Errors

journalctl: cannot open files — insufficient permissions. Add yourself to the systemd-journal group:

sudo usermod -aG systemd-journal $USER
# re-login

Logs empty after reboot — persistent storage not configured (see first section).

journalctl hangs — huge buffer. Start with -b or limit with --since.

No unit logs — check that the unit actually ran:

systemctl status nginx
journalctl -u nginx --no-pager -n 20

journalctl is built for fast searching. Don’t read logs manually — filter from the start.

29 - OpenSSL: TLS Certificate Verification and Parsing in CLI

Certificates expiring on prod at the worst moment — a familiar story. OpenSSL answers TLS certificate questions faster than any marketplace checker. Here are the key scenarios without the fluff.

Basic Certificate Parsing

The first command for any diagnostics is the text dump:

openssl x509 -text -noout -in cert.pem

Output shows Subject, Issuer, validity dates, signature algorithm, and public key. For a quick summary without the wall of text:

# Subject only
openssl x509 -noout -subject -in cert.pem

# Issuer only
openssl x509 -noout -issuer -in cert.pem

# Fingerprint only (SHA-256)
openssl x509 -noout -fingerprint -sha256 -in cert.pem

The -in flag accepts a file path. Certificates downloaded from browsers usually come in PEM or DER format. OpenSSL handles both, but DER requires an extra flag:

openssl x509 -inform DER -in cert.der -text -noout
Note

-noout suppresses the base64 block from output. Useful when you need only structured data, not a copy of the certificate.

Expiration: dates and Overdue Checks

For monitoring, getting only the dates is more convenient:

openssl x509 -noout -dates -in cert.pem

Typical output:

notBefore=Jan 15 00:00:00 2024 GMT
notAfter=Jan 14 23:59:59 2025 GMT

For automation, extracting the timestamp and calculating the difference is cleaner:

# Days remaining until expiry
not_after=$(openssl x509 -noout -enddate -in cert.pem | cut -d= -f2)
days_left=$(( ($(date -d "$not_after" +%s) - $(date +%s)) / 86400 ))
echo "$days_left days left"

If days_left is negative, the certificate has already expired.

Tip

For checking multiple hosts from inventory, a one-liner works well:

for host in api.example.com admin.example.com; do
  echo -n "$host: "
  echo | openssl s_client -servername "$host" -connect "$host":443 2>/dev/null \
    | openssl x509 -noout -enddate
done

-servername sends SNI — without it some hosts return the default certificate.

Chain Verification: s_client and verify

Connect and display certificates:

openssl s_client -connect example.com:443 -showcerts </dev/null

Output includes the chain from leaf certificate to root CA. To filter certificates only:

openssl s_client -connect example.com:443 -showcerts </dev/null \
  | awk '/-----BEGIN/,/-----END/{if(/-----BEGIN/)a=1;a;if(/-----END/)a=0}' > chain.pem

Verify chain against the system store:

openssl verify -CAfile /etc/ssl/certs/ca-certificates.crt chain.pem

If verification fails with error 20 at 0 depth lookup, the intermediate CA is missing. A common cause is incorrect chain configuration on the server.

Warning

openssl verify uses the system store by default. In Ubuntu this is /etc/ssl/certs/ca-certificates.crt, in Alpine it is a separate ca-certificates package. If verification fails, check that the package is installed.

Quick check without saving to file:

echo | openssl s_client -connect example.com:443 2>/dev/null \
  | openssl verify

Output Verify return code: 0 (ok) means success.

To stay out of interactive mode and fail the process on a bad chain:

openssl s_client -connect example.com:443 -servername example.com \
  -quiet -verify_return_error </dev/null

Force a protocol or cipher when you need to confirm the server still accepts a specific handshake:

openssl s_client -connect example.com:443 -servername example.com \
  -tls1_2 -cipher ECDHE-RSA-AES256-GCM-SHA384 -quiet </dev/null

For an internal CA, pass the bundle explicitly. Without -CAfile, a private PKI typically returns Verify return code: 21 (unable to get local issuer certificate):

openssl s_client -connect example.com:443 -servername example.com \
  -CAfile /etc/ssl/certs/ca-bundle.crt -verify_return_error -quiet </dev/null

Extracting CN and SAN

Common Name extracts directly:

openssl x509 -noout -subject -in cert.pem | grep -oP '(?<=CN = )[^,]+'

But CN has not been sufficient for years — modern certificates use Subject Alternative Names (SAN). OpenSSL 1.1.1+ pulls them cleanly:

openssl x509 -noout -ext subjectAltName -in cert.pem

Output:

X509v3 Subject Alternative Name:
    DNS:example.com, DNS:www.example.com, DNS:api.example.com, IP:192.0.2.1

To get only the DNS names:

openssl x509 -noout -ext subjectAltName -in cert.pem \
  | grep -oP '(?<=DNS:)[^,]+'
Note

If SAN is missing (old certificate), browsers fall back to CN. When checking API endpoints, this explains why curl complains but the browser opens the page.

Comparing Expiration Across Multiple Hosts

A script for checking a host list serves as a practical monitoring foundation:

#!/bin/bash
# check-certs.sh — check certificate expiration dates

check_host() {
  local host=$1
  local port=${2:-443}
  
  echo | timeout 5 openssl s_client -servername "$host" -connect "$host:$port" 2>/dev/null \
    | openssl x509 -noout -enddate 2>/dev/null \
    | cut -d= -f2 \
    | while read date; do
        ts=$(date -d "$date" +%s)
        now=$(date +%s)
        days=$(( (ts - now) / 86400 ))
        printf "%-30s %3d days  %s\n" "$host" "$days" "$date"
      done
}

# Example usage
for h in api.example.com admin.example.com legacy.internal; do
  check_host "$h"
done

Typical output:

api.example.com                   45 days  Jan 14 23:59:59 2025 GMT
admin.example.com               -12 days  Dec  1 23:59:59 2024 GMT
legacy.internal                -120 days  Aug  5 23:59:59 2024 GMT

Negative values are expired certificates. In production, wrapping this in cron with chat notifications at a 30-day threshold is practical.

Quick Flag Reference

CommandFlagPurpose
x509-textFull text dump
x509-nooutSuppress base64 block
x509-datesNotBefore, NotAfter
x509-subjectSubject (CN, O, OU)
x509-issuerIssuing CA
x509-fingerprint -sha256Certificate fingerprint
x509-enddateExpiration date only
x509-ext subjectAltNameAlternative names
s_client-connect host:portTLS connection
s_client-servername nameSNI (required for vhost)
s_client-showcertsDisplay full chain
s_client-quietSkip interactive mode
s_client-tls1_2 / -tls1_3Force a TLS version
s_client-verify_return_errorNon-zero exit on validation failure
verify-CAfile pathTrusted CA file
verify-partial_chainAccept partial chain
Tip

openssl s_client does more than read certificates. With -starttls smtp or -starttls pop3 it checks mail servers. -http retrieves HTTP headers over TLS. Useful for diagnosing miTM filters.

All commands work out of the box in any Linux distribution. No dependencies beyond OpenSSL itself — the tool is present on every server. If missing, install in seconds: apt install openssl or apk add openssl.

30 - ProxyJump and bastion hosts via ~/.ssh/config

Sometimes a server sits in a private network with no public IP. The only entry point is a bastion host with a public address. Typing ssh -J user@bastion user@private every time gets old fast. Here’s how to configure everything in ~/.ssh/config so you can reach private networks in one command.

Why you need a bastion host

A bastion (jump host, jump box) is an intermediate server with public access that proxies connections to infrastructure without external IPs. The typical topology:

Laptop → Bastion (public IP) → Private server (10.0.1.5)

The bastion doesn’t need to be hardened like a fortress — it’s just an open relay point. Access control lives on SSH keys and, if needed, security groups or firewall rules.

Note

The bastion host is not a terminal destination — it only proxies traffic. You don’t need to run a VPN or additional services on it.

ProxyJump — the modern syntax

-J (ProxyJump) arrived in OpenSSH 7.3. The parameter takes a host in [user@]host[:port] format and spins up a SOCKS5 proxy through the specified node.

Basic invocation:

ssh -J user@bastion.example.com user@10.0.1.5

Authentication on both hosts via keys. If the username matches, you can omit it:

ssh -J bastion.example.com 10.0.1.5

With a port other than 22:

ssh -J bastion.example.com:2222 10.0.1.5

The same thing in the config file:

Host private-server
    HostName 10.0.1.5
    ProxyJump bastion.example.com

After that, ssh private-server connects through the bastion automatically.

ProxyCommand — the classic approach

ProxyJump is a wrapper around ProxyCommand. When you need more control or are working with older OpenSSH, use ProxyCommand directly.

ssh -o ProxyCommand="ssh -W %h:%p bastion.example.com" 10.0.1.5

-W forwards stdin/stdout to the target host. In the config:

Host private-server
    HostName 10.0.1.5
    ProxyCommand ssh -W %h:%p bastion.example.com

The difference from ProxyJump is minimal, but ProxyCommand lets you inject variables, conditions, and command chains.

Multiple hops in a row

A chain of two bastion hosts:

ssh -J bastion1.example.com,bastion2.example.com 10.0.1.5

In the config:

Host private-server
    HostName 10.0.1.5
    ProxyJump bastion1.example.com,bastion2.example.com

OpenSSH connects hosts sequentially: laptop → bastion1 → bastion2 → private-server. Make sure your keys are present on each node.

For complex scenarios, ProxyCommand with nc (netcat) gives more flexibility:

Host dmz-server
    HostName 192.168.1.10
    ProxyCommand ssh -W %h:%p bastion.example.com

Host private-server
    HostName 10.0.1.5
    ProxyCommand ssh -W %h:%p dmz-server

The chain works, but each hop adds latency. For interactive work, more than two hops signals a network architecture problem.

Complete config example

# Bastion host (public entry point)
Host bastion
    HostName bastion.example.com
    User admin
    Port 22
    IdentityFile ~/.ssh/id_ed25519
    ForwardAgent yes
    ServerAliveInterval 60
    ServerAliveCountMax 3

# Private server via bastion
Host private-web
    HostName 10.0.1.5
    User appuser
    ProxyJump bastion
    IdentityFile ~/.ssh/id_ed25519
    ServerAliveInterval 60
    ServerAliveCountMax 3

# Private database
Host private-db
    HostName 10.0.2.10
    User dbadmin
    ProxyJump bastion
    IdentityFile ~/.ssh/id_ed25519
    LocalForward 5433 127.0.0.1:5432
Tip

ForwardAgent yes on the bastion lets your agent forward keys further down the chain. Don’t enable it if you don’t trust the bastion machine.

LocalForward in the example tunnels PostgreSQL from the private server to local localhost:5433. Useful for connecting IDEs or psql.

Verify the config parses without errors:

ssh -G private-web | grep -E '^(hostname|proxyjump)'

If you see the correct values — the config was picked up.

Common mistakes and solutions

Connection timeout during ProxyJump

Verify the bastion is reachable directly:

ssh bastion.example.com echo ok

If that fails — the problem is in the network, not the config.

Permission denied (publickey) on bastion

Make sure the key is loaded in the ssh-agent:

ssh-add ~/.ssh/id_ed25519
ssh-add -l

If the agent is empty — add the key and test with ssh -vT bastion.

Works one way, not the other

ProxyJump tunnels TCP. ICMP (ping) won’t pass through. Check connectivity with nc -zv host port or ssh -v.

Agent refused operation during agent forwarding

Check the SSH_AUTH_SOCK variable:

echo $SSH_AUTH_SOCK

If empty — start the agent:

eval "$(ssh-agent -s)"
ssh-add

Slow connection through a chain

Check MTU. Sometimes MTU in VPN/LAN is smaller than required for TCP-over-TCP. Add to the config:

Host *
    IPQoS lowdelay throughput

For very slow links, try compression:

Host *
    Compression yes

ProxyJump covers 90% of use cases. If you need visualization or a UI manager — look at ssh-config tools or Terminator, but for console work ~/.ssh/config with ProxyJump is sufficient.

31 - auditd: file access and syscall logging

Linux doesn’t write every access to /etc/shadow or every unlink call to syslog. For incident investigation and compliance this is critical. auditd solves this: the Linux Audit kernel subsystem records system calls, file access, and more.

Installation and Startup

auditd comes in the audit package available in any distribution.

# Debian/Ubuntu
apt install auditd

# RHEL/CentOS/Alma
yum install audit

# Arch
pacman -S audit

After installation, start the service via systemd.

systemctl enable --now auditd

Check status and current rules:

systemctl status auditd
auditctl -l
Note

On RHEL-based distributions with SELinux enabled, you may need to adjust policies for auditd to work with non-standard paths. Standard installation usually covers most cases.

File Monitoring: the -w Flag

The -w flag adds a watch rule for a file. By default it tracks open, read, write, truncate, chmod, and chown.

# Watch the password file
auditctl -w /etc/shadow -p rwxa -k shadow_access

# Watch the nginx config directory
auditctl -w /etc/nginx/ -p rwxa -k nginx_config
FlagMeaning
-wpath to watch
-ppermissions: r(read), w(write), x(execute), a(append)
-kkeyword for searching in logs

Verify rules:

auditctl -l

Remove a rule by key:

auditctl -W /etc/shadow -p rwxa -k shadow_access

Rules added via auditctl don’t survive reboots. For persistence, write rules to /etc/audit/rules.d/:

echo "-w /etc/shadow -p rwxa -k shadow_access" >> /etc/audit/rules.d/audit.rules

On RHEL, rules load from /etc/audit/audit.rules via the augenrules script. Same works on Debian/Ubuntu.

System Call Logging: the -S Flag

The -S flag records the specified system call for all processes or with filters.

# Log file deletions
auditctl -S unlink -S unlinkat -k file_deletion

# Log socket creation
auditctl -S socket -k network_socket

Check available system calls with ausyscall --dump. Not all calls are available on every architecture — on x86_64 some go through the compat layer.

Warning

Excessive syscall monitoring generates massive log volume. On production servers, limit rules with filters.

Combined rule — syscall plus path:

# Only deletions from /var/log/
auditctl -S unlink -S unlinkat -w /var/log/ -p wa -k log_deletion

Filtering by UID and Executable

Without filters, rules apply globally. Add conditions for targeted monitoring.

# Only rm executions by www-data user
auditctl -S execve -a always,entry -F arch=b64 -F uid=33 -F exe=/usr/bin/rm -k rm_by_www

# All access() calls to file from any uid
auditctl -a always,entry -S access -F path=/etc/shadow -F perm=r -k shadow_read

Core filter fields:

FlagDescriptionExample
-Ffield to compare-F uid=1000
archarchitecture (b32/b64)-F arch=b64
uidreal UID-F uid=33
euideffective UID-F euid=0
exefull path to executable-F exe=/bin/bash
permaccess permissions-F perm=awx

Combine filters into a chain via -a:

auditctl -a always,entry -S openat -F dir=/etc -F perm=w -F uid=0 -k etc_write_root

Reading Logs: ausearch

Logs live in /var/log/audit/audit.log. Binary format, read with ausearch.

# Search by keyword
ausearch -k shadow_access

# Search by time (today, last hour)
ausearch -k shadow_access -ts today
ausearch -k shadow_access -ts recent

# Search by user
ausearch -k shadow_access -ui 0

# Search by event type
ausearch -m SYSCALL -k file_deletion

# Filter by result (success/failure)
ausearch -k shadow_access -sv success
ausearch -k shadow_access -sv failed

Useful output formats:

# Raw text (default)
ausearch -k shadow_access -i

# CSV for parsing
ausearch -k shadow_access --format csv
Tip

For automation, use -if (input file) — read from a dump instead of the live log:

ausearch -if /tmp/audit_events.dump -k shadow_access

Reading Logs: aureport

aureport aggregates logs into readable reports.

# Summary of all events
aureport

# System call report
aureport -s

# File event report
aureport -f

# User report
aureport -u

# Timeline report
aureport -t

# Errors and failures only
aureport --failed

Typical aureport -s output:

Syscall Report
============================================
UID        Syscall    Count
--------------------------------------------
0          unlink      12
0          openat      8
33         unlink      3

Quick investigation combo — summary then details:

aureport -t -i | head -20
ausearch -k file_deletion -ts recent | less

For SIEM or ELK ingestion, convert logs to JSON or text:

ausearch -k shadow_access --format json > /var/log/audit/shadow_access.json

auditd doesn’t require complex setup to start capturing critical events. Install the package, add a few rules with keywords, and get comfortable with ausearch and aureport for log review.

32 - logrotate: automatic log rotation and archiving

Application logs fill up disk space within a week, and manually running rm *.log is a recipe for trouble. logrotate handles this automatically: it rotates, compresses, and deletes old files on a schedule. Let’s see how it works and how to set it up in five minutes.

How It Works

logrotate runs daily through cron. The default config lives in /etc/logrotate.conf, and additional configs are included from /etc/logrotate.d/. During rotation, the current file gets renamed, a new empty one is created, old copies are compressed and numbered.

The cycle looks like this:

app.log        →  app.log.1      (compressed: app.log.1.gz)
app.log.1.gz   →  app.log.2.gz
...
app.log.5.gz   →  deleted

The mechanism relies on rename or mv, so the process must hold the file descriptor open. If rotation isn’t picked up by the application, you end up with duplication or an empty log.

Configuration Structure

# /etc/logrotate.conf — global settings
weekly          # rotate weekly
rotate 4        # keep 4 copies
compress        # compress old logs
include /etc/logrotate.d/

Files from /etc/logrotate.d/ override global values for specific logs. The format is straightforward:

/path/to/log {
    directive value
    ...
}

Directives inherit from the global config unless overridden. You can specify multiple paths separated by spaces or use a glob pattern — convenient for rotating all application logs in one block.

Key Directives

Basic parameters covering 90% of use cases:

DirectivePurposeExample
rotate Nnumber of copies to keeprotate 7
size Nsize threshold to trigger rotationsize 100M
missingoknot an error if file is missing—
notifemptyskip empty logs—
compresscompress old logs (default gzip)—
dateextuse date instead of number in name—
dateformatdate formatdateformat -%Y%m%d
postrotate ... endscriptcommands after rotationreload service
prerotate ... endscriptcommands before rotationprepare directories

size overrides weekly/monthly/daily when set. Rotation happens when the file reaches the specified size AND the interval has passed. So size 100M with weekly means rotation no earlier than a week and only if the file exceeds 100M.

Note

dateext is incompatible with long filenames on some filesystems. If the log name plus date suffix exceeds 255 bytes — logrotate will fail.

Example Config for an Application

Assume your myapp service writes to /var/log/myapp/. Config:

/var/log/myapp/*.log {
    daily
    rotate 14
    size 50M
    missingok
    notifempty
    compress
    dateext
    dateformat -%Y%m%d-%s
    sharedscripts
    postrotate
        systemctl reload myapp > /dev/null 2>&1 || true
    endscript
}

sharedscripts ensures postrotate runs once for all rotated files, not once per file. Without it, the script executes as many times as files match the rotation criteria.

For Python applications using standard library rotation:

/var/log/myapp/app.log {
    su root myapp
    daily
    rotate 7
    size 200M
    missingok
    notifempty
    compress
    postrotate
        /usr/bin/pkill -HUP -f "python.*myapp" || true
    endscript
}

su changes the rotation process owner — useful when the application runs under a separate user with restricted log permissions.

Debugging and Manual Execution

Dry-run mode shows what would happen without making changes:

logrotate -d /etc/logrotate.d/myapp

Output contains each decision: which file gets renamed, which gets compressed, which commands run. Look for renaming and running postrotate script lines.

Force rotation bypassing the schedule:

logrotate -f /etc/logrotate.d/myapp

-f ignores the last rotation time and size conditions. Combining with -d is a safe way to verify before production:

logrotate -d -f /etc/logrotate.d/myapp

To rotate a specific log outside the schedule but respecting conditions — use the state file:

# check state
cat /var/lib/logrotate/status

# temporarily shift last rotation time
sed -i 's|/var/log/myapp/app.log.*|/var/log/myapp/app.log 2024-01-01-00:00:00|' /var/lib/logrotate/status
logrotate /etc/logrotate.d/myapp

Syntax check without execution:

logrotate -d /etc/logrotate.conf

If there are no errors — output is empty (without -d) or shows the action plan (with -d).

Warning

Do not edit /var/lib/logrotate/status manually in production without understanding the format. One mistake — and logrotate decides rotation already happened, skipping all files until the next cron run.

Default cron setup:

# /etc/cron.daily/logrotate
#!/bin/sh
test -x /usr/sbin/logrotate || exit 0
/usr/sbin/logrotate /etc/logrotate.conf

On most distros you don’t need to touch this file. If you need rotation more often than once a day — add to /etc/cron.hourly/ or write a separate cron job.

33 - sshd_config: baseline for a test stand

SSH access to a test stand often gets opened in a hurry, and then the logs fill with brute-force attempts. A baseline sshd_config that blocks common attack vectors fits into five parameters and twenty minutes.

Why Change Defaults

Distribution-provided sshd ships with permissive settings: root login via password, no user restrictions, three authentication attempts. On a test stand this is tolerable until the logs show:

Failed password for root from 1.2.3.4 port 42341 ssh2
Failed password for root from 1.2.3.4 port 42342 ssh2
Failed password for root from 1.2.3.4 port 42343 ssh2

A local network is not a trusted network. Default configuration means risk and noise in monitoring.

Checking Current Values

Before editing, inspect what’s already set:

sshd -T | grep -E '^(permitrootlogin|passwordauthentication|maxauthtries|allowusers)'

Output shows the actual values sshd will use at startup, including parameters from Match blocks.

Four Parameters for a Test Stand

ParameterValueRationale
PermitRootLoginnoRoot should not log in directly
AllowUsersdevops adminWhitelist, everyone else rejected
PasswordAuthenticationnoKeys only, passwords disabled
MaxAuthTries3Block after three failures
Warning

Changes apply after systemctl reload sshd. Make edits via ssh -t user@host "sudo nano /etc/ssh/sshd_config" so you do not lose your session on a mistake.

PermitRootLogin

Direct root login is the first thing attackers brute-force. Disable it:

sudo sed -i 's/^PermitRootLogin.*/PermitRootLogin no/' /etc/ssh/sshd_config

If you need root access, log in as a regular user and escalate via sudo. This logs your actions to auth.log.

AllowUsers

A whitelist excludes everyone not explicitly added. If a user is not in the list, sshd returns Permission denied before asking for a password.

echo "AllowUsers devops admin monitoring" | sudo tee -a /etc/ssh/sshd_config
Note

Specify users space-separated. For groups use AllowGroups. Both parameters support patterns: AllowUsers devops@10.0.0.* restricts login by subnet.

PasswordAuthentication

Keys cannot be brute-forced. Switch the setting:

sudo sed -i 's/^PasswordAuthentication.*/PasswordAuthentication no/' /etc/ssh/sshd_config

Before disabling, confirm your public key is in ~/.ssh/authorized_keys on the stand. Otherwise you lock yourself out.

MaxAuthTries

Protection against brute force. After three failed attempts the connection drops:

MaxAuthTries 3

Setting below one disables the limit. Use 3–5 depending on your network’s reliability.

Validating Configuration

Always validate syntax after editing:

sudo sshd -t

Empty output means sshd will start with the new parameters. Any error prints to the screen.

Then apply:

sudo systemctl reload sshd

Common Mistakes

Editing the wrong file. On some distributions sshd_config lives in /etc/ssh/sshd_config.d/. Include files are read in alphabetical order. The default /etc/ssh/sshd_config may be overwritten on package updates — put custom parameters in a .conf file with a meaningful name instead.

Spaces after the parameter. Syntax requires a space between key and value:

# Wrong
PermitRootLogin= no

# Correct
PermitRootLogin no

Comments instead of parameters. The line #PasswordAuthentication no is a comment, sshd ignores it. Remove the # or add a new line.

Match blocks override global settings. If a Match User root block exists at the end of the file, it may restore PermitRootLogin yes for that user. Check sshd -T output.

Additional

This covers a test stand. Production adds ClientAliveInterval 300, ClientAliveCountMax 2 for keepalive and X11Forwarding no if graphics are not needed. But deploying a minimal baseline takes four parameters and a couple of commands.

34 - systemd-run: Run Services Without Unit Files

Sometimes you need to run a process under systemd’s control without writing a unit file — maybe you’re in a container without systemd, on someone else’s machine, or just need a quick one-off. That’s where systemd-run comes in.

Why systemd-run

The tool creates a transient unit — a unit that exists only in systemd’s memory, with no file on disk. This is useful when you need:

  • Resource control (cgroup) over an ad-hoc process;
  • The process to survive terminal closure;
  • Isolation in a separate slice.

In practice, it’s a wrapper around systemctl start for units nobody persists.

Basic Syntax

systemd-run [OPTIONS] COMMAND [ARGUMENTS...]

Simple run:

systemd-run /bin/bash -c "while true; do :; done"

The process lands in user.slice, runs in the background, and is manageable via systemctl. Verify:

systemctl list-units --type=service

You’ll see something like run-u1234.service in the output.

Key Flags

FlagPurpose
--scopeCreate a scope unit instead of service (parent process stays in scope)
--unit=NAMEAssign a custom unit name instead of auto-generated
--uid=USERRun as the specified user
--gid=GROUPRun as the specified group
--nice=NNice value (-20 to 19)
--property=KEY=VALUEPass a unit property (CPUAccounting, MemoryMax, etc.)
--chdir=PATHWorking directory
--setenv=VAR=VALUEEnvironment variable
--tmpfs=PATH:OPTIONSMount a tmpfs at the specified path

Flags combine freely. Example with resource limits:

systemd-run \
  --uid=appuser \
  --gid=appgroup \
  --nice=10 \
  --property=MemoryMax=512M \
  --property=CPUAccounting=true \
  --property=CPUWeight=256 \
  -- /opt/myapp/bin/server

Isolation: CPU, RAM, Root via cgroup

The main value is resource control through cgroup v2.

Memory limit:

systemd-run \
  --property=MemoryMax=1G \
  --property=MemoryHigh=800M \
  /bin/memory-heavy-task

If the process exceeds the limit, systemd kills it (OOM via cgroup).

CPU limit:

systemd-run \
  --property=CPUQuota=50% \
  /bin/calculator

50% of one CPU core. On multi-core systems, calculate percentages against total units.

Isolation with alternate root:

systemd-run \
  --property=RootDirectory=/opt/jail \
  --property=RootImage=overlayfs \
  /bin/sh -c "whoami"

RootDirectory requires a correct directory structure inside. Without it, the process fails to start with “Failed to pivot root”.

Mount a tmpfs:

systemd-run \
  --tmpfs=/tmp/isolated:rw,noexec,size=256M \
  -- /bin/sh

Useful for processing temporary files with a size constraint.

Combined example:

systemd-run \
  --uid=nobody \
  --gid=nogroup \
  --nice=5 \
  --chdir=/var/data \
  --property=MemoryMax=256M \
  --property=CPUQuota=25% \
  --property=IOAccounting=true \
  -- /usr/local/bin/worker --queue=default

All process resources are tracked in cgroup, visible via systemd-cgtop and /sys/fs/cgroup.

Comparison with nohup, setsid, chroot

ToolResourcesIsolationSurvives logoutManagement
nohupNoNoYesNo
setsidNoNoYesNo
chrootNoFilesystem onlyDependsNo
systemd-runcgroupCPU, RAM, I/O, userYessystemctl

nohup and setsid work fine for background scripts. But if you need memory limits or CPU priority — cgroup is non-negotiable.

chroot solves only filesystem isolation. You can’t do systemd-run --chroot, but you can combine them:

systemd-run \
  --uid=nobody \
  --property=MemoryMax=100M \
  --chdir=/tmp \
  /usr/sbin/chroot --userspec=nobody /opt/jail /bin/daemon

Here chroot handles filesystem isolation, systemd-run limits resources.

Pitfalls: scope vs service

By default, systemd-run creates a service unit. The difference:

  • service — full unit, registered with systemd, has dependencies, managed normally.
  • scope — tied to the parent process. Specified with --scope. Parent dies — scope dies.

In day-to-day practice, service works for long-lived processes:

systemd-run --unit=my-daemon /usr/local/bin/daemon

Scope is useful for grouping related processes:

systemd-run --scope --uid=user1 /bin/bash -c 'spawn-worker-1 & spawn-worker-2 & wait'

If you run with --scope and close the terminal — processes die. Without --scope — they survive.

Transient Unit Limitations

Transient units don’t persist across reboots. Obvious, but there are other gotchas:

  • No dependencies. Units lack After=, Wants=, Requires=. The process starts immediately. If you need sequencing — chain commands in a script with sleep.
  • Limited property set. Not all unit file fields are available via --property. For example, Restart=always is not supported in transient units (tested on systemd 254).
  • Harder to debug. No file — no systemctl edit. All parameters are only in the command line.

For long-running services with dependencies and restart policies — a unit file remains the right choice. systemd-run is for one-off tasks, prototyping, and resource limiting on the fly.

35 - curl: HTTP Debugging in CLI

cURL is the standard tool for debugging HTTP in the terminal. It ships out of the box on Linux and macOS and is present in most Docker images. Need to quickly check an API, inspect response headers, or trace a redirect issue — one command line is enough.

Basic Debug Flags

The most common scenario: get a response and see what the server returned. The -i flag prints headers before the body, -v enables verbose mode with connection details.

curl -i https://api.example.com/health
curl -v https://api.example.com/health

The difference: -i shows headers plus body, -v adds DNS resolution, TLS handshake, and debug info before the request.

A HEAD request checks resource availability and metadata without downloading the body:

curl -I https://api.example.com/v2/large-file.zip
Note

HEAD does not guarantee the server supports ranges or caching — this depends on server configuration.

Methods and Request Body

By default, curl sends GET. For other methods, use the -X flag.

curl -X POST https://api.example.com/users
curl -X DELETE https://api.example.com/users/42

Request body is passed via -d. For JSON APIs, a typical pattern:

curl -X POST https://api.example.com/users \
  -H "Content-Type: application/json" \
  -d '{"name": "alice", "role": "admin"}'
Tip

Multiline JSON is easier to read in a heredoc when the body is large:

curl -X POST https://api.example.com/users \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
  {
    "name": "alice",
    "role": "admin"
  }
  EOF

To send form data or data from a file:

curl -X POST https://api.example.com/upload \
  -d "username=admin" \
  -d "password=secret"

Custom Headers

The -H flag adds or overrides a header. You can use -H multiple times.

curl -X GET https://api.example.com/orders \
  -H "Accept: application/json" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Cache-Control: no-cache"

Overriding the Host header is useful when debugging virtual hosts or proxies:

curl -X GET http://10.0.0.5/ \
  -H "Host: example.com"

To remove a default header, use -H "Accept:" — empty value after the colon.

Auth and Certificates

Basic HTTP auth via -u in user:password format:

curl -u admin:secret https://api.example.com/admin

For Bearer tokens, use a header:

curl -H "Authorization: Bearer eyJhbGci..." https://api.example.com/me
Warning

-u sends credentials in plain text if TLS is not used. Always use HTTPS for production servers.

When working with self-signed certificates, the -k flag disables verification:

curl -k https://dev.example.com/api

For known hosts and pinned certificates:

# specify CA bundle
curl --cacert /etc/ssl/certs/ca-certificates.crt https://secure.example.com
# check remote host certificate
curl -v https://secure.example.com 2>&1 | grep "Server certificate"

Timeouts and Saving Response

By default, curl waits indefinitely. For scripts and monitoring, set limits:

FlagPurpose
--max-time Ntotal timeout in seconds
--connect-timeout Nconnection timeout
curl --max-time 5 --connect-timeout 2 https://slow-api.example.com

Saving the response:

# to file with original name
curl -O https://example.com/reports/november.csv
# to specified file
curl -o report.csv https://example.com/reports/november.csv

Redirect stdout to pipe the response body:

curl -s https://api.example.com/health | jq .status

Redirects

Curl does not follow redirects by default. The -L flag enables automatic redirection:

curl -L https://bit.ly/api-status

To debug a redirect chain:

curl -Lv https://short.link/resource 2>&1 | grep -E "< HTTP|< Location"
Note

-L limits redirect depth (default 50). Infinite redirect loops are a common cause of curl hanging.

A flag combination for a full picture when debugging an API:

curl -X POST https://api.example.com/orders \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"item_id": 101, "qty": 2}' \
  -iv --max-time 10 -o response.json

This command shows request and response headers, saves the body to a file, and enforces a 10-second timeout.

36 - ethtool: network interface diagnostics and tuning

Network issues hide well — interface is up, IP is assigned, iptables is quiet, yet packet loss or micro-freezes only surface under load. ethtool gives direct access to hardware state, driver behavior, and offload mechanisms that neither ip nor netstat expose.

Installation is straightforward:

# RHEL/Alma/Rocky
sudo dnf install ethtool -y

# Debian/Ubuntu
sudo apt install ethtool

Running ethtool without flags prints a summary:

$ ethtool eth0
Settings for eth0:
    Supported ports: [ TP ]
    Supported link modes:   1000baseT/Full
    Supported pause frame use: Symmetric
    Supported auto-negotiation: Yes
    Advertised link modes:  1000baseT/Full
    Advertised pause frame use: Symmetric
    Advertised auto-negotiation: Yes
    Speed: 1000Mb/s
    Duplex: Full
    Port: Twisted Pair
    PHYAD: 0
    Transceiver: internal
    Auto-negotiation: on
    MDI-X: Unknown
    Link detected: yes

First thing I check when network complaints come in — Link detected. If no, cable or transceiver is dead. Speed and Duplex tell you whether the link dropped to 100Mb/s or half-duplex.

Driver and Hardware: -i and -a

-i shows driver information:

$ ethtool -i eth0
driver: ixgbe
version: 5.19.0
firmware-version: 0x8000095d
expansion-rom-version: [trimmed]
bus-info: 0000:01:00.0
supports-statistics: yes
supports-test: yes
supports-eeprom-access: yes
supports-register-dump: yes
supports-priv-flags: yes

Firmware version matters for Intel and Broadcom — older versions have known bugs. If you do not see error counters you expect, update firmware first, not the driver.

-a displays auto-negotiation and pause settings:

$ ethtool -a eth0
Pause parameters for eth0:
Autonegotiate:  on
RX:             on
TX:             on
Warning

Disabling flow control on one end of a link without agreement on the other causes packet loss during bursty traffic. If the switch has PAUSE disabled — disable it on the host too.

Speed and Duplex: -s and autoneg

-s changes interface parameters. To set fixed speed and duplex, disable auto-negotiation first, then specify the values:

# Fix 1G full-duplex, disable auto-negotiation
sudo ethtool -s eth0 speed 1000 duplex full autoneg off

# Re-enable auto-negotiation
sudo ethtool -s eth0 autoneg on

After changing parameters, verify the link again — not all cards re-negotiate cleanly without an interface bounce.

ethtool flagPurpose
speed NSpeed in Mb/s (100, 1000, 10000, …)
duplex full|halfDuplex mode
autoneg on|offAuto-negotiation control
port tp|fiber|aui|bnc|miiPort type (not available on all cards)
advertise NBitmask of modes for auto-negotiation

All flags combine in a single call. Persisting across reboots requires writing to /etc/sysconfig/network-scripts/ifcfg-eth0 on RHEL or using a systemd override:

# RHEL-style: /etc/sysconfig/network-scripts/ifcfg-eth0
ETHTOOL_OPTS="speed 1000 duplex full autoneg off"

Offload Flags: -k, -K and Common Pitfalls

-k shows current offload flags, -K modifies them:

$ ethtool -k eth0 | head -20
Features for eth0:
tcp-segment-offload: on
tcp-segment-offload: on [fixed]
generic-segment-offload: on
generic-receive-offload: on
generic-segment-offload: on [fixed]
large-receive-offload: on
rx-vlan-offload: on
tx-vlan-offload: on [fixed]
ntuple-filters: off [fixed]
receive-hashing: off

The [fixed] label means the flag is hardware-enforced and cannot be changed.

Disabling offload flags is a frequent cause of VPN, monitoring, and virtualization issues:

# Disable TSO to force the kernel to send raw segments
sudo ethtool -K eth0 tso off

# Disable VLAN offload if the 802.1Q driver filter is buggy
sudo ethtool -K eth0 rxvlan off txvlan off

# Verify changes
ethtool -k eth0 | grep -E 'tcp-segment|vlan'
Note

GRO (generic-receive-offload) and TSO (tcp-segment-offload) work as a pair. Disabling one without the other causes fragmentation at the kernel level — CPU usage spikes for no good reason.

Statistics: -S and Packet Loss Investigation

-S outputs driver statistics. Format and available counters vary by driver:

$ ethtool -S eth0 | grep -E 'error|drop|miss'
     rx_errors: 0
     tx_errors: 0
     rx_dropped: 0
     tx_dropped: 0
     multicast: 42
     rx_no_buffer_count: 0
     rx_missed_errors: 0

For Intel drivers (ixgbe, i40e), relevant counters include:

$ ethtool -S eth0 | grep -iE 'flow-director|rss|mbus|over'
     rx_fifo_errors: 0
     rx_pause_pfc_ignored: 0
     tx_fifo_errors: 0
Tip

rx_fifo_errors and tx_fifo_errors indicate congestion. If they grow despite normal CPU utilization, the problem lies with memory or the bus.

For scripted problem detection:

#!/bin/bash
IFACE=${1:-eth0}
ethtool -S $IFACE | awk '/error|drop|miss|overflow|fifo|discard/ {if ($2 > 0) print}'

Run periodically via cron — peak loss spikes become impossible to explain retroactively.

Taskip linkethtool
Bring interface up/downip link set eth0 up/downno
MAC addressip link show eth0no
MTUip link set eth0 mtu 9000no
Speed/duplexnoethtool -s eth0 speed 1000 duplex full
Auto-negotiationnoethtool -s eth0 autoneg off
Offload flagspartially via ethtool -kethtool -K eth0 tso off
Error statisticsip -s link show eth0ethtool -S eth0 (more detailed)
Driver informationnoethtool -i eth0
Wake-on-LANnoethtool eth0 (shown in output)

ethtool does not replace ip, it complements it. Network configuration stack: ip link → ip addr → ethtool → tc.

Troubleshooting: Link/Duplex Mismatch

Classic scenario: server and switch failed to agree on parameters. Symptoms — link is up, pings work, but under load there are sudden drops.

Diagnostic sequence:

# 1. Check what the host sees
ethtool eth0 | grep -E 'Speed|Duplex|Auto-negotiation|Link'

# 2. Inspect error counters
ethtool -S eth0 | grep -iE 'error|drop|miss|fifo'

# 3. Compare with the peer
# On the switch (Cisco): show interfaces GigabitEthernet0/1

# 4. Lock parameters on both sides
# On the host:
sudo ethtool -s eth0 speed 1000 duplex full autoneg off

# On the switch (Cisco):
interface Gi0/1
  speed 1000
  duplex full
  no negotiate
Warning

Always negotiate both ends. If the switch runs fixed mode without autoneg and the host has autoneg enabled — the 802.3 standard mandates fallback behavior, but vendors implement it poorly.

If the interface still drops packets after parameter alignment, investigate:

  • Driver and firmware: update the network card firmware
  • Cable or transceiver: a 10G SFP module plugged into a 1G port without auto-negotiation guarantees loss
  • RSS and IRQ issues: cat /proc/interrupts | grep eth0, check IRQ balancing

37 - lsof: which processes listen on port and hold file

Service won’t start — port 8080 is already bound. You dig into who’s holding it, and discover the same process has your config file open while you’re trying to edit it. lsof answers both questions: which processes opened which files and sockets.

Listening Ports

The classic task — find who is listening on a specific port.

lsof -i -n -P
FlagEffect
-iShow internet sockets
-nSkip DNS resolution (show IP instead of hostname)
-PSkip port-to-service conversion (show 80 instead of http)

Without -n -P, lsof wastes time on DNS lookups and resolves ports through /etc/services. On production hosts that’s unnecessary seconds.

For a specific port:

lsof -i :8080 -n -P
Tip

To find what process is listening on port 443, lsof -i :443 -n -P is sufficient. Output shows PID, user, and socket type (IPv4/IPv6, TCP/UDP).

To filter by protocol:

lsof -i TCP:22 -n -P    # TCP only
lsof -i UDP:53 -n -P    # UDP only

Processes in a Directory

Need to see which processes are working with files inside a directory? lsof +D recursively walks the directory and lists all open files.

lsof +D /var/log/
Warning

On directories with thousands of files (for example, /tmp), this command runs slowly. It traverses the filesystem rather than querying the kernel — this is an O(n) operation.

For non-recursive search (files directly in the directory, no subdirectories), use find + xargs:

find /etc/nginx -maxdepth 1 -type f -exec lsof {}

The result shows all processes holding open files from the specified directory. Typical scenario — cannot unmount a partition because someone is working with files inside it.

Single Process Inventory

When the PID is known, list all open files:

lsof -p 1234

Output includes regular files, libraries (.so), sockets, and pipes. To grep by type:

lsof -p 1234 | grep REG      # regular files only
lsof -p 1234 | grep FIFO     # pipes
lsof -p 1234 | grep IPv      # network sockets
Note

REG — Regular file, DIR — directory, FIFO — named pipe, IPv4/IPv6 — network sockets. The TYPE in lsof output matches the type in /proc/PID/fd.

The reverse operation — find PID by file:

lsof /var/log/syslog

If the file is locked (log rotation fails, unmount does not work), this command shows the culprit.

Truncated Command Names

By default, lsof truncates command names to 9 characters. For long names (java, python) this may not be enough:

lsof +c 0 -i -n -P    # show full command name
lsof +c 20 -p 1234    # up to 20 characters
lsof +c 0 -p $(pgrep -f nginx)
Tip

+c 0 means “no limit”. Useful when working with Java processes where the command line contains dozens of characters of classpath.

Quick Reference

CommandPurpose
lsof -i :PORTWho listens on PORT
lsof -i TCPAll TCP connections
lsof -i UDPAll UDP connections
lsof -p PIDFiles opened by PID
lsof +D DIRProcesses in directory
lsof /path/to/filePID that opened file
lsof +c NLimit command name to N characters
lsof -u USERAll open files for user
lsof -c CMDFiles for processes named CMD

Common Issues

lsof not installed — minimal images require installation:

apt install lsof    # Debian/Ubuntu
yum install lsof    # RHEL/CentOS

No read access to /proc — viewing other users’ processes requires root or group membership with access. Usually means running through sudo.

lsof hangs — kernel not responding to file descriptor requests (NFS issues, stalled filesystem). Ctrl+C and restart with a timeout.


lsof is one of those tools you return to every time you troubleshoot locked resources. Three commands cover 90% of the tasks: -i :PORT for ports, +D /path for directories, -p PID for processes.

38 - mc — MinIO Client S3 CLI

S3-compatible object storage is the default choice for buckets, backups, and static assets. When AWS CLI feels excessive and the web console is too clunky, MinIO Client (mc) fills the gap. This CLI tool works with any S3-compatible storage: MinIO, Yandex Cloud, AWS S3, Backblaze B2. Zero dependencies, works out of the box, configured in under a minute.

Installation

Download the binary and make it executable:

curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc \
  -o /usr/local/bin/mc && chmod +x /usr/local/bin/mc

Verify the version:

mc --version

For macOS, use Homebrew:

brew install minio-stable/mc/mc
Note

mc is a single static binary with no dependencies. Works perfectly in containers and minimal base images.

Adding an Alias

An alias is a named connection to an S3 endpoint. Without it, every command requires the full URL.

mc alias set myminio https://minio.example.com \
  AKIAIOSFODNN7EXAMPLE \
  wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY

After this, myminio replaces the URL in all commands. List existing aliases:

mc alias list
Tip

In production, avoid storing keys in command history. Use environment variables: mc alias set prod ${S3_ACCESS_KEY} ${S3_SECRET_KEY} --api API-S3v4 and substitute via env.

For S3-compatible services with self-signed certificates:

mc alias set myminio https://minio.example.com \
  minioadmin minioadmin --api S3v4 --insecure

Browsing and Navigation

List buckets:

mc ls myminio

Recursive listing with contents:

mc ls --recursive myminio/backups/

Object metadata:

mc stat myminio/backups/db-2024-01.sql.gz
Warning

mc ls without --recursive shows only top-level buckets. Use prefixes to navigate folders inside a bucket.

Working with Buckets

Create a bucket:

mc mb myminio/app-logs

If the bucket already exists, mc reports an error. The --ignore-existing flag suppresses it:

mc mb --ignore-existing myminio/app-logs

Remove an empty bucket:

mc rb myminio/app-logs

For non-empty buckets, use force:

mc rb --force myminio/app-logs

Working with Objects

Print object contents to stdout without downloading:

mc cat myminio/config/latest.yaml

Remove an object:

mc rm myminio/backups/old.sql.gz

Recursive removal by pattern:

mc rm --recursive --force myminio/temp/*
Note

mc rm without --force asks for confirmation. Always use --force in scripts.

Find objects by criteria:

mc find myminio --name "*.log" --older-than 30d

Combine with deletion:

mc find myminio --name "*.tmp" --exec "mc rm {path}" {}

Uploading and Downloading

Copy a local file to a bucket:

mc cp /tmp/dump.sql myminio/backups/

Multiple upload:

mc cp ./uploads/* myminio/static/

Download an object locally:

mc cp myminio/backups/latest.tar.gz /tmp/

Copy between buckets (or across aliases):

mc cp myminio/archive/2024/ mys3/backup-2024/ --recursive

Flags for flow control:

FlagPurpose
--recursiveProcess directories recursively
--forceOverwrite without prompting
--preserveKeep file attributes (mtime, ACL)
--if-not-existsSkip already existing objects
--disable-multipartUpload in a single PUT request

Mirroring

One-way directory synchronization:

mc mirror /data/myminio/uploads

Flags for production use:

mc mirror --overwrite --delete \
  /data/myminio/uploads
FlagPurpose
--overwriteReplace changed files
--deleteRemove files in destination missing from source
--watchMonitor for changes in real time
--md5Verify MD5 after upload
Warning

--delete is dangerous: it removes files in the destination that don’t exist in the source. Test with --dry-run or use --preserve for backup scenarios.

Dry-run — show what would happen without making changes:

mc mirror --overwrite --delete --dry-run \
  /data/myminio/uploads

Access Policies

Set public access on a bucket:

mc anonymous set download myminio/public

Common policies:

# Read-only for everyone
mc anonymous set download myminio/public

# Fully public
mc anonymous set public myminio/public

# Private (keys only)
mc anonymous set private myminio/private

# No listing, allow downloads by exact link
mc anonymous set uploadOnly myminio/uploads

View current policy:

mc anonymous list myminio

Generate a presigned URL (works for private buckets too):

mc share download --expire 48h \
  myminio/backups/db-2024-01.sql.gz

Output contains the signed URL and expiration time.

Useful Flags

Global flags working with any command:

FlagPurpose
--debugVerbose HTTP request/response output
--jsonJSON output (easy to parse in scripts)
--no-colorDisable colored output
--insecureSkip TLS certificate verification
--config-dirConfig path (default ~/.mc)
--limitRate limit (e.g., --limit 10MiB/s)

JSON output for automation:

mc ls --json myminio | jq -r '.key'

Progress for large file copies:

mc cp --progress large.iso myminio/backups/

Shell Completion

Bash/zsh autocompletion saves time:

# Bash
mc completion bash > /etc/bash_completion.d/mc

# Zsh
mc completion zsh > "${fpath[1]}/_mc"

# Fish
mc completion fish > ~/.config/fish/completions/mc.fish

After enabling, type mc and press Tab twice to see available commands and aliases.


mc covers 90% of S3 tasks. For complex scenarios (versioning, lifecycle policies, encryption) — use the API or a Terraform provider. But basic bucket and object operations are faster with this tool than with any SDK.

39 - nftables: Modern Linux Firewall

Warning

Before changing nftables, make sure you have physical or console access to the server. A misconfigured input chain can block SSH and lock you out.

nftables replaced iptables in the Linux kernel starting with version 3.13. If you’re still writing rules in iptables style, it’s time to reconsider. nftables performs better, has built-in dual-stack IPv4/IPv6 support, and lets you manage the entire ruleset as a whole instead of entering commands one by one.

Why Migrate from iptables

iptables has fundamental design problems. Each table (filter, nat, mangle) is a separate rule set with its own semantics. There is no built-in support for simultaneous IPv4 and IPv6 handling—you end up writing two separate rule sets. Performance degrades with a large number of rules due to linear lookup.

nftables takes a different approach. All protocols work within a single data structure—the ruleset. The kernel compiles rules into efficient lookup structures. Atomic ruleset replacement eliminates race conditions during rule updates.

Warning

RHEL 8 and newer redirect iptables to nftables by default. Ubuntu 22.04 and later do the same. Check with update-alternatives --display iptables.

Basic Commands: Viewing and Flushing Rules

The first command to memorize:

nft list ruleset

Output shows all tables, chains, and rules. Without tables, output is empty—this is normal.

Create a table named filter for packet handling:

nft add table inet filter

inet means the table handles both protocols. For IPv4-only use ip, for IPv6 use ip6.

Add a chain for incoming traffic:

nft add chain inet filter input { type filter hook input priority 0 \; policy accept \; }

Flags breakdown:

FlagPurpose
type filterChain type—packet filtering
hook inputAttachment point—incoming packets
priority 0Processing order relative to other hooks
policy acceptDefault action—allow everything

Flush rules in a chain:

nft flush chain inet filter input

Delete an entire table:

nft delete table inet filter

Adding and Deleting Rules by Handle

Create several rules and inspect their handles:

nft add rule inet filter input tcp dport 22 accept
nft add rule inet filter input tcp dport 80 accept
nft add rule inet filter input tcp dport 443 accept
nft list ruleset

Output shows something like:

table inet filter {
    chain input {
        type filter hook input priority 0; policy accept;
        tcp dport 22 accept
        tcp dport 80 accept
        tcp dport 443 accept
    }
}

Handles are hidden by default. To work with specific rules:

nft -a list ruleset

The -a flag reveals the handle for each rule. Now you can delete by number:

nft delete rule inet filter input handle 3
Tip

Handles change whenever you add or delete a rule. If your script modifies the ruleset, save nft -a list ruleset output to a file for tracking.

Add a rule with priority before existing ones—at the chain start:

nft insert rule inet filter input tcp dport 2222 accept

The add command appends to the end, insert adds at the start. For insertion at a specific position:

nft add rule inet filter input position 2 tcp dport 8080 accept

Atomic Ruleset Replacement

Adding rules one by one creates a window where some rules are active and others are not yet applied. This is unacceptable for production.

The solution: write the complete ruleset to a file and load it atomically:

nft list ruleset > /etc/nftables.conf

/etc/nftables.conf is the standard location across most distributions. Now edit it:

#!/usr/sbin/nft -f

flush ruleset

table inet filter {
    chain input {
        type filter hook input priority 0; policy drop;
        
        ct state established,related accept
        ct state invalid drop
        
        iif lo accept
        
        tcp dport 22 accept
        tcp dport 80 accept
        tcp dport 443 accept
        
        counter drop
    }
}

Load it:

nft -f /etc/nftables.conf
Note

flush ruleset clears everything before loading. If you need to append to existing rules, remove this line.

Check without applying:

nft -c -f /etc/nftables.conf

The -c flag performs syntax validation without changing state. Useful in CI/CD before deployment.

Enable persistence on systemd distros:

systemctl enable --now nftables

nftables.service loads /etc/nftables.conf on boot. After editing the file, systemctl restart nftables applies it.

Forward Chain

If the host is not a router, leave forward empty with a drop policy. When IP forwarding is on, the minimum is established replies plus transit between interfaces:

nft add chain inet filter forward { type filter hook forward priority 0 \; policy drop \; }
nft add rule inet filter forward ct state established,related accept
nft add rule inet filter forward iifname "eth0" oifname "eth1" accept

For debugging, log before the implicit drop:

nft add rule inet filter forward log prefix "nft-forward-drop: " level warn

Logs go to journalctl -k or /var/log/kern.log.

Common Scenarios: Blocking Ports and IPs

Block incoming connection from a specific IP:

nft add rule inet filter input ip saddr 1.2.3.4 drop

Block outgoing to a specific IP:

nft add rule inet filter output ip daddr 5.6.7.8 drop

Block an IP range (CIDR):

nft add rule inet filter input ip saddr 10.0.0.0/8 drop

Block a port for everyone:

nft add rule inet filter input tcp dport 25 drop

Allow a port only for a specific subnet:

nft add rule inet filter input ip saddr 192.168.1.0/24 tcp dport 5432 accept

Log dropped packets:

nft add rule inet filter input counter drop

Counters appear in nft list ruleset output—they show packet and byte counts.

NAT via Masquerading

For sharing a single IP with a local network:

nft add table ip nat
nft add chain ip nat postrouting { type nat hook postrouting priority 100 \; }
nft add rule ip nat postrouting ip saddr 192.168.0.0/24 masquerade

Masquerading automatically substitutes the external IP of the interface. For port forwarding NAT:

nft add chain ip nat prerouting { type nat hook prerouting priority -100 \; }
nft add rule ip nat prerouting tcp dport 8080 dnat to 192.168.0.100:80
Warning

NAT in nftables works only for IPv4. For IPv6, use stateless NAT66 or routing-level addressing.

iptables-nft Compatibility Mode

Some distributions keep iptables “working” by translating to nftables:

update-alternatives --set iptables /usr/sbin/iptables-nft
update-alternatives --set ip6tables /usr/sbin/ip6tables-nft

The problem is these are two different worlds. iptables-nft translates commands to nftables, but reverse compatibility does not work. Rules created via iptables will not appear directly in nft list ruleset.

Warning

Do not use iptables and nftables simultaneously. The result is unpredictable. Either fully migrate to nftables or stay on iptables. Check current mode: iptables -V shows whether iptables-legacy or iptables-nft is in use.

For migration from iptables, use the iptables-translate utility:

iptables-translate -A INPUT -p tcp --dport 22 -j ACCEPT

Output: nft add rule ip filter input tcp dport 22 accept. Manual verification of output is mandatory—automatic translation is not perfect.

nftables is not the future—it is the present. If you administer Linux servers, spend an evening on migration. The file /etc/nftables.conf with the complete ruleset is your backup and deployment in one.

40 - ss: socket statistics instead of deprecated netstat

When netstat hangs on a server with tens of thousands of connections, it’s time to switch to ss. Part of the iproute2 package, ss queries the kernel directly via netlink instead of parsing /proc/net/*. The result is instant output with minimal overhead.

Why switch from netstat

netstat from net-tools relies on a deprecated approach: it reads from /proc/net/tcp, /proc/net/unix and converts numeric IDs to symbolic names. On a server with active connections, this takes seconds and spikes CPU usage.

ss communicates with the kernel through a netlink socket. One call, structured data returned. For 10,000 connections, the difference is 0.02 seconds versus 3–5 seconds.

Note

netstat is officially marked as deprecated in most distributions. iproute2 is the current standard for network management in Linux.

Separate installation is rarely needed: ss ships with iproute2, which is present in every Linux by default.

Basic flags: the -tulnp equivalent

No need to relearn everything — flags are similar, just with more flexible ordering:

# Listening ports, show processes
ss -tlnp
FlagWhat it shows
-tTCP sockets
-uUDP sockets
-lListening sockets only
-nNumeric addresses and ports (no DNS)
-pProcess owner (PID, name)
-aAll sockets (not just listening)
-eExtended info (uid, inode)
-oTimer information

Flag order doesn’t matter — -tlnp and -ltnp are the same. -p needs root to show processes owned by other users.

Output differs structurally from netstat:

State      Recv-Q   Send-Q   Local Address:Port   Peer Address:Port   Process
LISTEN     0        128      0.0.0.0:22           0.0.0.0:*          users:(("sshd",pid=1234,fd=3))
LISTEN     0        511      127.0.0.1:6379       0.0.0.0:*          users:(("redis-server",pid=5678,fd=6))

Key columns:

ColumnMeaning
Recv-QBytes in receive buffer, not yet read by application
Send-QBytes in send buffer, not acknowledged by peer
Local Address:PortLocal end of the connection
Peer Address:PortRemote end
Tip

Non-zero Recv-Q or Send-Q on an established connection signals a problem. The application isn’t keeping up with reads, or the network is congested.

Full set of basic filters:

ss -t       # TCP only
ss -u       # UDP only
ss -w       # raw sockets
ss -x       # Unix sockets
ss -a       # all (listening and established)
ss -l       # listening only

Filtering by state and port

This is where ss beats netstat. Filters are native, not piped through grep:

# Established connections only
ss -t state established

# Active connections (not listening)
ss -t state connected

# TIME_WAIT — the classic use case
ss -t state time-wait

# Everything except listening
ss -t state connected -s

State combinations:

# ESTABLISHED only, with process info
ss -t state established -p

# SYN-SENT, SYN-RECV — handshake issues
ss -t state syn-sent

# FIN-WAIT-1, FIN-WAIT-2
ss -t state fin-wait1,fin-wait2

# CLOSE-WAIT — connection hanging, waiting to close
ss -t state close-wait

# Grouped states
ss -tan 'state established or state time-wait'

Port filter — one of the most common:

# Who is listening on 443
ss -tlnp 'sport = :443'

# Who is connected to 5432 (PostgreSQL)
ss -tp 'dport = :5432'

# All connections to any port 80 or 443
ss -t 'sport = :80 or dport = :80 or sport = :443 or dport = :443'
Warning

The sport and dport filters work with numeric values. For ranges, use >= and <=: 'dport >= 3000 and dport <= 4000'.

Address filter:

# Connections from a specific IP
ss -tp 'src 192.168.1.100'

# Outbound connections (not from local network)
ss -tp 'not src 192.168.0.0/16'

Extended output: -e, -i, -s

For queue diagnostics and statistics:

# Detailed information (extended)
ss -teln

# Adds:
# - uid (user)
# - inode
# - timers (for keepalive, TIME_WAIT)
# - timeout

Output with -e for an established connection:

ESTAB 0 0 10.0.0.5:22 10.0.0.100:52431 users:(("sshd",pid=1820,fd=3)) uid=1000 ino=35234 sk=0xffff88003a2c8000 <->

Interface information (-i):

ss -ti 'dst 10.0.0.1'
ESTAB 0 0 10.0.0.5:22 10.0.0.100:52431
         ts sack hbrs pmtu cwnd rtt rttvar unacked
         wscale:7,7 pmtu:1500 rcvmss:1448 advmss:1448 cwnd:10
         send 0.4Mbps rcv_space:43690

Key metrics:

MetricDescription
cwndCongestion window
rttRound-trip time
pmtuPath MTU
rcv_spaceReceive buffer size
sendCurrent send rate

State statistics (-s):

ss -s
Total: 124 (kernel 128)
TCP:   45 (estab 38, closed 2, orphaned 0, synrecv 0, timewait 2)

Transport Total     IP          IPv6
*         128       -           -
RAW       0         0           0
UDP       12        8           4
TCP       43        38          5
INET      55        46          9
FRAG      0         0           0
Note

The timewait 2 in ss -s output is a quick way to assess connection buildup. If the number grows each time you run it — something isn’t closing connections properly.

Common use cases

Which process is listening on a port and which interface:

ss -tlnp 'sport = :3306'
State   Recv-Q  Send-Q  Local Address:Port  Peer Address:Port  Process
LISTEN  0       128     0.0.0.0:3306        0.0.0.0:*         users:(("mysqld",pid=2345,fd=18))
LISTEN  0       128     127.0.0.1:3306      0.0.0.0:*         users:(("mysqld",pid=2345,fd=17))

Listening twice — once on all interfaces, once on localhost. If you need external only — check bind-address in the config.

How many sockets in TIME_WAIT — and who they belong to:

ss -s | grep timewait
# or
ss -ant | awk '/TIME-WAIT/ {count++} END {print count}'

# Top destinations leaking TIME-WAIT
ss -tan state time-wait | awk '{print $5}' | sort | uniq -c | sort -rn | head -20
Tip

High TIME_WAIT is usually normal. If it causes issues — on the client side use setsockopt with SO_LINGER, or SO_REUSEADDR on the server. Flag -ttu shows timers.

Whether a connection is stalled:

ss -ti 'dst 10.0.0.50'

Check rtt and unacked. If unacked grows but rtt stays the same — packets aren’t arriving but aren’t being lost either. Most likely the remote side stopped reading from the socket.

Find the process that opened a connection to a specific host:

ss -tp 'dst 192.168.1.50'
State   Recv-Q  Send-Q  Local Address:Port  Peer Address:Port  Process
ESTAB   0       0       10.0.0.5:45678      192.168.1.50:443   users:(("curl",pid=9876,fd=3))

Aggregated statistics by process:

ss -tnp | awk 'NR>1 {print $6}' | sort | uniq -c | sort -rn | head -10

Shows how many connections each process holds. Useful for finding processes opening too many sockets.

Quick reference for everyday tasks:

# netstat -tulnp  →  ss -tlnp
# netstat -ulnp   →  ss -ulnp

# What is listening
ss -tlnp

# Active connections
ss -tnp

# Connections to port
ss -t 'dport = :80'

# TIME_WAIT count
ss -s | grep timewait

# Process on port
ss -tlnp 'sport = :8080'

# Keep-alive / TIME_WAIT timers
ss -tno

41 - SSH Config: Wildcards and Dynamic Variable Substitution

SSH reads ~/.ssh/config line by line, but without variables the file quickly becomes copy-paste hell. Here’s how Host patterns, Match exec, and substitution tokens like %h, %r, %l cut config size by orders of magnitude while covering real scenarios — from dynamic routing to agent forwarding through bastion hosts.

Templates and wildcards in SSH config

SSH supports glob-like patterns in the Host directive. The most common is Host *, but compound patterns work too.

# All hosts without explicit configuration get these defaults
Host *
    ServerAliveInterval 60
    ServerAliveCountMax 3
    IdentitiesOnly yes

# All hosts in the example.com domain
Host *.example.com
    User ubuntu
    Port 22

# Hosts matching a pattern: web-01, web-02, web-03
Host web-0?
    IdentityFile ~/.ssh/id_ed25519_web
Note

SSH reads config top-down and applies the first matching Host directive. Put more specific rules above general ones.

# Correct order: specific before general
Host bastion
    HostName bastion.example.com
    User admin

Host *.internal
    ProxyJump bastion

Host *
    ServerAliveInterval 60

Patterns can be combined in a single Host line via spaces — this works as logical OR.

Host dev-* staging-*
    User deploy
    IdentityFile ~/.ssh/id_ed25519_deploy

Substitution variables: %h, %r, %l

SSH substitutes variables in directive values during config processing. Three core ones:

VariableValueExample
%hHostname from command linessh web-01 → web-01
%rRemote usernameubuntu
%lLocal usernamealex

Substitution works in most directives — ProxyCommand, IdentityFile, LocalCommand, RemoteCommand.

Host jump-*
    HostName %h.internal
    ProxyJump bastion.example.com
    User deploy

Host backup-*
    HostName %h.backup.local
    User backup
    # Login matches local username
    IdentityFile ~/.ssh/id_ed25519_%r

Variable %h substitutes what you passed after ssh. This lets you write one rule for hundreds of hosts.

# Instead of five separate blocks
Host server-{01..05}
    HostName %h.internal.corp
    User ansible
    ProxyJump bastion
Tip

%h substitutes what you typed, not the resolved HostName. ssh web-01 substitutes web-01, not its IP.

Match exec and dynamic routing

The Match directive lets you apply rules based on conditions. Without exec, it checks user, host, localuser. With Match exec, you get arbitrary logic via shell command.

# Route through bastion only for internal hosts
Match host 10.* exec "echo true"
    ProxyJump bastion.example.com

The exec condition runs on your local machine. The command must return 0 (success) for the Match block to apply.

# Forward agent only when connecting to prod
Match exec "[ '%h' = prod-* ]"
    AddKeysToAgent yes
    IdentityAgent SSH_AUTH_SOCK
# Different routing depending on network
Match exec "hostname -I | grep -q 192.168.1"
    ProxyCommand none  # local network, direct access

Match exec "[ '%h' != prod-* ]"
    ProxyJump office-jump
Warning

Match exec runs through the shell. Backticks and variables expand. For complex logic, extract the check into a separate script.

Match exec "/home/alex/.ssh/route-check.sh %h"
    ProxyJump bastion
# route-check.sh
#!/bin/bash
HOST="$1"
if [[ "$HOST" == prod-* ]]; then
    exit 1  # prod needs different logic
fi
exit 0

Escaping the percent sign

When you need to pass a literal % character — in RemoteCommand, ProxyCommand, or LocalCommand — use %%.

# SSH into Docker container with eval
Host docker-*
    HostName %h
    User docker
    RemoteCommand docker exec -it %h /bin/sh -c "eval $(ssh-agent -s) && exec /bin/sh"
# Add timestamp in remote command
Host *
    RemoteCommand echo "%%Y-%%m-%%d %%H:%%M:%%S connected to %h" >> /tmp/ssh-log
# Port forwarding via nc on remote host
Host relay
    HostName relay.example.com
    ProxyCommand ssh -W %h:%p user@bastion
    # or classic nc
    ProxyCommand nc -Z %h 22

Examples for dev, staging, prod

Typical setup: one file, three environments, shared base.

# ============================================
# Base settings for all hosts
# ============================================
Host *
    ServerAliveInterval 60
    ServerAliveCountMax 3
    IdentitiesOnly yes
    AddKeysToAgent yes

# ============================================
# Environments: dynamic HostName
# ============================================
Host dev-*
    HostName %h.dev.internal
    User developer
    IdentityFile ~/.ssh/id_ed25519

Host staging-*
    HostName %h.staging.internal
    User deploy
    IdentityFile ~/.ssh/id_ed25519_deploy
    # Additional jump through staging-bastion
    ProxyJump staging-bastion

Host prod-*
    HostName %h.prod.internal
    User deploy
    IdentityFile ~/.ssh/id_ed25519_prod
    ProxyJump prod-bastion
    # Agent only, no key file on disk
    IdentityAgent SSH_AUTH_SOCK

# ============================================
# Match: dynamic logic
# ============================================
# Forward agent only on prod and staging
Match exec "[ '%h' = prod-* ] || [ '%h' = staging-* ]"
    ForwardAgent yes

# Disable ForwardAgent for dev (paranoid setting)
Match host dev-* exec "true"
    ForwardAgent no

# ============================================
# Substitution variables in commands
# ============================================
Host db-*
    HostName %h.internal
    User dbadmin
    RemoteCommand psql -U postgres -c "SELECT now(), '%r'"

# ============================================
# Escaping %% in arguments
# ============================================
Host log-*
    HostName %h
    RemoteCommand echo "Connected at %%Y-%%m-%%d" >> /var/log/ssh-connections.log

After this config, ssh prod-web-03 connects to prod-web-03.prod.internal via prod-bastion with key id_ed25519_prod, while ssh dev-app-01 connects directly.

# Verify SSH sees your config correctly
ssh -G dev-app-01 | grep -E "^hostname|^user|^proxyjump"

# Test connection without executing remote command
ssh -v prod-db-01 echo "connected"
Note

-G outputs all parameters SSH will use after parsing the config for the given host. Use it for debugging before making real connections.

Combining Host patterns, variable substitution, and Match exec transforms SSH config from a collection of copy-paste blocks into a dynamic routing system. One file covers dev/staging/prod without duplication, and %h, %r, %l eliminate manual edits when adding new hosts.

42 - strace: System Call Tracing for Diagnosing Hangs and Leaks

When a service hangs, standard tools like top, htop, and ps show the state but not the cause. If a process is in state D (uninterruptible sleep), it’s waiting on a syscall. strace attaches to a live process and outputs every system call in real time. This turns a mysterious hang into a specific syscall, its arguments, and return code.

Note

strace uses ptrace, the kernel’s debugging mechanism. On production, tracing slows a process by 2–10x. Use it sparingly, targeting a single PID.

Basic Flags and Syntax

Installation:

# Debian/Ubuntu
apt install strace

# RHEL/CentOS
yum install strace

# Alpine
apk add strace

Running:

# Trace a new process
strace ls -la /tmp

# Attach to a running process
strace -p 12345

# Attach including subprocesses (fork/clone)
strace -fp 12345

Core flags for diagnostics:

FlagPurpose
-fFollow child processes
-cCount calls and time (summary)
-ttMicrosecond timestamps
-TTime spent in each syscall
-e trace=openat,read,writeTrace only specified calls
-e write=1,2Trace writes to fd 1 and 2 only
-o output.logWrite output to file
-s 1024Truncate strings longer than N characters

Diagnosing Blocking Calls

Scenario: process is hung, ps shows state D. Need to find what it’s blocked on.

# Check process status
ps aux | grep nginx
# root     12345  0.0  0.1 ... S    pts/0    0:00 nginx: worker

# Attach for 5 seconds, output everything
strace -p 12345 -f -tt -T 2>&1 | head -50

Typical output when blocked on a file:

14:23:45.123456 read(15, "", 1024)  = 0 <2.345678>
14:23:47.469134 openat(AT_FDCWD, "/var/data/large-file.db", O_RDONLY) = 15 <0.000023>
14:23:47.469157 read(15, "", 1024)  = 0 <5.678901>

If you see read(...) <time> = 0 with large time values, the process is waiting for data. If <time> is in seconds, you’ve found the bottleneck.

Tip

Blocking on epoll_wait, poll, select is normal for an idle process. Look for read, write, openat, sendto with times exceeding 100ms.

For network sockets:

# Trace network calls only
strace -p 12345 -e trace=network,read,write -f

Hang on connect() to an unreachable host:

14:30:01.234 connect(14, {sa_family=AF_INET, sin_port=htons(5432), sin_addr=inet_addr("10.0.0.100")}, 16) = -1 EINPROGRESS <3.456789>

EINPROGRESS means non-blocking socket, but long time indicates a network issue or timeout.

Finding File Descriptor Leaks

Scenario: process won’t open files, error “too many open files”. Need to find who’s holding the descriptors.

# Check limits and current usage
ps aux | grep myservice
# 12345  5123  0.0  /opt/myservice

cat /proc/12345/limits | grep "Max open files"
# Max open files            1024                 1024                 files

# Count open fd
ls /proc/12345/fd | wc -l
# 987

Attach strace with a filter on file operations:

strace -p 12345 -e trace=openat,open,close,clone -f 2>&1 | tee /tmp/strace.log

Analyze the output:

# What was opened but not closed
grep -E "openat|close" /tmp/strace.log | awk '{print $2}' | sort | uniq -c | sort -rn | head -20

Alternative — summary mode:

# Stats over 30 seconds
timeout 30 strace -p 12345 -c -f 2>&1
% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- -------
 45.23    1.234567          123       1000         openat
 30.12    0.823456           45      18234       close
 20.45    0.559123         8905        63         read
------ ----------- ----------- --------- --------- -------
100.00    2.617146               19297        63 total

If close is called less than openat — you found a leak.

Warning

With high call frequency (thousands per second), strace generates enormous output. Limit time with -tt and filter by syscall via -e trace=.

Analyzing Slow Requests

Scenario: API endpoint responds in 5 seconds instead of 200ms. Need to find where time is lost.

# Start process with tracing and timings
strace -f -tt -T -o /tmp/slow.log ./myservice

# Or attach to running process
strace -p 12345 -f -tt -T 2>&1 | tee /tmp/live.log

Find calls with high execution time:

# Highlight syscalls taking more than 500ms
grep -E "\> [0-9]\.[0-9]{3}" /tmp/live.log | sort -t '>' -k2 -rn | head -20
14:45:23.123456 write(7, "HTTP/1.1 200 OK\r\n"..., 512) = 512 <0.000045>
14:45:23.890123 openat(AT_FDCWD, "/opt/app/cache.json", O_RDONLY) = 12 <0.523456>
14:45:24.413679 read(12, "{\"key\":\"value\"}"..., 4096) = 4096 <0.000089>

openat with 523ms — investigate: file on NFS, missing permissions, remote filesystem.

For SQL-like queries (PostgreSQL, MySQL), trace the socket:

# Find socket fd of process
ls -la /proc/12345/fd | grep socket
# lr-x 14 -> socket:[1234567]

# Trace specific fd
strace -p 12345 -e write=14 -f -tt -T

Slow database query looks like a series of write/read with long time between them:

14:50:01.123 write(14, "SELECT * FROM orders"..., 45) = 45 <0.000234>
14:50:05.890 read(14, "", 4096)          = 2048 <4.766890>

4.7 seconds between sending the query and receiving data — problem is on the database side or network path to it.

Summary

strace turns a hang with no visible cause into a specific syscall. Attach to PID, filter calls via -e trace=, check execution time with -T. For leaks — compare open/close counts in summary mode. For slow requests — find calls with time exceeding 100ms.

Tip

On production use -e trace=write,read,openat instead of tracing all calls — reduces overhead by 3–5x.

43 - tcpdump and tshark: Packet Capture in CLI

When debugging network issues in Linux infrastructure, ping and curl are not enough. Sometimes you need to see what is actually traveling over the wire. tcpdump is the standard tool for capturing packets from the CLI. tshark is its sibling from the Wireshark suite, convenient for scripting.

Quick Start with tcpdump

Check that packets are reaching the host:

tcpdump -i eth0 host 10.0.0.5

The utility puts the interface into promiscuous mode and prints one line per packet passing through. By default it works with the first interface it finds, but specifying explicitly is better.

For a quick test without DNS resolution (to avoid timeout when there is no network):

tcpdump -i eth0 -nn host 10.0.0.5

-nn prevents resolving both hostnames and ports.

Key Flags

FlagPurpose
-i ifaceInterface
-c NCapture N packets and exit
-nDo not resolve hostnames
-nnDo not resolve hostnames and ports
-v, -vv, -vvvIncrease output verbosity
-w fileWrite raw dump to file (pcap)
-r fileRead dump from file
-XShow hex + ASCII packet body
-s NTruncate each packet to N bytes (0 = full)
-CRotate output file when it reaches N MB

The -v flags are useful for debugging: the first level shows TTL and ID, the second shows flags and window size, the third adds ACK and displays the full IP header.

# Capture 100 packets on port 443, verbosity -vv
tcpdump -i eth0 -nn -c 100 -vv port 443

BPF Filters

tcpdump uses Berkeley Packet Filter. The syntax reads left to right.

# TCP only on port 80 or 443
tcpdump -i eth0 -nn tcp port 80 or port 443

# Host as source OR destination
tcpdump -i eth0 -nn host 192.168.1.10

# Do not show SSH (cut the noise)
tcpdump -i eth0 -nn not port 22

# TCP packets with SYN flag (connection start)
tcpdump -i eth0 -nn 'tcp[tcpflags] == tcp-syn'

# Packets larger than 1000 bytes
tcpdump -i eth0 -nn 'ip[2:2] > 1000'

# ICMP ping (type 8 code 0)
tcpdump -i eth0 -nn 'icmp[icmptype] == 8'

Filters combine with and, or, not. Quotes are needed when the expression contains spaces or special characters.

Warning

Do not run tcpdump -i any in production without restricting by host or port. You will get a flood of traffic and waste disk space for no reason.

Writing to File and Reading

Capturing to a file is mandatory practice. A live sniffer dumps data dozens of lines per second; analyzing in the terminal is impossible.

# Write 10,000 packets to a file
tcpdump -i eth0 -nn -w /tmp/capture.pcap -c 10000

# Read from file (interactive)
tcpdump -r /tmp/capture.pcap

# Read with a filter
tcpdump -r /tmp/capture.pcap -nn 'tcp port 443'

The pcap format is binary. The -C option limits file size:

tcpdump -i eth0 -nn -w /tmp/capture -C 10 -W 5

This creates files /tmp/capture-0, /tmp/capture-1 … up to 5 files, each up to 10 MB. -W sets the number of files.

tshark as a tty-free Alternative

tshark is the command-line part of Wireshark. It produces structured output convenient for parsing in scripts:

# Installation
apt install tshark   # Debian/Ubuntu
yum install wireshark-cli   # RHEL/CentOS

# Capture with field output
tshark -i eth0 -f 'host 10.0.0.5' -c 100 -T fields -e ip.src -e ip.dst -e tcp.port

-T fields -e extracts specific fields from each packet:

# HTTP requests: method and host
tshark -i eth0 'tcp port 80' -Y http.request -T fields -e http.host -e http.request.method -e http.request.uri

Reading pcap files with tshark is more convenient than with tcpdump:

# Show protocol hierarchy statistics
tshark -r /tmp/capture.pcap -z io,phs -q

tshark does not support -C rotation like tcpdump, but it handles live filters better (full Wireshark dissector engine).

Reading Dumps in Wireshark

A pcap created by tcpdump opens in Wireshark without conversion:

# Transfer file to local machine
scp user@server:/tmp/capture.pcap /tmp/

# Or start a web server on the remote host
python3 -m http.server 8080 --directory /tmp

In Wireshark Apply as display filter you enter the same BPF syntax: tcp.port == 443 && ip.src == 10.0.0.1.

Tip

If the file is large (>100 MB), do not open it entirely. Use tcpdump -r with a pre-filter to extract the needed slice: tcpdump -r big.pcap -nn 'host 10.0.0.5' -w small.pcap.

Common Mistakes

  • Forgot -c — the process hangs, capturing traffic indefinitely. Add the limit immediately.
  • No permissions — you need root or sudo tcpdump. In modern distributions you can grant capabilities: setcap cap_net_raw,cap_net_admin=eip /usr/sbin/tcpdump.
  • -w does not work with -l (line-buffering) simultaneously. If you need progress — write to file and read in parallel via tail -f.
  • File was written but reads empty — possibly there was no traffic matching the filter. Check tcpdump -i eth0 -nn without a filter.

44 - CasaOS: Web Dashboard for Home Lab

CasaOS is a lightweight web dashboard for managing Docker containers on a single host. If you’re currently accessing each container through its own port, this replaces that mess with a unified interface where you can install apps with a couple of clicks, monitor resources, and manage storage — without touching Nginx or dealing with Portainer’s complexity.

Installation

CasaOS installs on a clean system with one command. Supported distros: Debian 10+, Ubuntu 18.04+, Raspbian. If Docker is already running, remove it first or accept that CasaOS will take over the Docker daemon.

curl -fsSL https://get.casaos.io | bash

The script outputs the dashboard address when done. Default is http://<server-IP>:80. Web UI is ready immediately, logged in as root.

Note

The installer deploys its own Docker stack. If you already have dockerd running, CasaOS will overwrite network configs. Make sure you have snapshots or a backup ready.

Check the service status after installation:

systemctl status casaos

Interface and Features

The main screen shows installed apps as tiles. Each card displays status (running / stopped), an icon, and quick actions: start, stop, restart, delete.

Left sidebar: file manager, container list, app store, and settings. Out of the box CasaOS provides:

  • volume creation and USB drive mounting;
  • port forwarding through the web interface;
  • CPU, RAM, disk, and network usage metrics;
  • compose file editing directly in the browser.

For a home lab, this covers about 80% of needs. The remaining 20% is manual compose files.

Installing Apps from the Catalog

The CasaOS catalog has several dozen pre-built templates. Find what you need, click Install — CasaOS generates the compose file and starts the container. Under the hood it runs AppStore, essentially a docker-compose wrapper.

Typical flow:

  1. Open App Store in the left sidebar.
  2. Pick an app, for example Plex or AdGuard Home.
  3. Set the data path and port mappings if needed.
  4. Click Install.

For AdGuard Home, CasaOS runs this under the hood:

docker run -d \
  --name adguardhome \
  --restart unless-stopped \
  -v /opt/adguardhome/work:/opt/adguardhome/work \
  -v /opt/adguardhome/conf:/opt/adguardhome/conf \
  -p 53:53/tcp \
  -p 53:53/udp \
  -p 3000:3000/tcp \
  --network host \
  adguard/adguardhome

All parameters are set through the web form — no manual editing required.

Running Containers Manually

When an app isn’t in the catalog, CasaOS lets you upload your own compose file.

# example: Cloudflare Tunnel for external access
docker run -d \
  --name cloudflared \
  --restart unless-stopped \
  -e TUNNEL_TOKEN=your_token \
  docker.io/cloudflare/cloudflared:latest \
  tunnel run

In CasaOS: App Store → Upload — drop in the YAML. It parses the file, lets you adjust variables, and launches. The container appears on the main screen.

Tip

Use /DATA/AppData/ for all volumes. This makes backups straightforward — all state lives in one directory.

Storage Layout

CasaOS creates three key directories:

PathPurpose
/DATA/AppData/Application configs
/DATA/Storage/User files
/var/lib/casaos/Panel system data

When you attach a new disk or USB drive, it shows up in File Manager and mounts automatically. You can set any directory as the default storage location for new apps: Settings → Storage → Default Location.

Updates and Maintenance

Updating the panel itself:

wget https://github.com/IceWhaleTech/CasaOS/releases/latest/download/casaos-update.sh
bash casaos-update.sh

Container images update through the web UI: app card → menu (⋮) → Update. CasaOS pulls the fresh image and recreates the container, leaving data untouched.

Service logs are accessible via terminal:

journalctl -u casaos -f
docker logs casaos -f

To prune unused images and networks:

docker system prune -af

When to Choose Something Else

CasaOS works well for a single host. If you’re running a multi-machine cluster, it won’t help — it’s not a replacement for Kubernetes or Docker Swarm.

Portainer gives more control over networks, images, and stacks. If you prefer detailed configuration and write compose files by hand, CasaOS will feel limiting.

Yacht is even more minimal, with no enforced directory structure. Good if you only want a web UI for containers without a file manager or app catalog.

Warning

CasaOS is actively developed. Minor releases break template compatibility more often than ideal. Keep your working compose files in a Git repository so you can rebuild stacks from scratch if needed.

For a home server running Pi-hole, Home Assistant, Jellyfin, and a few utilities — CasaOS handles the job without overhead. Install it, add your apps, forget about it.

45 - cockpit-ufw-module: Uncomplicated Firewall in Cockpit

UFW on a home or small server is usually configured over SSH: ufw status numbered, then ufw allow 443/tcp. cockpit-ufw-module covers the same cycle in the browser: package, status, policies, rules. It is one panel from the cockpit-modules group — UFW under Tools.

The UI is Russian, in PatternFly v5. You need Cockpit 264+ and administrator rights for writes. MIT license; current release is tag v.1.0.1.

Why a panel if ufw already exists

The client is already there: system ufw. What is missing is a view without a terminal and safe input: port, CIDR, rule number.

The panel does not replace iptables with its own engine. HTML and JS call host commands through cockpit.spawn as argv arrays, not shell strings. Input is validated on the client: port (including a range and a list), IPv4/IPv6 and CIDR, action allow / deny / reject / limit. Without Cockpit admin rights, package and rule operations do not run.

What the page shows

Four blocks, top to bottom:

BlockJob
ufw packageinstalled or not, version, manager (APT, DNF, YUM, Pacman), install and remove
Statusactive / inactive, logging level, incoming/outgoing policies, enable / disable / reload
Rulestable from ufw status numbered: number, to, action, from, delete
Add ruleaction, protocol, port, source IP/CIDR

While the firewall is off, the rules table is hidden: UFW does not apply them, even if the config still has entries. The counter then reads “N rules (inactive)”.

Installing the module

Use the store if it is already on the host. Otherwise install by hand:

git clone https://gitlab.com/cockpit-modules/cockpit-ufw-module.git
cd cockpit-ufw-module
sudo ./install.sh

Files go to /usr/share/cockpit/ufw. Refresh Cockpit — UFW appears under Tools.

For the current user without root:

./install.sh --user

The module lands in ~/.local/share/cockpit/ufw.

Note

install.sh installs only the panel. The ufw package is installed from the UI, with Install ufw.

Typical flow on a clean host

Order matters more than the buttons. After the package install the module runs ufw --force reset, sets allow for incoming and outgoing, opens 22/tcp, and enables the ufw systemd unit. It does not enable the firewall — that is a separate button with an SSH warning.

  1. Tools → UFW → Install ufw. Confirm. On APT this starts with apt-get update.
  2. Add explicit rules for what you manage from this host. Cockpit listens on 9090/tcp — after the package install that rule is missing, only SSH 22 is present. If you later set incoming to deny without 9090, the panel disappears.
  3. Services you need: 80/tcp, 443/tcp, a range such as 8000:8010. The source can be a CIDR, for example 10.0.0.0/8.
  4. Before adding, the UI shows a command preview (ufw allow 443/tcp) — the same argv that will run on the host.
  5. Enable. While default incoming is allow, access is not cut: only explicit deny/reject/limit rules take effect.
  6. Once SSH, Cockpit, and sites still work — switch the incoming policy to deny. Leave outgoing on allow in the usual case.

Examples the form accepts:

22/tcp
443,80
8000:8010
from 10.0.0.0/8 to any port 443 proto tcp

limit is UFW’s rate limit on new connections, useful on 22. reject replies with ICMP/TCP reset; deny drops silently.

Warning

Install ufw runs ufw --force reset. Existing rules on the host are wiped. On a host that already has a tuned firewall, install the package by hand and use the panel only for status and small edits.

What the panel does not do

This is not the full ufw from the man page. The UI has no:

  • application profiles (ufw allow OpenSSH, ufw app list);
  • interface binding (on eth0);
  • rule comments;
  • in/out direction on the add form — new rules are inbound;
  • logging level changes (the Logging line is read-only);
  • routed policy, even though ufw status verbose is parsed.

Removing the package disables UFW (ufw --force disable) and uninstalls it through the system manager; APT uses --purge.

Local check without touching a live host — Docker Compose from the repository: a privileged container with systemd, Ubuntu, Cockpit, and ufw, default http://localhost:9090, login admin / admin. That is a demo, not a production pattern.

Sources, issues, and tags live at gitlab.com/cockpit-modules/cockpit-ufw-module. Neighbouring panels in the same group: fail2ban, cron, CertManager.

46 - Creating a Custom Systemd Service

Your application needs to start on boot, restart on crash, and log output. Shell scripts in /etc/rc.local give you none of that. Systemd solves all three with a single declarative file.

Why Write a Custom Unit File

Supervisord and init scripts are overkill for most cases. Systemd provides a unified interface for service management: socket-based activation, dependency tracking, resource limits, and built-in logging via journald. You get all of it without additional tooling.

Unit File Structure

A unit file is an ini-style text file placed in /etc/systemd/system/ for persistent configuration or /run/systemd/system/ for runtime-only changes. Naming convention is name.service.

[Unit]
Description=My Application
After=network.target

[Service]
Type=simple
ExecStart=/opt/myapp/bin/start.sh
Restart=on-failure
User=myapp

[Install]
WantedBy=multi-user.target

Required Sections and Directives

SectionKeyPurpose
[Unit]DescriptionHuman-readable name
[Unit]AfterStartup ordering relative to other units
[Service]TypeHow the service demonizes
[Service]ExecStartCommand to execute on start
[Install]WantedByTarget that enables this unit

Type

  • simple — process stays in foreground, systemd monitors it directly
  • forking — process forks and parent exits (classic daemon pattern)
  • oneshot — runs once and exits, useful for one-off tasks
  • exec — like simple, but waits for ExecStartPre to finish first

For modern applications written in Go, Node.js, or similar, simple is almost always correct.

Example: Application Service

[Unit]
Description=Backend API Service
Documentation=https://internal.example.com/docs
After=network-online.target postgresql.service
Wants=network-online.target

[Service]
Type=simple
User=appuser
Group=appgroup
WorkingDirectory=/opt/api
ExecStart=/opt/api/start.sh
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5s
TimeoutStartSec=30s
TimeoutStopSec=60s
Environment=NODE_ENV=production
EnvironmentFile=/etc/default/api
StandardOutput=journal
StandardError=journal
SyslogIdentifier=api-backend

# Hardening
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/opt/api /var/log/api
ProtectKernelTunables=true
ProtectControlGroups=true

[Install]
WantedBy=multi-user.target
Note

The --user flag creates a user-level service that lives with the user’s session instead of the system. Useful for personal tooling without root access.

Python: venv and gunicorn

Point ExecStart at the venv interpreter, not system python. That keeps service dependencies off the host packages:

[Service]
Type=simple
User=appuser
Group=appuser
WorkingDirectory=/opt/myapp
ExecStart=/opt/myapp/venv/bin/python -m myapp
EnvironmentFile=/etc/myapp/env
Restart=on-failure
RestartSec=5

For a WSGI app:

ExecStart=/opt/myapp/venv/bin/gunicorn --workers 3 --bind 127.0.0.1:8000 myapp.wsgi:application

Keep secrets in EnvironmentFile (KEY=VALUE), mode 600, owner root:appuser. If the unit fails to start, journalctl -u myapp usually shows a traceback, a missing WorkingDirectory, or a missing env file.

Production-Ready Directives

Restart Behavior

Restart=on-failure      # on non-zero exit code
Restart=on-abnormal     # on signal or timeout
Restart=always          # unconditionally, even after clean stop
RestartSec=5

Dependencies

After=network.target         # after network is up
After=postgresql.service     # after specific service
Wants=network-online.target  # weak dependency, continue if unavailable
Requires=postgresql.service   # hard dependency, fail if unavailable

Resource Limits

LimitNOFILE=65536
LimitNPROC=4096
MemoryMax=512M
CPUQuota=50%

Logging

StandardOutput=journal
StandardError=journal
SyslogIdentifier=myapp

Reading logs:

journalctl -u myapp -f
journalctl -u myapp --since "1 hour ago"
journalctl -u myapp -p err

Activation and Management

# Reload unit files from disk
sudo systemctl daemon-reload

# Start the service
sudo systemctl start myapp

# Check current status
sudo systemctl status myapp

# Enable on boot
sudo systemctl enable myapp

# Reload config without restarting
sudo systemctl reload myapp

# Full restart
sudo systemctl restart myapp

# Stop
sudo systemctl stop myapp

# Disable from boot
sudo systemctl disable myapp

Syntax check without applying changes:

systemd-analyze verify /etc/systemd/system/myapp.service

Common Mistakes

Missing ExecStart. Service fails immediately with Unit entered failed state.

Type not specified. Defaults to simple, but if your process self-daemonizes, you need forking.

Wrong path to script. Verify the file exists and is executable. Systemd does not validate paths at parse time — it simply fails to start.

Runtime edit without reload. Changed the file, forgot daemon-reload. Systemd continues using the old version.

WorkingDirectory does not exist. If you specify it, the directory must be present.

Restart=always without ExecStop. If the process exits cleanly, always restarts it anyway. This means systemctl stop may not behave as expected for long-running services.

Warning

Never edit package-provided unit files directly in /usr/lib/systemd/system/. Updates overwrite them. Use drop-in files in /etc/systemd/system/<name>.service.d/ instead.

Drop-in example:

mkdir -p /etc/systemd/system/myapp.service.d
# /etc/systemd/system/myapp.service.d/override.conf
[Service]
Environment=DEBUG=1
RestartSec=10s
sudo systemctl daemon-reload
sudo systemctl restart myapp

47 - kubectl whoami and Service Account Permission Checks

When deploying an application to Kubernetes, the most common failure is a service account that cannot do what it should. Permission denied when creating a secret, rejection on list pods, refusal on update. kubectl whoami and kubectl auth can-i let you quickly identify exactly who cannot do what.

kubectl whoami plugin

Note

kubectl whoami is not a built-in kubectl command. It is a plugin from krew or a standalone binary. Install with kubectl krew install whoami or download from GitHub.

After installation the command shows the current context and associated service account:

kubectl whoami

Output:

system:serviceaccount:default:myapp

Without the plugin, the same information is available via:

kubectl auth can-i --list

The beginning of the output shows User: system:serviceaccount:default:myapp — that is your whoami.

kubectl auth can-i: checking arbitrary subject permissions

The built-in kubectl auth can-i command checks permissions without impersonating the target subject. Syntax:

kubectl auth can-i <verb> <resource> --as=<subject>

The --as format for a service account:

# Within a namespace
kubectl auth can-i list pods \
  --as=system:serviceaccount:default:myapp

# Cross-namespace
kubectl auth can-i list pods \
  --as=system:serviceaccount:production:myapp \
  -n production

Responses: yes or no.

To check the full permission set of an SA:

kubectl auth can-i --list --as=system:serviceaccount:default:myapp
Tip

Combining --as with --list is a fast SA permission audit. The output contains a table with resources, versions, and permitted actions.

kubectl auth can-i flags

FlagPurpose
--as <subject>Subject to check
-n <namespace>Subject’s namespace (for SA)
--listAll permissions of the subject
--quiet / -qExit code only, no output
--no-headersNo table headers

Exit code: 0 for yes, 1 for no. Useful for scripts:

kubectl auth can-i delete secrets --as=system:serviceaccount:default:myapp
if [ $? -eq 0 ]; then
  echo "SA can delete secrets — review the policy"
fi

Binding ClusterRole to a Service Account: RoleBinding vs ClusterRoleBinding

A service account is bound to a role through two resources. The difference is scope.

RoleBinding binds a Role or ClusterRole to an SA within a single namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: myapp-pod-reader
  namespace: default
subjects:
  - kind: ServiceAccount
    name: myapp
    namespace: default
roleRef:
  kind: ClusterRole        # reference to ClusterRole
  name: pod-reader        # ClusterRole name
  apiGroup: rbac.authorization.k8s.io
Note

RoleBinding can reference either Role or ClusterRole. With Role, permissions are limited to that namespace. With ClusterRole, permissions apply within the RoleBinding’s namespace but inherit all ClusterRole rules.

ClusterRoleBinding binds a ClusterRole cluster-wide:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: myapp-cluster-reader
subjects:
  - kind: ServiceAccount
    name: myapp
    namespace: default
roleRef:
  kind: ClusterRole
  name: cluster-reader
  apiGroup: rbac.authorization.k8s.io
ScenarioResource
SA needs permissions in one namespace onlyRoleBinding → ClusterRole
SA needs cluster-wide permissionsClusterRoleBinding
Shared role for different SA across namespacesClusterRoleBinding with multiple subjects

Subject in CSR: reading and verifying

A CSR (Certificate Signing Request) contains spec.username and spec.groups. For an SA it looks like this:

kubectl get csr <csr-name> -o jsonpath='{.spec}' | jq .

Output:

{
  "request": "base64-encoded-certificate-request",
  "username": "system:serviceaccount:default:myapp",
  "groups": ["system:serviceaccount", "system:serviceaccounts", "system:serviceaccounts:default"],
  "uid": "...",
  "extra": {
    "authentication.kubernetes.io/credential-id": ["..."],
    "authentication.kubernetes.io/node-name": ["..."]
  }
}

Three components in username separated by colons: system:serviceaccount:<namespace>:<name>.

To check permissions of a specific SA from a CSR:

# Extract SA from CSR
SUBJECT=$(kubectl get csr <csr-name> -o jsonpath='{.spec.username}')
kubectl auth can-i list pods --as="$SUBJECT"
Warning

CSR is created by kubelet when a node connects or through ServiceAccount admission. If an SA uses the TokenRequest API (the modern approach), no CSR is generated — the token is issued directly.

Quick check in CI/CD

In a pipeline you need to confirm the deploy service account has sufficient rights before running manifests:

#!/bin/bash
SA="system:serviceaccount:ci-runner:deployer"
RESOURCES=("pods" "services" "configmaps" "secrets")

for RES in "${RESOURCES[@]}"; do
  kubectl auth can-i create "$RES" --as="$SA" -n "$NAMESPACE" || {
    echo "ERROR: SA cannot create $RES"
    exit 1
  }
done

echo "Permissions confirmed"
kubectl auth can-i get pods --as="$SA" -n "$NAMESPACE"
Tip

Add this check after kubectl apply or helm template — an early exit is cheaper than a crashed pod with ImagePullBackOff due to missing imagePullSecrets.

Useful output for diagnostics in CI logs:

kubectl auth can-i --list --as="system:serviceaccount:default:myapp" --no-headers

Lines without headers are easy to grep for suspicious wildcards like */* or */delete.

48 - ngrep: grep for Network Packets in Real Time

Ngrep applies grep-style pattern matching to network packets. When you need to see exactly what two services are exchanging over the wire and tcpdump drowns you in noise, ngrep isolates the payload content you care about.

Installation

Ngrep ships in the standard repositories of most distributions.

# Debian/Ubuntu
apt install ngrep

# RHEL/CentOS/Alma
yum install ngrep

# macOS
brew install ngrep

Running ngrep requires root privileges or the CAP_NET_RAW and CAP_NET_ADMIN capabilities.

Basic Syntax

ngrep [options] [pattern] [BPF filters]

Minimal invocation catches all packets on a port:

ngrep -i 'password' port 80

Breaking it down:

  • -i — case-insensitive search
  • 'password' — regex pattern matched against payload
  • port 80 — BPF filter: traffic on port 80 only

Output resembles tcpdump with decoded payload:

T 10.0.1.15:54321 -> 10.0.2.10:80 [AP]
GET /api/v1/users HTTP/1.1..
Host: api.example.com....

Common Flags

FlagDescription
-iCase-insensitive search
-W bylineLine-wrap output, strip escape sequences
-qQuiet: show matches only, suppress metadata
-tPrefix each packet with timestamp
-d eth0Listen on a specific interface
-n 1Exit after first match
-c NLimit output to N characters per line
-xShow hexdump instead of ASCII

Practical example with timestamps:

ngrep -i -t 'POST' port 8080

Port and Protocol Filtering

Ngrep accepts standard tcpdump BPF filters. Common patterns:

# Specific port
ngrep 'GET' port 80

# Port range
ngrep 'error' portrange 8000-9000

# By host
ngrep 'token' host 10.0.1.50

# Combined filter
ngrep -i 'auth' host 10.0.1.50 and port 443

# Outbound traffic only
ngrep 'response' dst port 8080
Note

Ngrep parses BPF filters identically to tcpdump. The port, host, and, and or syntax works the same way in both tools.

HTTP Requests in Docker Containers

A frequent task is tracking HTTP traffic between containers. Two approaches work.

First: run ngrep inside the container

Works if the container is Alpine or Debian-based and you have access:

docker exec -it <container_id> sh -c "apt-get update && apt-get install -y ngrep"
docker exec -it <container_id> ngrep -i 'Content-Type' port 80

Second: listen on docker0

When containers communicate over the host’s bridge network:

# Find the docker network interface
ip addr show docker0

# Listen on it
ngrep -W byline 'HTTP' host 172.17.0.2 and port 80 -d docker0

Check a container’s IP:

docker inspect -f '{{.NetworkSettings.IPAddress}}' <container_name>

Limitations and Alternatives

Ngrep only works with plaintext traffic. It cannot decrypt TLS/HTTPS — you will see binary garbage instead of content.

Other limitations:

  • Does not understand HTTP/2 or HTTP/3 — these protocols are binary
  • No built-in JSON or XML parsing, only text-based matching
  • Performance lags behind tcpdump under high load

Alternatives by use case:

TaskTool
Quick pcap capturetcpdump -i eth0 -A port 80
Detailed protocol analysistshark -Y http -i eth0
HTTPS monitoring (key required)Wireshark with decryption
gRPC tracinggrpcurl or Wireshark with Protobuf dissector

For most debugging tasks, ngrep plus tcpdump covers the bases. If you need to parse a specific protocol in depth, tshark with filters gives more control — but requires learning Wireshark’s syntax.

49 - socat: Forwarding Unix Sockets Over TCP

Sometimes you need to reach a Unix socket from a host where that socket doesn’t physically exist. SSH tunnels won’t help — they only work with TCP ports. socat solves this: it opens a TCP listener and forwards connections to a Unix socket, and the client just connects over the network.

Installation

The package is available in every major distribution. On Debian/Ubuntu:

apt install socat

On RHEL/CentOS:

yum install socat
# or
dnf install socat

Alpine:

apk add socat

Verify:

socat -V
# socat version 1.7.4.4

Basic Forwarding: TCP-LISTEN + UNIX-CONNECT

Server side. Listen on a TCP port and redirect traffic to a Unix socket on connection:

socat TCP-LISTEN:2375,fork UNIX-CONNECT:/var/run/docker.sock

Flags:

  • TCP-LISTEN:2375 — opens port 2375
  • fork — spawns a child process for each connection; without it socat accepts one connection and exits

On the client side, work as usual — for example, curl the Docker API:

curl http://localhost:2375/version

If the client is on a remote host, specify the server IP:

curl http://192.168.1.100:2375/version
Warning

Docker listens on the local socket by default. Exposing TCP-LISTEN externally without TLS or firewall is a risk. Restrict the bind to an interface: TCP-LISTEN:2375,bind=127.0.0.1.

Stop the forward — Ctrl+C or kill by PID.

Client Test via STDIO

To quickly verify socket availability or send a manual command, use STDIO on the client side:

socat STDIO UNIX-CONNECT:/var/run/docker.sock

After starting, enter raw HTTP requests. Example session:

socat STDIO UNIX-CONNECT:/var/run/docker.sock
GET /version HTTP/1.0

HTTP/1.1 200 OK
Content-Type: application/json
{"ApiVersion":"1.45","Version":"24.0.7"...}

Exit — Ctrl+D or Ctrl+C. This is handy for debugging APIs without curl and without setting environment variables.

For a TCP connection over the network, the client runs symmetrically:

socat STDIO TCP:192.168.1.100:2375

Abstract vs Filesystem Sockets

Unix sockets come in two types. The difference matters for socat operation.

Filesystem sockets — bound to the filesystem. Path starts with /:

/var/run/docker.sock
/run/user/1000/pulse/runtime/native
/tmp/mysql.sock

Abstract sockets — live in kernel memory, have no filesystem representation. Path starts with \0 or @ (ASCII zero and at-sign). Docker in rootless mode uses these:

# Displaying abstract socket in ls
ls -la /run/user/1000/docker.sock
# srwxr-xr-x 1 user user 0 Jan 15 10:00 /run/user/1000/docker.sock

# Actual path in kernel starts with \0
# In socat, write:
socat TCP-LISTEN:2375,fork UNIX-CONNECT:@/docker.sock

Check socket type:

ss -x | grep docker
# u_str  LISTEN  0  4096  /run/user/1000/docker.sock  12345  * 0

# If path starts with @ — abstract

In socat syntax:

  • @/path/to/socket — abstract socket
  • /path/to/socket — filesystem socket
Note

Abstract sockets are invisible to processes without namespace access. This is an advantage for isolation but complicates forwarding between containers.

Timeout Flags

By default, socat waits forever. For automation and scripts, you need timeouts.

FlagDescription
readtimeout=SECONDSRead timeout
writetimeout=SECONDSWrite timeout
timeout=SECONDSTimeout for both operations

Example with a general timeout:

socat TCP-LISTEN:2375,fork,timeout=30 UNIX-CONNECT:/var/run/docker.sock

Connection closes after 30 seconds of inactivity.

Separate timeouts for client and server:

# Server waits 10 sec for write, client waits 5 sec for read
socat TCP-LISTEN:2375,forever,writewait=10 UNIX-CONNECT:/var/run/docker.sock,readtimeout=5

In cron scripts or systemd units, set a timeout or the process will hang on network disruption:

socat TCP-LISTEN:2375,fork,timeout=60 UNIX-CONNECT:/var/run/docker.sock

For infinite waiting without fork, forever works, but in production combine it with system limits:

socat TCP-LISTEN:2375,reuseaddr,timeout=0 UNIX-CONNECT:/var/run/docker.sock
Tip

reuseaddr lets you quickly restart socat without “Address already in use” errors.

Common Errors

Permission denied accessing the socket

# Check permissions
ls -la /var/run/docker.sock
# srw-rw---- 1 root docker

# Add user to the group
usermod -aG docker username

Connection refused

Verify socat is running and listening on the port:

ss -tlnp | grep 2375
# LISTEN 0 5 *:2375 *:*  users:(("socat",pid=1234))

Firewall:

iptables -L -n | grep 2375
# ACCEPT  tcp  --  0.0.0.0/0  0.0.0.0/0  tcp dpt:2375

One request — and socat dies

Missing fork. Each socat instance handles one connection and exits. Add the flag:

socat TCP-LISTEN:2375,fork UNIX-CONNECT:/var/run/docker.sock

Abstract socket not found

Ensure the correct prefix. Docker in rootless uses @, but this is ASCII 0:

# Shows actual path in kernel
cat /proc/$(pgrep dockerd)/net/unix | grep docker

In socat, write @/docker.sock, not @@/docker.sock.

Systemd Service for Persistent Forwarding

For a permanent forward, wrap it in a systemd unit:

[Unit]
Description=socat Docker socket forwarder
After=network.target

[Service]
ExecStart=/usr/bin/socat TCP-LISTEN:2375,fork,reuseaddr,timeout=60 UNIX-CONNECT:/var/run/docker.sock
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
systemctl enable socat-docker-forward
systemctl start socat-docker-forward

Don’t forget to restrict the bind to an interface if you don’t want the port exposed externally.

In Closing

socat is the Unix way for transparent forwarding of anything to anywhere. For Unix sockets over TCP, two processes and a minute of configuration are enough. Keep timeouts in mind, don’t forget fork, and don’t expose ports without authentication on public networks.

50 - SSH Escape Sequences: Reviving a Frozen Terminal

SSH session froze, Ctrl+C does nothing, Ctrl+D spits out garbage — familiar situation. Before closing the terminal and losing the session, try built-in escape sequences. They operate at the SSH client level before data reaches the remote host.

How to invoke escape sequences

The default escape character is tilde (~). The combination works only at the beginning of a line. Press Enter, then ~, then the desired symbol. For example, ~. terminates the connection.

Note

If tilde does not work — make sure you pressed Enter before it. In the middle of terminal output the sequence is ignored.

Escape sequence reference

ssh> ~?

This command prints the list of available escape sequences directly into the terminal.

Supported escape sequences:
 ~.  - terminate connection (and any multiplexed sessions)
 ~B  - send a BREAK to the remote system
 ~C  - open a command line
 ~R  - request rekey
 ~V/~v  - decrease/increase verbosity (LogLevel)
 ~^Z  - suspend ssh
 ~#  - list forwarded connections
 ~&  - background ssh (when waiting for connections to terminate)
 ~?  - this message
 ~  - send the escape character by typing ~~

You do not need to memorize everything. Keep three scenarios in mind — they cover 90% of problems.

Emergency disconnect

# Example: frozen scp or sftp
# Press Enter, then:
~.

~. closes the SSH connection immediately, including all multiplexed sessions. Works even if the remote host is not responding. Do not confuse it with ~^Z (suspend) — that leaves the session alive as a background process.

Warning

Forced disconnect does not send SIGHUP to the remote side. If an important process was running without nohup — it will die.

Breaking the data flow

# Output of a command is stuck or SCP is slow
~V

~V (Shift+v) lowers the logging level. The reverse command — ~v — increases verbosity. In practice useful when you see a stream of meaningless debug messages and want to hide them.

A more radical option — ~B. Sends BREAK to the remote system. This can interrupt a stuck process that is listening on a serial port. Use deliberately: not all systems handle BREAK correctly.

Built-in command mode

# Press Enter, then:
~C

You enter the SSH client command line. Here you can manage tunnels on the fly without reconnecting.

ssh> -L 8080:localhost:80
# Added local port 8080 -> remote:80

ssh> -R 2222:localhost:22
# Added reverse proxy on remote host

ssh> -D 1080
# SOCKS proxy on port 1080

ssh> -KL 8080
# Remove forwarded local port

ssh> help
Tip

Command mode is convenient when you forgot to forward a port before connecting. No need to reconnect — add it directly from the session.

Format: -L [bind_addr:]port:host:hostport and -R [bind_addr:]port:host:hostport. Bind address localhost restricts forwarding to the local interface only.

Forced rekey

~R

Forcibly initiates re-keying of the session. Sometimes useful when you notice strange delays or suspect encryption issues. In practice rare, but worth knowing.

Escape sequence summary

SequenceAction
~.Emergency disconnect
~VDecrease verbosity
~vIncrease verbosity
~CCommand line (tunnels)
~RForced rekey
~BSend BREAK
~?Show help
~^ZSuspend ssh
~#List forwarded connections
~&Background mode (on disconnect)
~~Send literal tilde

Changing the escape character

If you need tilde in the remote session (for example, connecting to Cisco equipment), change the escape character:

ssh -e '^Z' user@host

Now escape sequences are invoked via ^Z instead of ~. Found in specific scenarios — do not touch without need.

Note

You cannot change the character on the fly in an active session. Only at connection time via the -e flag.

Common pitfalls

  1. ~ does not work — forgot to press Enter before it.
  2. Session is frozen but you are not sure SSH is still alive — try ~. blind. It will not make things worse.
  3. Multiplexing session (ssh -M) is bound to the control socket. ~. closes all related sessions at once.
  4. If you use PuTTY — escape sequences are different. There ~. sends a literal tilde and dot. Use Ctrl+] and then close.

When escape sequences will not help

  • Network is physically unavailable — only disconnect.
  • Remote process consumed stdin — ssh will receive nothing.
  • Problem is at the terminal emulator level — killall ssh, then check tmux/screen.

SSH escape sequences are the first-line tool when a session freezes. Keep ~. for disconnect and ~C for tunnels in mind — that is enough for everyday use.

51 - What is Self-Hosted and Why It's So Popular

You’re paying for Notion, Dropbox, and Google Drive. Then Notion raises prices, Dropbox caps storage, and Google “improves” the Docs interface. Self-hosted is taking infrastructure into your own hands and breaking the vendor dependency loop.

Definition: Not Your Cloud

Self-hosted means deploying and operating applications on your own servers or VPS instead of using SaaS alternatives. The server can sit at home, in a data center, or be a VM at a hosting provider—the point is that the hardware is under your control, not a third party’s.

The difference from classic hosting is straightforward: with virtual hosting you rent part of a server with pre-installed software. With self-hosted you install whatever you want and configure it however you like. It’s closer to VPS or bare metal, but focused on applications rather than websites.

Note

Self-hosted ≠ self-administered. You can hire an admin, use managed servers, or deploy ready-made images. The point is that data and the service belong to you, not a cloud provider.

Typical Self-Hosted Stack

Most self-hosted setups revolve around Docker and Docker Compose. It’s the de facto standard for isolation and reproducibility.

# Install Docker on Ubuntu
curl -fsSL https://get.docker.com | sh
sudo usermod -aG docker $USER

# Minimal docker-compose.yml for Nextcloud
cat << 'EOF' > docker-compose.yml
version: "3.8"
services:
  app:
    image: nextcloud:latest
    restart: unless-stopped
    ports:
      - "8080:80"
    volumes:
      - ./data:/var/www/html
    environment:
      - MYSQL_HOST=db
      - MYSQL_DATABASE=nextcloud
      - MYSQL_USER=nextcloud
      - MYSQL_PASSWORD=changeme
  db:
    image: mariadb:latest
    restart: unless-stopped
    volumes:
      - ./db:/var/lib/mysql
    environment:
      - MYSQL_ROOT_PASSWORD=rootpw
      - MYSQL_DATABASE=nextcloud
      - MYSQL_USER=nextcloud
      - MYSQL_PASSWORD=changeme
EOF

docker compose up -d

Nextcloud usually sits behind a reverse proxy—Traefik or Caddy. Traefik auto-discovers containers via labels, Caddy works with zero-config TLS.

# Traefik with Docker provider
cat << 'EOF' >> docker-compose.yml
  traefik:
    image: traefik:v3.0
    restart: unless-stopped
    ports:
      - "80:80"
      - "443:443"
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - ./traefik/acme.json:/acme.json
    command:
      - "--providers.docker"
      - "--providers.docker.exposedbydefault=false"
EOF

Common home setup stack:

CategoryExamplesPurpose
FilesNextcloud, SyncthingStorage and sync
MediaJellyfin, PlexPersonal streaming
PasswordsBitwarden, VaultwardenPassword manager
CodeGitea, ForgejoGit repositories
Smart homeHome AssistantIoT controller
MonitoringGrafana, PrometheusMetrics and alerts

Why It’s Gone Mainstream

The main driver is data control. When clouds regularly leak data and corporate privacy policies shift with profit margins, the idea of “my server, my rules” appeals beyond just the paranoid.

Financial factors matter too. One server for $10–20/month replaces a stack of subscriptions. Bitwarden Premium costs $10/year, Nextcloud AIO is $29/year. Even a personal NAS with multiple disks pays for itself in a couple years versus Dropbox and Google One.

Less obvious: vendor lock-in. Notion exports data poorly, Slack doesn’t provide proper exports, and Discord keeps closing APIs. When you run your own instance, migration is just a backup and spin-up on a new server.

Tip

Keep your configuration in Git. Twenty docker-compose.yml and compose-override.yml files in a repo isn’t just a backup—it’s a reproducible setup that deploys to new hardware in minutes.

The last factor is the learning curve. Self-hosted has gotten easier: ready-made images, auto-TLS, Compose files on GitHub. This attracts people who want to understand infrastructure, not just use a service.

What Holds People Back

The main issue is maintenance burden. SaaS updates without your involvement. Self-hosted requires regular apt update && docker compose pull && docker compose up -d and monitoring for outdated images.

Security is the second bottleneck. Your server is your responsibility. Vulnerability in Nextcloud? That’s your CVSS, your patches, your downtime. Clouds at least patch faster and notify users.

# Check for outdated images
docker images --format "{{.Repository}}:{{.Tag}}"
# Or via Trivy
trivy image --severity HIGH,CRITICAL nextcloud:latest

Uptime is the third barrier. For a home setup it’s not critical. For a work tool you need backup connectivity, monitoring, and a process supervisor. Systemd unit or Watchtower with auto-restart is the minimum.

TaskToolComplexity
Auto-update imagesWatchtowerLow
Container monitoringcAdvisor + GrafanaMedium
Downtime alertsHealthchecks.io, Uptime KumaLow
Data backupRestic, DuplicatiMedium

When It Makes Sense

Self-hosted pays off in three scenarios.

Homelab and personal use. You effectively get a private cloud for the cost of a server. Files, passwords, photos, media library—all under your control. For many it’s both a hobby and hands-on practice.

Small teams without compliance requirements. Gitea for code, Vaultwarden for passwords, Outline or HedgeDoc for documents replace GitHub, 1Password, and Confluence. Budget is $20–50/month on a VPS.

Specific requirements. Custom integrations, above-average privacy, data that can’t live with third parties. This can be a personal project, a cleaning company handling personal data, or an early-stage startup.

Warning

Don’t try to replace managed services with self-hosted if your team has no time budget for operations. An SRE engineer on half-time will cost more than a Notion subscription and burn through goodwill faster.

Typical path: start with one service, add a second, then a third. By the sixth you realize you have a homelab with ten containers and a 200-line Ansible playbook. That’s fine—many people found their way into DevOps exactly this way.

52 - cockpit-modules: web panels for day-to-day operations

Cockpit covers basic Linux administration in the browser: services, logs, networking, accounts. Firewall, fail2ban, cron, and Let’s Encrypt sit outside that set — either there is no panel, or it is too generic.

The cockpit-modules group is a set of separate modules for those jobs, plus a store that installs them on the host. Each module lives in its own repository. This is a map of the group, not a walkthrough of UI and commands. Individual panels get their own articles.

Why extra modules

Cockpit grows through packages in /usr/share/cockpit/ (or ~/.local/share/cockpit/ for the current user). A module is a manifest.json, HTML, and JS that call host tools via cockpit.spawn.

The group targets a typical home or small server:

  • perimeter: UFW and fail2ban;
  • scheduling: crontab and systemd timers;
  • TLS: certbot and the Cockpit panel certificate;
  • delivery: a module store backed by the same GitLab group.

The UI is in Russian and matches the rest of Cockpit (PatternFly v5). You need Cockpit 264+ and administrator rights for writes.

Shared repository rules

Modules share the same layout so the store and install scripts stay consistent:

  • sources in pkg/<id>/;
  • install.sh — system-wide or --user;
  • releases tagged v.M.m.p;
  • Docker Compose for local development (localhost:9090);
  • MIT license.

Host commands run as argv arrays, not shell strings. Input is validated on the client. Without Cockpit admin rights the panel stays read-only.

A project appears in the store catalog only if it lives in the cockpit-modules group. The catalog is not arbitrary GitLab — it is an explicitly trusted set.

Module store

cockpit-modules-store is the entry point. Under Tools you get Module store: the group project list, tags, status (available / installed / update ready), install of a release archive into /usr/share/cockpit/, and removal of user modules. Built-in Cockpit pages are left alone.

You can still install the other panels by hand with install.sh, but the store removes a clone/copy cycle per repository.

Quick install of the store itself:

curl -fsSL https://gitlab.com/cockpit-modules/cockpit-modules-store/-/raw/master/install.sh | sudo bash

Panels in the group

ModuleJob
UFWufw package, status, policies, allow/deny/reject/limit rules
Fail2banpackage and service, jails, ban/unban IPs
Cronuser crontabs, /etc/crontab and /etc/cron.d/, systemd timers
CertManagerLet’s Encrypt / certbot, Cockpit panel TLS, renew

UFW is a full Uncomplicated Firewall cycle without ufw status numbered over SSH: the package via APT/DNF/YUM/Pacman, enable/disable, default policies, rule list, and add-by port, protocol, and source. Panel write-up: cockpit-ufw-module.

Fail2ban is the neighbouring perimeter panel: daemon install, jails, banned addresses, ban and unban in a selected jail or globally.

Cron is the scheduler people usually edit in nano. User crontabs, system files, and systemd timers in one place. Editing timer unit files is out of scope.

CertManager covers certificates for sites and for the Cockpit panel itself. HTTP-01 and DNS-01, binding a lineage to the panel, upload of your own crt+key, auto-renew via certbot.timer.

How it fits together

On a clean host the natural order is: Cockpit → store → UFW and fail2ban → CertManager for HTTPS on the panel → Cron if needed. Modules are independent: cron does not depend on certbot.

Sources, issues, and releases live at gitlab.com/cockpit-modules. First per-panel write-up: UFW.

53 - dig: DNS Query Debugging in CLI

DNS resolvers return the wrong address, clients don’t see updates, or it’s unclear which server is handling requests. dig (Domain Information Groper) is the standard CLI tool for DNS diagnostics. Works on Linux, macOS, and Windows via WSL.

Installation

# Debian/Ubuntu
apt install dnsutils

# RHEL/CentOS/Alma
dnf install bind-utils

# macOS — ships with the system
# Windows — via WSL or ISC official binaries

Basic Flags

dig has two classes of options: short flags (start with -) control query behavior, while keywords with + control output format.

dig example.com

Without arguments the output is verbose — lots of header information. For debugging you pick the pieces you need.

FlagPurpose
-b <addr>Outbound IP address (multi-interface hosts)
-f <file>Read queries from file, one per line
-p <port>Non-standard DNS server port
-t <type>Record type: A, AAAA, MX, TXT, SOA, NS, CNAME, ANY
-c <class>Network class (default: IN — Internet)
-x <addr>Reverse lookup (PTR)
-6Force IPv6
-4Force IPv4

Quick Response: +short

For scripts and quick checks:

dig example.com +short
# 93.184.216.34
dig mail.example.com +short
# 10 mailgateway.example.com.

If the record doesn’t exist — empty output. For CNAMEs you see the final address but not the chain.

For AAAA records:

dig example.com AAAA +short
# 2606:2800:220:1::248a:2873
Note

+short doesn’t show TTL and doesn’t guarantee it’s the final answer. A CNAME loop will return the last record in the chain, not an error.

Full Response: +noall +answer

When you need TTL, canonical name, and all records at once:

dig example.com +noall +answer
# example.com.         86400   IN      A       93.184.216.34
dig example.com MX +noall +answer
# example.com.         3600    IN      MX      10 mailstore1.example.com.
# example.com.         3600    IN      MX      20 mailstore2.example.com.

TTL in seconds. Small values (60–300) mean the record changes frequently.

Verbose response with timing:

dig example.com +stats

Reverse Lookup: -x

Reverse zone: IP → hostname.

dig -x 93.184.216.34 +short
# domain.example.com.
dig -x 8.8.4.4 @1.1.1.1 +short
# dns.google.
Tip

Not all PTR zones are populated. Empty response with -x is normal, especially for client addresses.

IPv6: -6

Force IPv6 transport to the DNS server:

dig -6 @2001:4860:4860::8888 example.com AAAA +short
# 2606:2800:220:1::248a:2873

If you want an AAAA record over IPv4 transport — just request the type:

dig @1.1.1.1 example.com AAAA +short

Chain Tracing: +trace

Shows the path from root servers to the final answer:

dig example.com +trace +noall +answer

Output is split into sections: . (root), TLD (.com), authoritative NS, response.

dig internal.example.com +trace +noall +answer
# .                       518400  IN      NS      a.root-servers.net.
# com.                    172800  IN      NS      a.gtld-servers.net.
# example.com.            172800  IN      NS      a.iana-servers.net.
# internal.example.com.   300     IN      A       10.0.1.50

+trace is slow — walks the hierarchy recursively. Use it for NXDOMAIN and SERVFAIL diagnosis.

Specific Resolver: @server

By default dig uses the system resolver from /etc/resolv.conf. Specifying explicitly compares responses or bypasses local cache:

dig @8.8.8.8 example.com +short
dig @1.1.1.1 example.com +short
dig @9.9.9.9 example.com +short
dig @ns1.example.com example.com AXFR +short

AXFR transfer works only if the NS allows it.

Warning

Full zone AXFR is sensitive. Don’t do this on third-party NS without reason.

Common Scenarios

Checking SOA and NS records

dig example.com SOA +short
# ns1.example.com. admin.example.com. 2024011501 7200 3600 1209600 86400

dig example.com NS +short
# ns1.example.com.
# ns2.example.com.

Serial in SOA — if you updated records but it didn’t increase, transfer hasn’t happened.

TTL of a specific record

dig example.com A +noall +answer +ttlid
# example.com.         300     IN      A       93.184.216.34

300 seconds is low TTL, normal for frequently changing records. Static records usually sit at 3600+.

CNAME chain

dig www.example.com +trace +noall +answer

The chain displays in full. If a redirect broke somewhere — NXDOMAIN or SERVFAIL at which step becomes clear from the output.

Comparing resolvers

for ns in 8.8.8.8 1.1.1.1 9.9.9.9; do
  echo "=== $ns ===";
  dig @$ns example.com +short;
done
# Output
=== 8.8.8.8 ===
93.184.216.34
=== 1.1.1.1 ===
93.184.216.34
=== 9.9.9.9 ===
93.184.216.34

If addresses differ — the issue isn’t on your service side, it’s with that specific resolver or record propagation.

SERVFAIL analysis

dig @ns1.example.com _sip._tcp.example.com SRV +noall +answer +stats

Check flags: SERVFAIL in the response means the NS couldn’t fetch data. Verify the requested record type actually exists in the zone.

ANY query (carefully)

dig example.com ANY +short

Deprecated in production. But for quick diagnostics of all record types on a new NS — it works.

54 - MkDocs: Project Documentation from Markdown

Documentation in a repository ages faster than anyone reads it: README links lead nowhere, sections are scattered across docs/, wiki/, and Confluence, and site search doesn’t work. MkDocs solves this predictably — it takes a folder of .md files and builds a static site. One config, one command for the deploy, familiar Markdown.

What is MkDocs

MkDocs is a static site generator for documentation written in Python. Input: a directory of Markdown files and a YAML config. Output: a ready site/ directory with HTML, served by any web server or hosted on GitHub Pages, GitLab Pages, S3. MkDocs core handles rendering; the theme defines look and features. The de-facto standard is Material for MkDocs.

Note

MkDocs does not use Jinja templates and does not require a database. It is static content that builds locally or in CI in a few seconds.

Installation

The minimum requirement is Python 3.8+. Install into a virtual environment to avoid polluting the system pip.

python3 -m venv .venv
source .venv/bin/activate
pip install mkdocs

Verification:

mkdocs --version

Typical output is mkdocs, version 1.6.x. The version matters because themes and plugins often require a specific range.

Useful packages installed alongside the core or separately:

PackagePurpose
mkdocs-materialMaterial theme, navigation, search, tabs
mkdocstringsDocumentation generated from Python docstrings
pymdown-extensionsAdditional Markdown extensions for Material
mkdocs-minify-pluginHTML/CSS/JS minification in site/
pip install mkdocs-material mkdocstrings[pymdownx]
Tip

Pin theme and plugin versions in requirements.txt. Material breaks compatibility between minor releases, as does mkdocstrings.

Creating a Project

mkdocs new scaffolds the project:

mkdocs new my-docs
cd my-docs

This creates a docs/ directory with index.md and an empty mkdocs.yml. That is the working minimum — nothing else is mandatory.

tree my-docs
my-docs
├── docs
│   └── index.md
└── mkdocs.yml

Directory Structure

docs/ is the single source of Markdown. The directory hierarchy maps directly to URLs. The file docs/guide/install.md becomes /guide/install/. An index.md in the root of docs/ is the landing page.

docs/
├── index.md
├── guide/
│   ├── install.md
│   └── config.md
├── reference/
│   └── cli.md
└── about.md

The site builds into the site/ directory next to mkdocs.yml. This directory is a build artifact — it is committed only for manual deploys, and usually built by CI.

Configuration mkdocs.yml

A minimal working config:

site_name: My Project Docs
site_url: https://example.com/docs/
docs_dir: docs
site_dir: site

theme:
  name: material

The full set of keys actually used in production:

KeyPurpose
site_nameSite title and default <title>
site_urlCanonical URL, required for sitemap.xml and robots.txt
site_descriptionDescription, goes into meta tags
docs_dirDirectory with Markdown, defaults to docs
site_dirWhere HTML is built, defaults to site
themeTheme and its parameters
navExplicit navigation, overrides auto-discovery
pluginsPlugins in load order
markdown_extensionsEnabled Markdown extensions
extraArbitrary variables read by the theme

Example with navigation, extensions, and plugins:

site_name: Service Docs
site_url: https://docs.example.com/
repo_url: https://github.com/example/service

theme:
  name: material
  features:
    - navigation.tabs
    - navigation.sections
    - search.highlight
    - content.code.copy
  palette:
    - scheme: default
      toggle:
        icon: material/brightness-7
        name: Dark theme
    - scheme: slate
      toggle:
        icon: material/brightness-4
        name: Light theme

nav:
  - Home: index.md
  - Guide:
      - Installation: guide/install.md
      - Configuration: guide/config.md
  - Reference:
      - CLI: reference/cli.md

markdown_extensions:
  - admonition
  - tables
  - toc:
      permalink: true
  - pymdownx.highlight:
      anchor_linenums: true
  - pymdownx.superfences
  - pymdownx.tabbed:
      alternate_style: true

plugins:
  - search
Warning

Enable search explicitly when using the plugins list. In newer Material versions it is no longer pulled in automatically from the theme.

Content

Markdown files are standard CommonMark with extensions. Useful constructs that work out of the box with the extensions enabled in the example above.

Admonitions:

> [!NOTE]
> A brief note for the reader.

> [!WARNING]
> This action may cause data loss.

Tabs with pymdownx.tabbed:

=== "Linux"

    ```bash
    sudo apt install foo
    ```

=== "macOS"

    ```bash
    brew install foo
    ```

Code highlighting with language specified:

```python
from mkdocs import config
print(config.DEFAULT_SCHEMA.keys())
```
Tip

Use heading anchors to link between pages. Material renders a # icon next to the heading when toc.permalink: true is set.

Internal links are relative paths from the current file:

See the [configuration section](config.md).

External links open in the same tab by default. To open in a new tab:

[Material for MkDocs](https://squidfunk.github.io/mkdocs-material/){target=_blank}

Build and Local Server

Local development — run the live server:

mkdocs serve

By default it listens on http://127.0.0.1:8000. Useful flags:

FlagEffect
--dev-addr 0.0.0.0:9000Change address and port, useful in a container
--strictBuild fails on any warning, including broken links
--livereloadPage reloads in the browser without F5 (on by default)
--no-livereloadDisable auto-refresh
--cleanRemove site/ before building

Build the artifact for deployment:

mkdocs build --clean --strict

--strict is mandatory in CI. Otherwise typos in links and missing files in nav will be silently built. The exit code on warning is zero, so without --strict the pipeline passes green with broken documentation.

Warning

Do not run mkdocs serve in production. It is a dev server with no authentication and file reload enabled.

Common errors on first run:

  • WARNING - A relative path to '...' is included in the 'nav' config. A file is listed in nav but missing from docs/. Check case and path.
  • WARNING - Documentation file 'x.md' is not included in the 'nav' configuration. The file exists but is not in navigation. Either add it to nav or rely on auto-navigation by removing the nav section entirely.
  • ERROR - Config value 'theme': The theme 'mkdocs' is not installed. The theme package is not installed, or the name is misspelled.

Material for MkDocs: Navigation, Search, Tabs

Material extends base MkDocs with three things that are almost always needed.

Navigation. Enabled through features in the theme section:

theme:
  name: material
  features:
    - navigation.tabs          # top-level tabs
    - navigation.sections      # tabs stick on scroll
    - navigation.top           # "back to top" button
    - navigation.indexes       # index.md becomes a section
    - navigation.tracking      # anchor in URL on navigation
    - toc.follow               # right-side TOC scrolls with text
Tip

The combination of navigation.tabs + navigation.sections gives the familiar “sticky” menu. Without sections the tabs scroll away with the content.

Search. The search plugin ships with Material but registers separately:

plugins:
  - search:
      separator: '[\s\-\.\_]+'

For Russian documentation, stemming is important — set search.lang: ru in the theme config. The list of supported languages is published in Material’s documentation and new languages are added regularly.

Tabs within a page. Implemented via the pymdownx.tabbed extension, connected above. The alternative syntax using !!! example blocks does not work in some themes — this is a frequent reason “why tabs don’t render.”

Code highlighting and copying. The “copy” button is enabled by the content.code.copy feature. Line numbers come from the pymdownx.highlight extension with linenums: true or globally via markdown_extensions:

markdown_extensions:
  - pymdownx.highlight:
      anchor_linenums: true
      line_spans: __span
      pygments_lang_class: true
  - pymdownx.inlinehilite
  - pymdownx.snippets
  - pymdownx.superfences
Warning

pymdownx.superfences is required for highlighting inside admonitions and tabbed blocks. Without it, code renders as a plain block without syntax highlighting.

Dark theme is a must-have for documentation read at night on-call. The toggle is configured in the example config above through palette with two schemes: default and slate. To make Material serve the correct scheme on first visit, add a custom script to <head> or use theme.palette.toggle with media — it works without JS and respects the user’s system settings.

55 - Squid: Internet Forwarding to Remote VM

Your VM in the cloud has no public IP or internet access is blocked via NAT, but the deployment needs wget/curl from inside. Squid on an intermediate host with a decent uplink solves this in ten minutes.

Why This Is Needed

I forward internet through Squid when a VM sits in an isolated network segment. An intermediate host with a public IP and network access becomes the proxy server. The application on the remote machine routes traffic through the tunnel.

Real-world cases: test environments without external connectivity, CI/CD agents in a private subnet, temporary traffic routing for debugging.

Installing and Basic Squid Setup

Install on the intermediate host (Ubuntu/Debian):

apt update && apt install -y squid

The default config lives in /etc/squid/squid.conf. Minimum working config:

# /etc/squid/squid.conf
http_port 3128

acl localnet src 10.0.0.0/8
acl localnet src 172.16.0.0/12
acl localnet src 192.168.0.0/16

http_access allow localnet
http_access deny all

Enable and start:

systemctl enable --now squid

Verify the port is listening:

ss -tlnp | grep 3128

ACL and Port Configuration

If access should be only via SSH tunnel from a specific address, replace localnet with the exact IP:

# Proxy accessible only from this address
acl tunneled src 203.0.113.50

http_access allow tunneled
http_access deny all

To change the port:

http_port 8080

After config changes, reload without restarting:

squid -k reconfigure
Note

If Squid fails to start after changes, check the log: journalctl -u squid -n 50.

SSH Tunnel to VM

On the remote machine, establish a tunnel to the intermediate host:

ssh -N -L 3128:localhost:3128 user@proxy-host

-N — don’t open a shell, forwarding only. -L binds local port 3128 to localhost:3128 on the remote host.

For background:

ssh -N -L 3128:localhost:3128 user@proxy-host &

Or via a systemd user service:

# ~/.config/systemd/user/proxy-tunnel.service
[Unit]
Description=SSH tunnel to proxy-host

[Service]
ExecStart=/usr/bin/ssh -N -L 3128:localhost:3128 user@proxy-host
Restart=always
RestartSec=10

[Install]
WantedBy=default.target
systemctl --user enable --now proxy-tunnel.service

Client Proxy Setup

On the remote VM, set environment variables for applications that respect http_proxy:

export http_proxy=http://localhost:3128
export https_proxy=http://localhost:3128

Or permanently in /etc/environment:

http_proxy=http://localhost:3128
https_proxy=http://localhost:3128

For curl/wget, variables suffice. For apt — additionally:

echo 'Acquire::http::Proxy "http://localhost:3128";' | tee /etc/apt/apt.conf.d/99proxy

Test:

curl -s --max-time 10 https://ifconfig.me

If it returns the intermediate host IP — it’s working.

Verification and Logging

Squid access log:

tail -f /var/log/squid/access.log

Format: time client/status code size method URL

Sample entry:

1703123456.123  1024 192.168.1.100 TCP_MEM_HIT/200 5123 GET http://example.com/file.tar.gz

Response codes: TCP_HIT — served from cache, TCP_MISS — fetched from network, TCP_DENIED — access denied by ACL.

To clear the cache before testing:

squid -k shutdown && rm -rf /var/spool/squid/* && squid -z && systemctl start squid
FlagDescription
http_portListening port
acl name src IP/maskAccess rule by IP
http_access allow|denyPermit or deny ACL
-k reconfigureReload config
-k shutdownGraceful stop
-zInitialize cache directories
Warning

Squid caches responses by default. For debugging, disable caching: add cache deny all to the config, then run squid -k reconfigure.

Nine minutes of setup — and the isolated VM has internet via proxy. If you need HTTPS transparent mode with certificate substitution — that’s a different story involving SSL-bump and CA generation.

56 - SSH certificates instead of authorized_keys

authorized_keys works fine for a handful of servers. Once you hit a dozen, it becomes a liability. Onboarding a new developer means manually distributing their public key across every machine. SSH certificates flip this model: one CA signs all public keys, and authorized_keys stays empty.

Why authorized_keys breaks at scale

authorized_keys requires your public key to exist on every target server. Scaling this creates:

  • a separate deployment step for key distribution during onboarding
  • no centralized revocation — removing a key means touching each host
  • key rotation touches every machine
  • no expiration means stale access accumulates

An SSH certificate is a CA signature on your public key. The server only needs to trust the CA — your key never needs to be present locally.

How SSH certificates work

Two entities: the CA (Certificate Authority) and the signed key. The CA is a standard SSH keypair, typically ed25519 or RSA. Signing creates the certificate with ssh-keygen -s <ca_private> -I <identifier> <key.pub>, producing <key-cert.pub>.

The server needs only two things: the CA public key in TrustedUserCAKeys (for users) or a HostCertificate directive (for hosts). Authentication succeeds if the signature is valid and the certificate hasn’t expired.

Note

A certificate does not replace the key. The key is still required — the CA signs it. The certificate adds metadata: TTL, principals, extensions, serial number.

Generating CA keys

ssh-keygen -t ed25519 -f /etc/ssh/ca_user -C "CA for user certificates"
ssh-keygen -t ed25519 -f /etc/ssh/ca_host -C "CA for host certificates"
FlagPurpose
-t ed25519Key type; ed25519 recommended per RFC 8709
-fOutput file path
-CComment; use for CA identification

Store CA private keys securely — ideally on a dedicated build machine or in an HSM. The CA private key never goes to target servers.

Signing user certificates: one-liner

ssh-keygen -s /etc/ssh/ca_user \
  -I "john@devops" \
  -n ubuntu,deploy \
  -V +52w \
  -z 1 \
  ~/.ssh/id_ed25519.pub
FlagPurpose
-s ca_privateCA private key for signing
-I identifierString logged during authentication
-n principalsComma-separated list of authorized identities
-V +52wValidity: 52 weeks from now
-z serialSerial number; useful for audit trails

The result is id_ed25519-cert.pub next to your private key. The user keeps both files. Signing another employee’s key takes the same command with a different key.

Tip

Duration suffixes: h (hours), d (days), w (weeks). +1d is one day, -1d means yesterday — already expired.

Signing host certificates

Generate host keys on each server if they don’t exist:

ssh-keygen -t ed25519 -f /etc/ssh/ssh_host_ed25519_key -N ""

Sign on the CA machine:

ssh-keygen -s /etc/ssh/ca_host \
  -I "prod-web-01" \
  -h \
  -n prod-web-01,10.0.1.5 \
  -V +52w \
  /etc/ssh/ssh_host_ed25519_key.pub

The -h flag marks this as a host certificate. The -n value lists the hostname and IP the client will verify on connect.

Place the certificate next to the host key:

cp /tmp/ssh_host_ed25519_key-cert.pub /etc/ssh/ssh_host_ed25519_key-cert.pub

sshd_config: CertFile, TrustedUserCAKeys, HostCertificate

Configuration on the target server:

# /etc/ssh/sshd_config.d/certs.conf

# Enable certificate authentication
PubkeyAuthentication yes

# CA public key trusted for user authentication
TrustedUserCAKeys /etc/ssh/ca_user.pub

# Host certificate path
HostCertificate /etc/ssh/ssh_host_ed25519_key-cert.pub

# Optional: per-user principal files
# AuthorizedPrincipalsFile /etc/ssh/%u.principals

Validate and reload after changes:

sshd -t && systemctl reload sshd
Warning

TrustedUserCAKeys expects the CA public key, not a certificate. Certificates are only needed for host keys.

Principals and source restrictions

You can restrict a certificate to specific source IPs at signing time:

ssh-keygen -s /etc/ssh/ca_user \
  -I "jenkins@ci" \
  -n deploy \
  -O source-address=10.8.0.0/16 \
  ~/.ssh/id_ed25519.pub

Or enforce from via AuthorizedPrincipalsFile on the server:

# /etc/ssh/deploy.principals
deploy from="10.8.0.0/16"

A certificate can carry multiple principals. sshd checks if any principal matches those listed in AuthorizedPrincipalsFile. This lets you issue a cert with ubuntu,deploy,admin and grant access through different principals on different hosts.

Time-to-live and rotation

Set TTL at signing. Practical guidelines:

RoleRecommended TTLRationale
CI/CD, automation24–72 hoursPipeline credentials are short-lived
Developers1–6 monthsBalance between security and convenience
Hosts12 monthsHost certificates bind to hostname/IP

Rotation means signing a new certificate with a fresh serial. The old one becomes invalid after expiry. No centralized revocation list needed — an expired certificate simply fails validation.

Fingerprints: the host-key problem

Traditional known_hosts stores host key fingerprints. Host certificates break this: the fingerprint in known_hosts won’t match the certificate. Two solutions:

Use ssh_known_hosts with ssh-keyscan:

ssh-keyscan -t ed25519 prod-web-01 >> /etc/ssh/ssh_known_hosts

Or explicitly prefer certificate algorithms in your client config:

Host prod-web-01
    HostKeyAlgorithms ssh-ed25519-cert-v01@openssh.com
    UpdateHostKeys ask

On first connect, ssh will prompt to accept the certificate, add it to known_hosts, and won’t ask again.

Troubleshooting

Debug in order:

# Check sshd accepted the config
sshd -t

# Watch authentication logs
journalctl -u sshd -f

# Verbose client output
ssh -vvv user@host

In -vvv output, look for Certificate lines and Authentications that can continue. no matching identity means the certificate isn’t next to the key. certificate refused indicates the CA isn’t trusted or the principal didn’t match.

Inspect certificate contents:

ssh-keygen -Lf ~/.ssh/id_ed25519-cert.pub

Output shows Valid: from ... to ..., Principals:, Serial:. First place to check when something breaks.

Note

too many authentication failures with a valid certificate usually means ssh is cycling through all keys before reaching the cert. Add -o PubkeyAuthentication=no before explicit key specification, or remove extraneous keys from .ssh.

SSH certificates eliminate key distribution across hosts. One CA, signed keys with TTL, principals for access control. If you manage more than five servers, this isn’t optional — it’s infrastructure.

57 - systemd-timer: scheduling instead of cron

cron works, but its logs are flat text files with no structure, and service dependencies require workarounds like embedding Requires= logic inside shell scripts. systemd-timer fixes this: unified management interface, logs in journald, dependencies through the familiar After= and WantedBy= directives — all in one stack.

Structure: .service and .timer

A timer is a separate unit that triggers a .service. The separation is intentional: the service can be invoked manually or on a schedule.

/etc/systemd/system/
├── backup.service
└── backup.timer

backup.service is a regular unit, startable via systemctl start backup.service.

backup.timer is the trigger. Without it, the service will not run on a schedule.

# /etc/systemd/system/backup.service
[Unit]
Description=Backup to storage
After=network.target

[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
# /etc/systemd/system/backup.timer
[Unit]
Description=Run backup daily

[Timer]
OnCalendar=daily
Persistent=true

[Install]
WantedBy=timers.target

Persistent=true — if the machine was off when the timer fired, it catches up after boot.

Calendar Timers (OnCalendar)

OnCalendar= format is the closest analog to cron expressions, but with different syntax.

ExampleFires
OnCalendar=dailyEvery day at 00:00
OnCalendar=*-*-01 03:00First day of each month at 03:00
OnCalendar=*-*-* 02:00Every day at 02:00
OnCalendar=09..17:00Every hour from 09:00 to 17:00
OnCalendar=*:0/15Every 15 minutes
OnCalendar=Mon..Fri 09:30Weekdays at 09:30

Multiple values can be specified comma-separated:

[Timer]
OnCalendar=09:00,12:00,18:00

To validate syntax before applying:

systemd-analyze calendar '*-*-01 03:00'

Output shows the next firing date. Useful when “every second Tuesday” becomes *-*-1..31 03:00 — verify before deployment.

Monotonic Timers

Monotonic timers count from an event, not the clock.

DirectiveFires
OnBootSec=5min5 minutes after boot
OnStartupSec=10min10 minutes after systemd manager starts
OnUnitActiveSec=1h1 hour after the service last ran
OnUnitInactiveSec=1d1 day after the service stopped

OnBootSec and OnStartupSec are similar, but OnBootSec resets on each boot while OnStartupSec counts from when the systemd manager started. In practice, the difference shows in containers and during live migration.

Combination OnBootSec + OnUnitActiveSec implements “every hour, but not before boot”:

[Timer]
OnBootSec=10min
OnUnitActiveSec=1h

Verification: systemctl list-timers

After enabling and starting:

sudo systemctl enable --now backup.timer
sudo systemctl list-timers --all
NEXT                        LEFT     LAST                        PASSED  UNIT            ACTIVATES
Mon 2024-11-18 00:00:00 MSK  6h left  Sun 2024-11-17 00:00:08 MSK 18h ago backup.timer   backup.service

Without --all shows only active timers. NEXT is when it fires, LEFT is how long until then.

If the timer does not appear in the list — check status:

systemctl status backup.timer
systemctl status backup.service
journalctl -u backup.service -n 50

Common cause of a silent timer is a missing WantedBy=timers.target.

User-Level Timers (systemd –user)

Not every task needs root. Deployment scripts, home-directory cache cleanup, periodic git fetch — better under a user.

Units go in ~/.config/systemd/user/:

mkdir -p ~/.config/systemd/user
# ~/.config/systemd/user/sync.service
[Unit]
Description=Git sync

[Service]
Type=oneshot
WorkingDirectory=%h/projects/monorepo
ExecStart=/usr/bin/git fetch --all

[Install]
WantedBy=default.target
# ~/.config/systemd/user/sync.timer
[Unit]
Description=Git sync every hour

[Timer]
OnBootSec=2min
OnUnitActiveSec=1h

[Install]
WantedBy=timers.target

Activation:

systemctl --user enable --now sync.timer

User timers need linger if they should run without login:

sudo loginctl enable-linger username

Common Mistakes

Missing OnCalendar or monotonic timer. Without a [Timer] directive, the timer will never fire. systemctl start backup.timer starts the unit, but without a schedule it just sits there.

Forgot WantedBy. Without [Install], the unit does not persist across reboots. systemctl enable backup.timer completes without errors, but it will not appear in list-timers.

ExecStart in .timer. A timer only triggers its linked service. Placing ExecStart in .timer causes systemd to ignore it and pull from .service.

Persistent with no prior run. On first activation, Persistent=true does nothing — there is no “last run” in history. The service fires only at the next scheduled time.

Logging: cron vs timer in journald

Cron sends output to syslog or cron.log, structure is a text line with timestamp. Parsing requires grep or awk.

systemd-timer writes service stdout/stderr directly to journald:

journalctl -u backup.service -f

Filtering by time, unit, severity — standard journalctl flags:

FlagEffect
-u backup.serviceThis unit only
-n 100Last 100 lines
-fFollow in real time
--since "1 hour ago"Time range filter
-p errErrors only

Actual firing time is recorded in metadata. Build execution history without parsing text logs.

journalctl -u backup.timer -o short-iso -n 20

Output includes real start time, simplifying debugging of missed triggers.

Summary

Switching from cron to systemd-timer pays off when the workload already lives in a systemd environment. Unified management, dependencies via After=, logs in journald — gains are tangible. For a crontab one-liner, systemd-timer is overkill, but for scripts with dependencies, logging, and boot persistence — a mature tool.

58 - Cron setup: a practical walkthrough

Cron is the standard job scheduler in Linux, found in every infrastructure. Jobs pile up, logs accumulate, and the environment breaks things. Let’s walk through how to configure cron reliably and where it falls short.

When to Use Cron and When to Avoid It

Cron fits simple recurring tasks: backups, log rotation, temp file cleanup, periodic notifications. It’s a daemon that sleeps between runs — no resource consumption.

Avoid cron for:

  • tasks requiring millisecond precision (cron fires at the minute level)
  • tasks with hard dependencies on other services (use systemd units with After=)
  • long-running processes that could overlap (you need locking)

Anatomy of a Crontab Entry

Format:

* * * * * command
- - - - -
| | | | |
| | | | └── day of week (0-7, 0 and 7 = Sunday)
| | | └──── month (1-12)
| | └────── day of month (1-31)
| └──────── hour (0-23)
└────────── minute (0-59)

Special values:

@reboot   — at boot
@yearly   — once a year (0 0 1 1 *)
@monthly  — once a month (0 0 1 * *)
@weekly   — once a week (0 0 * * 0)
@daily    — once a day (0 0 * * *)
@hourly   — once an hour (0 * * * *)

Examples:

0 3 * * *          /opt/scripts/backup.sh       # daily at 03:00
15,45 * * * *      /usr/local/bin/check.sh      # every 15 and 45 minutes
0 */4 * * *        /opt/metrics/collect.sh       # every 4 hours
0 9-17 * * 1-5     /opt/reports/daily.sh         # hourly during business hours on weekdays

crontab -e and crontab -l: Common Commands

crontab -l              # show current user crontab
crontab -e              # edit crontab (opens in EDITOR)
crontab -r              # remove crontab (no confirmation!)
crontab -l -u username  # view another user's crontab (from root)
crontab filename        # load jobs from file
Note

Default editor is vi. Change it: export EDITOR=nano.

Where System Schedules Live

Beyond user crontabs, there are system files:

/etc/crontab                # system crontab (format differs — includes username field)
/etc/cron.d/                # drop-in directory
/etc/cron.daily/            # daily jobs (run-parts)
/etc/cron.hourly/           # hourly jobs
/etc/cron.monthly/          # monthly jobs
/etc/cron.weekly/           # weekly jobs
/var/spool/cron/crontabs/   # user crontab files

Lines in /etc/crontab and /etc/cron.d/* include a username field:

SHELL=/bin/bash
PATH=/usr/local/sbin:/usr/local/bin:/sbin:/bin:/usr/sbin:/usr/bin
MAILTO=root

* * * * * root /opt/scripts/check.sh
Warning

Don’t add jobs directly to /etc/crontab. Use /etc/cron.d/ instead. It’s safer across package updates.

Environment and PATH: Why Jobs Break in Cron

Cron runs commands with a minimal environment. Common failure:

# works in terminal
/opt/scripts/backup.sh

# in cron — "command not found"

Reason: cron sets PATH to only /usr/bin:/bin. Solutions:

Specify full paths explicitly:

0 3 * * * /usr/bin/python3 /opt/scripts/backup.py

Set PATH in crontab:

PATH=/usr/local/bin:/usr/bin:/bin:/opt/scripts
0 3 * * * backup.sh

Use a wrapper script:

#!/bin/bash
# /opt/scripts/run_backup.sh
source /etc/profile
cd /opt/project || exit 1
./backup.sh
Tip

Always verify variables: env | sort in terminal vs. job * * * * * env | sort > /tmp/cron_env.txt.

Output Redirection and Log Rotation

By default, cron emails stdout and stderr to the user. Set MAILTO="" to disable.

# send output to file
0 3 * * * /opt/scripts/backup.sh >> /var/log/backup.log 2>&1

# add timestamp to log (readability)
0 3 * * * /opt/scripts/backup.sh >> /var/log/backup.log 2>&1

# rotate old logs
0 3 * * * /opt/scripts/backup.sh >> /var/log/backup.log 2>&1 && \
  find /var/log -name "backup.log*" -mtime +7 -delete

Or use logger for syslog:

0 3 * * * /opt/scripts/backup.sh 2>&1 | logger -t backup

Timezones and TZ

Cron uses the system timezone. To override:

# option 1: variable in crontab
TZ=Europe/Moscow
0 9 * * * /opt/scripts/report.sh

# option 2: wrapper
0 9 * * * TZ=Europe/Moscow /opt/scripts/report.sh
Warning

TZ only affects the schedule. Inside the script, use $TZ explicitly: date +%Z shows system timezone.

Check schedule in UTC:

crontab -l | while read line; do
  if [[ ! "$line" =~ ^# ]] && [[ ! -z "$line" ]]; then
    echo "$line" | awk '{print $1":"$2" UTC  |  "$5" "$6" "$7" "$8" "$9" "$10}'
  fi
done

Common Pitfalls and Pre-Deployment Checks

1. Percent sign in commands

% in crontab means newline. Escape it:

# Wrong:
0 3 * * * /opt/scripts/report.sh "Report for $(date +%Y-%m-%d)"

# Right:
0 3 * * * /opt/scripts/report.sh "Report for $(date +\%Y-\%m-\%d)"

2. Overlapping jobs

If a script may run longer than the interval, use locking:

# /opt/scripts/long-task.sh
LOCKFILE=/var/run/long-task.lock

if [ -f "$LOCKFILE" ]; then
  echo "Already running" >&2
  exit 1
fi

trap "rm -f $LOCKFILE" EXIT
touch "$LOCKFILE"

# main logic
sleep 30

3. Pre-deployment checks

# show next run time for each job
for f in /etc/cron.d/*; do
  if [ -f "$f" ]; then
    echo "=== $f ==="
    head -1 "$f"
    # nearest run time
    next=$(echo "0 3 * * *" | sed 's/\*/0/g' | xargs -I{} date -d "{}" '+%Y-%m-%d %H:%M')
    echo "Next: $next"
  fi
done

# dry run
cat /etc/cron.d/my-task | grep -v "^#" | grep -v "^$" | while read schedule cmd; do
  echo "Would run: $cmd"
done

4. Syntax and validation

# check crontab format
crontab -l | grep -v "^#" | grep -v "^$" | awk '{print $1" "$2" "$3" "$4" "$5}' | \
  while read min hour dom mon dow; do
    # basic check
    echo "$min $hour $dom $mon $dow"
  done

Ansible and Cron: Idempotent Job Setup

- name: Add backup cron job
  community.general.cron:
    name: "backup database"
    minute: "0"
    hour: "3"
    job: "/opt/scripts/backup.sh >> /var/log/backup.log 2>&1"
    user: "root"
    state: present
    cron_file: "backup"
# Remove job
- name: Remove old cron job
  community.general.cron:
    name: "obsolete task"
    state: absent
    user: "root"
# Drop a file into /etc/cron.d/
- name: Deploy cron file
  ansible.builtin.copy:
    src: files/my-cron-job
    dest: /etc/cron.d/my-cron-job
    owner: root
    group: root
    mode: "0644"
  notify: restart cron

systemd Timers as an Alternative to Cron

Timers are more precise, support service dependencies, randomiseddelaysec, and calendar specs.

# /etc/systemd/system/backup.service
[Unit]
Description=Backup database

[Service]
Type=oneshot
ExecStart=/opt/scripts/backup.sh

[Install]
WantedBy=multi-user.target
# /etc/systemd/system/backup.timer
[Unit]
Description=Run backup daily at 3am

[Timer]
OnCalendar=*-*-* 03:00:00
Persistent=true

[Install]
WantedBy=timers.target
systemctl daemon-reload
systemctl enable --now backup.timer
systemctl list-timers --all | grep backup

Timer benefits:

  • dependencies (After=network.target)
  • logging via journal (journalctl -u backup.service)
  • randomiseddelaysec for staggering (avoid thundering herd)
  • one-shot and monotonic timers

Cron stays simpler for basic tasks. systemd timers are the right choice for services with dependencies and monitoring through journald.

59 - GNU Screen: sessions that survive an SSH drop

A long apt upgrade, a migration, a build — then the laptop sleeps. SSH dies, the process gets SIGHUP and dies with it. GNU Screen keeps the terminal on the server: disconnect, come back, the work is still there.

It is not a nohup replacement and not “another SSH”. It is a multiplexer: named sessions, several windows inside one, a shared session for two people. On older RHEL/Debian boxes screen is often already installed when tmux is not.

The minimal loop

On the server:

screen -S deploy

That is a normal shell. To leave without killing processes: Ctrl-a, release, then d. The session stays Detached.

List and resume:

screen -ls
screen -r deploy

If it is still Attached on another terminal (left it open at the office):

screen -d -r deploy

Detach there first, attach here.

Note

The prefix is Ctrl-a. Do not hold both keys together: Ctrl-a, release, then the letter. Ctrl-a a sends a real Ctrl-a into the program inside (needed by emacs, rarely by bash).

Install

# Debian / Ubuntu
sudo apt install screen

# RHEL / Alma / Rocky
sudo dnf install screen

# Alpine
sudo apk add screen

Check with screen -v. Vertical splits (Ctrl-a |) exist since GNU Screen 4.1; very old boxes will not have them.

Command-line flags

FlagWhat it doesExample
-S namecreate or address a session by namescreen -S logs
-ls / -listlist sessionsscreen -ls
-r [name]attach a detached sessionscreen -r logs
-d -r namedetach elsewhere, attach herescreen -d -r logs
-Rattach if it exists, otherwise createscreen -R logs
-x [name]attach without detaching the other sidescreen -x logs
-dmS name commandstart in the background, already Detachedscreen -dmS backup /opt/backup.sh
-Lwrite screenlog.N in the current directoryscreen -L -S migrate
-Logfile filelog path (with -L)screen -L -Logfile /tmp/migrate.log -S migrate
-X commandsend a command into a live sessionscreen -S logs -X stuff 'tail -f /var/log/nginx/error.log\n'
-wipedrop dead sockets from the listscreen -wipe

The -S name is what you look for in screen -ls. Without a name you get something like 12345.pts-0.hostname — awkward to recall.

Typical background job that must survive an SSH drop:

screen -dmS pg-dump pg_dump -Fc -f /backup/app.dump app
screen -ls
# later
screen -r pg-dump

When the command exits, the session usually disappears. Keep a shell afterwards:

screen -dmS build bash -lc 'make -j"$(nproc)"; exec bash'

Windows inside a session

One session, several windows: build, logs, a second shell. Switch without a new SSH connection.

KeyAction
Ctrl-a cnew window
Ctrl-a n / Ctrl-a pnext / previous
Ctrl-a 0 … Ctrl-a 9jump by number
Ctrl-a "window list, pick with arrows
Ctrl-a 'jump by number or name
Ctrl-a Arename the current window
Ctrl-a Ctrl-alast active window
Ctrl-a kclose the window (asks to confirm)
Ctrl-a \kill every window and quit screen

Window titles help once you have more than two: Ctrl-a A → nginx-log.

Session, copy, split

KeyAction
Ctrl-a ddetach; processes keep running
Ctrl-a D Dpower detach (other displays too)
Ctrl-a ?key help
Ctrl-a :screen command line (quit, sessionname, …)
Ctrl-a asend Ctrl-a into the window
Ctrl-a [copy mode / scrollback
Ctrl-a ]paste screen’s paste buffer
Ctrl-a Escsame as Ctrl-a [
Ctrl-a Ssplit horizontally (region above/below)
Ctrl-a |split vertically
Ctrl-a Tabfocus the other region
Ctrl-a Xclose this region (window stays)
Ctrl-a Qkeep only this region
Ctrl-a Htoggle screenlog.N
Ctrl-a Mmonitor the window for activity (bell)
Ctrl-a xlock the session (user password)

Scrollback: Ctrl-a [, then arrows or PageUp / PageDown. Select: Space to start, arrows, Enter to copy into screen’s buffer. Leave the mode with Esc. Paste with Ctrl-a ]. This is not the system clipboard: the buffer lives inside screen.

After a split the new region is empty until you focus it (Ctrl-a Tab) and pick a window (Ctrl-a n or Ctrl-a ").

Server-side workflows

Long deploy. Named session, log on disk, close the lid:

screen -L -Logfile ~/migrate.log -S migrate
# inside: ansible-playbook -i prod site.yml
# Ctrl-a d

In the morning: screen -r migrate, or just tail -f ~/migrate.log.

Several jobs in one SSH. Session ops, windows build, journal, sql:

screen -S ops
# Ctrl-a c  — another window
# Ctrl-a A  — name it

Paired view. Same user, both attached:

# first
screen -S incident
# second, without kicking the first off
screen -x incident

Both see the same terminal. Faster than “paste me the output” during an incident.

Push a command into a session already running — no interactive attach:

screen -S ops -X screen bash
screen -S ops -X stuff 'systemctl status nginx\n'

-X screen opens a window; stuff types into it. \n is Enter.

A short .screenrc

Default scrollback is tiny, the startup banner gets in the way, and window names are easy to miss. In ~/.screenrc:

startup_message off
vbell off
defscrollback 20000
shell -$SHELL

hardstatus alwayslastline
hardstatus string "%{= kw}%-w%{= BW}%n %t%{= kw}%+w %= %H %l %Y-%m-%d %c"

defscrollback is how many lines Ctrl-a [ can walk. hardstatus is the bar at the bottom: windows, host, load, time. The file is read when a session is created; live sessions pick it up only after you recreate them.

You do not have to touch /etc/screenrc: the user file extends it.

Common failures

SymptomWhyWhat to do
There is no screen to be resumedwrong name, or the session already diedscreen -ls, then the exact name
Attached and -r refusessession still on another ptyscreen -d -r name
Session vanished after a commandthe only window’s process exitedwrap with bash -lc '…; exec bash'
Ctrl-a “eats” emacs/tmux insidescreen’s prefix winsCtrl-a a for a literal; or another escape in .screenrc: escape ^Bb
No vertical splitScreen < 4.1horizontal Ctrl-a S, or a newer package
No log fileno -L and nobody pressed Ctrl-a Henable it explicitly, check the session cwd

Sockets live in /run/screen/S-$USER/ or ~/.screen/. Another user’s session is not something you attach to by accident: the directory permissions block it.

When screen, when not

A single non-interactive background command is enough for systemd-run --user, tmux, or even nohup. A standing workspace on a bastion, a deploy you start and then close the lid on, shared log tailing — screen covers that without extra dependencies.

tmux is nicer for splits and config. Screen wins when it is already there on the box you just SSHed into. A named session and Ctrl-a d are the reason it stays in muscle memory.

60 - Kafka: Cluster Health Check

A Kafka cluster in KRaft mode doesn’t forgive neglect until the first incident. Health checks need to be regular and quick — no graphs or dashboards, just the terminal. Here’s the command set that covers the typical checklist: processes, quorum, partition leaders, ISR, and a quick status report in one shot.

Checking KRaft Processes

KRaft mode has no separate ZooKeeper — the controller role is either co-located with the broker or isolated on dedicated nodes. First, verify the JVM processes are alive and see which mode each node started in.

ps -ef | grep -E 'kafka.Kafka|QuorumControllerMain' | grep -v grep

For a managed cluster you expect three QuorumControllerMain processes on dedicated controllers and N kafka.Kafka processes on brokers. Mixed nodes (combined mode) run both roles in a single process — normal for smaller installations.

Then check startup logs and configuration:

grep -E 'process.roles|node.id|controller.quorum.voters|listeners=' config/kraft/server.properties
Note

The process.roles parameter accepts broker, controller, or broker,controller. If empty, the cluster is still in legacy mode with ZooKeeper — adapt the commands below.

Verify all nodes agreed on the quorum:

/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server localhost:9092 describe --status

In the output look for leaderId, votedLeaders, and quorum size 2/3 (for three controllers) or N/N for a fully stabilized cluster.

Broker and Controller Status

Next, determine which brokers are actually responding and which dropped from the registry. Use kafka-broker-api-versions.sh — it returns supported API versions and simultaneously shows whether TCP connectivity reaches the broker.

for h in kafka1 kafka3 kafka5; do
  echo "=== $h ==="
  /opt/kafka/bin/kafka-broker-api-versions.sh \
    --bootstrap-server $h:9092 2>&1 | head -n 3
done

If a node is unreachable you’ll see a timeout. This is the fastest way to distinguish “broker stuck in JVM” from “network issue.”

Full list of registered brokers and their state:

/opt/kafka/bin/kafka-broker-api-versions.sh \
  --bootstrap-server kafka1:9092 | head -n 50

The command shows a JSON-like listing, but for a tabular report kafka-metadata-quorum.sh is more convenient:

/opt/kafka/bin/kafka-metadata-quorum.sh \
  --bootstrap-server kafka1:9092 describe --replicas

The LEADER column shows the current quorum leader, REPLICAS shows all active nodes. Unresponsive controllers drop from the list.

Tip

The lastCaughtUpTime field in describe --status shows how far a controller lags behind the leader. 0 or fresh now — healthy. Lag in minutes — reason to check GC and network latency.

Topics and Partition Leaders

The cluster can be alive but without partition leaders — producers won’t write, consumers won’t read. So the next step is topics and their leaders.

List all topics with partition and replica counts:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal

The output table contains Leader, Replicas, Isr. If Leader equals -1, the partition has no active leader — this is an emergency, producers will receive NotLeaderForPartitionException.

To get only “bad” partitions in one pipeline:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '$5 == -1 || $5 == "none" {print}'

To view leaders for a specific topic:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --topic orders.events

ISR and Unavailable Replicas

ISR (in-sync replicas) determines write reliability. Recommended setting is min.insync.replicas >= 2 for critical topics. Check for out-of-sync:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --under-replicated-partitions

The command returns partitions where Isr is smaller than Replicas. Empty output — good. Any lines in the list — incident.

Full picture across all partitions with problem filtering:

/opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server kafka1:9092 \
  --describe --exclude-internal \
  | awk '{
    replicas=$5; isr=$7;
    # Replicas and Isr are comma-separated id lists; length by commas
    rep_n = split(replicas, a, ",");
    isr_n = split(isr, b, ",");
    if (rep_n != isr_n) print "UNSYNC:", $0;
    if (a[1] == "-1") print "NO_LEADER:", $0;
  }'
Warning

kafka-topics.sh --describe on a topic with thousands of partitions outputs many lines and loads the controller. In production run with --partitions N or filter with awk, otherwise the health check itself becomes a problem.

Additionally — state of a specific replica on the broker side. If you suspect one of the disks is lagging:

/opt/kafka/bin/kafka-log-dirs.sh \
  --bootstrap-server kafka1:9092 \
  --describe --broker-list 1,3,5 \
  | jq '.[] | .logDirs[] | {broker: .broker, dir: .dir, partitions: (.partitions | length)}'

In the output check the partition.error field — if non-empty, the replica has issues (offline log dir, disk full, fs in read-only).

Quick Diagnostics in One Command

For daily rounds it’s convenient to bundle all checks into a single script with clear exit codes. Below is a minimal status indicator in bash.

#!/usr/bin/env bash
set -u
BOOTSTRAP="${BOOTSTRAP:-kafka1:9092}"
KAFKA_BIN="${KAFKA_BIN:-/opt/kafka/bin}"
fail=0

echo "== Quorum status =="
if ! $KAFKA_BIN/kafka-metadata-quorum.sh --bootstrap-server "$BOOTSTRAP" \
    describe --status 2>&1 | grep -q 'isLeader: true'; then
  echo "WARN: quorum leader not confirmed"; fail=1
fi

echo "== Brokers reachability =="
for h in $(echo "$BOOTSTRAP" | tr ',' ' '); do
  if ! timeout 5 bash -c "echo > /dev/tcp/${h%:*}/${h##*:}"; then
    echo "FAIL: $h unreachable"; fail=1
  fi
done

echo "== Under-replicated partitions =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --under-replicated-partitions 2>/dev/null)
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: 0 under-replicated partitions"
fi

echo "== Partitions without leader =="
out=$($KAFKA_BIN/kafka-topics.sh --bootstrap-server "$BOOTSTRAP" \
      --describe --exclude-internal 2>/dev/null \
      | awk '$5 == -1 {print}')
if [[ -n "$out" ]]; then
  echo "$out"; fail=1
else
  echo "OK: every partition has a leader"
fi

exit $fail

Save as kafka-health.sh, make executable, and wrap in cron or a systemd timer every 60 seconds:

chmod +x kafka-health.sh
*/1 * * * * /usr/local/bin/kafka-health.sh \
  >> /var/log/kafka-health.log 2>&1
Tip

For alerting, replace echo blocks with logger -p local0.err and configure rsyslog to your SIEM. No need for Prometheus node_exporter — events fly into the common channel.

Typical errors on first run and how to interpret them:

SymptomProbable CauseAction
Connection to node -1 could not be establishedBroker not registered in cluster but process is aliveCheck advertised.listeners, node.id, network ACL
isLeader: false for all controllersQuorum lost, controllers < __.min.insync.replicasCheck controller.quorum.voters and disk state on controllers
--under-replicated-partitions shows entries for >5 minBroker lagging, slow disk or GC pausesTake jstack, check iostat -x and LogFlushRate metrics
kafka-log-dirs.sh returns LogDirOfflineDisk full or failedFree space, check FS for read-only, restart broker
Empty --describe output on a running clusterWrong bootstrap address passedVerify advertised.listeners and DNS

The five command sets above cover 90% of operational “is the cluster alive?” questions. If everything is green — no need to dig deeper; if something is red — kafka-log-dirs.sh and kafka-metadata-quorum.sh describe --status will show which direction to go.

61 - kind: Local Kubernetes in Docker

kind creates a Kubernetes cluster from Docker containers: control-plane and worker nodes are kindest/node images. Primary use cases are local development and CI. GitHub Actions has an official create-kind action that makes pipelines with K8s tests straightforward.

Compared to minikube, kind has no hypervisor dependency and natively supports multi-node topologies. Compared to k3d, it requires no Rancher and talks to CRI/containerd directly.

Installation

macOS, Linux, and Windows (via WSL2) — download the binary from GitHub Releases.

# macOS
brew install kind

# Linux
curl -Lo /usr/local/bin/kind https://kind.sigs.k8s.io/dl/v0.22.0/kind-linux-amd64
chmod +x /usr/local/bin/kind

# Verify
kind version
# kind v0.22.0 go1.21.8 linux/amd64

Docker is required (or Podman with kind use docker driver). Make sure Docker has at least 4 GB of memory allocated for all containers.

First Cluster in One Command

kind create cluster
# Creating cluster "kind" ...
# ✓ Ensuring node image (kindest/node:v1.29.0) ✓
# ✓ Preparing nodes ✓
# ✓ Writing configuration ✓
# ✓ Starting control-plane ✓
# ✓ Installing CNI ✓
# ✓ Installing StorageClass ✓
# ✓ Waiting for node readiness ✓
# Successfully created cluster "kind"!

kind created a cluster named kind and wrote kubeconfig to ~/.kube/config. Verify:

kubectl get nodes
# NAME                 STATUS   ROLES           AGE   VERSION
# kind-control-plane   Ready    control-plane   2m    v1.29.0

kubectl get pods -A
# NAMESPACE            NAME                                         READY
# kube-system          coredns-...                                  1/1
# local-path-storage   local-path-provisioner-...                   1/1

Done. Single-node cluster is up in a minute.

Config: Multi-node and Runtime

For a cluster with multiple worker nodes, use a YAML config. Let’s create three workers:

# kind-config.yaml
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
  extraPortMappings:
  - containerPort: 80
    hostPort: 8080
    protocol: TCP
- role: worker
- role: worker
- role: worker
kind create cluster --name multi --config kind-config.yaml

Cluster creation flags:

FlagPurposeExample
--nameCluster namekind create cluster --name prod
--configPath to YAML--config ./kind.yaml
--imageCustom node image--image kindest/node:v1.28.0
--waitReadiness timeout--wait 5m
--kubeconfigAlternative kubeconfig--kubeconfig ~/.kube/dev
Tip

Instead of --image, you can set node.internalImage in the config — useful for air-gapped environments.

Loading Images into the Cluster

kind uses a separate container runtime inside the node. Images from your local Docker daemon are not visible. To load them:

# Build the image
docker build -t myapp:v1.0 ./myapp

# Load into kind nodes
kind load docker-image myapp:v1.0 --name multi

# For specific nodes
kind load docker-image myapp:v1.1 --name multi --nodes kind-worker,kind-worker2

For CI, you often load from a tar archive:

docker save myapp:v1.0 > myapp.tar
kind load image-archive myapp.tar --name multi

After loading, the image is available in the cluster without a registry.

extraMounts and kubeadm patches

Mount a host directory into a node — useful for a local registry or fixture files:

nodes:
- role: control-plane
  extraMounts:
  - hostPath: /tmp/registry
    containerPath: /var/lib/registry

Tune kubelet or kube-proxy via kubeadm patches:

nodes:
- role: control-plane
  kubeadmConfigPatches:
  - |
    kind: InitConfiguration
    nodeRegistration:
      kubeletExtraArgs:
        node-labels: "env=test"
  - |
    kind: kube-proxy
    apiVersion: kubeproxy.config.k8s.io/v1alpha1
    mode: ipvs

If the API server port 6443 is already taken, set another in the cluster config:

networking:
  apiServerPort: 6444

kind with kubeconfig

By default, kind merges the context into ~/.kube/config. For isolation:

# Separate kubeconfig
KUBECONFIG=~/.kube/kind-config kind create cluster --name isolated

# Or export after creation
kind get kubeconfig --name multi > ./kubeconfig
export KUBECONFIG=./kubeconfig
kubectl get nodes

Managing multiple clusters:

kind get clusters
# kind
# multi
# isolated

# Delete specific one
kind delete cluster --name isolated

Cleanup

kind delete cluster --name multi
# Deleting cluster "multi" ...

Without --name, the default cluster (kind) is deleted. All Docker resources are removed with the nodes. If Docker was stopped while the cluster was running, nodes remain in NotReady on the next start. Fix by recreating the cluster.

Gotchas

containerd inside the node. kubectl talks to containerd directly. Familiar docker ps and docker exec won’t show pods — they live inside the kind node. For debugging:

docker exec -it multi-control-plane crictl ps
docker exec -it multi-control-plane crictl logs <container-id>

HostNetworking and PortMappings. For external access to pods, you need extraPortMappings in the config. Without it, hostPort does not work — pod networking in kind is isolated.

PersistentVolumes. kind creates a kind-node StorageClass backed by local-path-provisioner. Data lives on the host at /var/local-path-provisioner. For a clean environment, just delete the cluster.

cgroups v2. In newer distros (Ubuntu 22.04+, Fedora), Docker may require configuration:

docker info | grep cgroup
# Cgroup Driver: systemd
# Cgroup Version: 2

kind works with both versions, but rare issues with memory limits occur with cgroups v2 on Arch Linux. Solution — pass in the config:

kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
featureGates:
  "MemoryManager": true
runtimeConfig:
  "memorymanager.k8s.io/v1alpha1": true

Versions. kind lags behind upstream Kubernetes by 1-2 minor versions. Check the compatibility matrix in the project README before setting up a prod-like environment.

Port already allocated. kind could not bind the API server or a mapped host port. Check that 6443 is free, or change networking.apiServerPort. For Ingress, extraPortMappings must not collide with host services.

No space left on device. Docker ran out of disk. Clean unused images, then recreate the cluster:

docker system prune -a
kind delete cluster --name multi

kind is a fast way to spin up Kubernetes on a developer machine or in CI without virtualization. Main loop: kind create cluster, work, kind delete cluster. For air-gapped or multi-node scenarios — YAML config and kind load docker-image. Known limitations: no GPU support, no real network stack, performance below bare metal. For everything else — it works.

62 - Too many authentication failures: SSH ran out of tries

Received disconnect from 10.0.0.5 port 22:2: Too many authentication failures followed by Permission denied (publickey) is not a broken server, and it is not necessarily a wrong password. The client spent the attempt budget while walking agent keys and never reached the method you meant to use.

The budget is MaxAuthTries on the server (6 by default). Each public key offered is one try. Five keys in ssh-agent plus one more wrong key — the connection dies before the right key or a password prompt.

Why this happens

SSH asks the server which methods it accepts, then the client walks its own order. publickey is first by default. While the agent keeps offering keys, the server is already counting failures.

You wantedWhat actually happened
Key loginThe key is missing from authorized_keys, it is not in the agent, or it is fifth in line — the limit hits first
Password loginThe server advertised publickey, the client started the key walk, the password prompt never appeared
PAM / external passwordThe client sends password, the server wants keyboard-interactive

A file sitting in ~/.ssh is not enough: OpenSSH offers agent identities, not “the file you had in mind”. What will actually be sent:

ssh-add -l

Empty agent or the wrong key — read -vvv instead of guessing.

Client log first

Debug level 3 shows the method order and every key offered:

ssh -vvv deploy@10.0.0.5

Look for:

  • Authentications that can continue — what the server allowed;
  • Offering public key / get_agent_identities — which keys go out;
  • Too many authentication failures — attempt budget, not “wrong password”.

Typical picture: the agent hands over id_rsa, id_ed25519, a GitLab key, a bastion key — four rejects, the fifth is still wrong, the sixth is no longer accepted.

Stop offering every key

IdentitiesOnly yes stops the client from mixing in every agent identity. After that, only IdentityFile or -i is used.

Per host, not in the global /etc/ssh/ssh_config:

Host prod
    HostName 10.0.0.5
    User deploy
    IdentitiesOnly yes
    IdentityFile ~/.ssh/id_ed25519_prod

One-shot:

ssh -o IdentitiesOnly=yes -i ~/.ssh/id_ed25519_prod deploy@10.0.0.5

-vvv should then show a single Offering public key. If too many turns into Permission denied (publickey), the walk is over and the key was simply rejected: it is not in authorized_keys, or it is the wrong file.

Note

IdentitiesOnly without IdentityFile keeps the default id_ed25519 / id_rsa in the home directory and does not dump the whole agent. For a stand with its own key, always name the file.

Password when the agent is full

The client will not ask for a password until publickey is exhausted. Turn keys off and put password first:

ssh -o PubkeyAuthentication=no \
    -o PreferredAuthentications=password \
    -o PasswordAuthentication=yes \
    deploy@10.0.0.5

In ~/.ssh/config:

Host prod-password
    HostName 10.0.0.5
    User deploy
    PubkeyAuthentication no
    PasswordAuthentication yes
    PreferredAuthentications password

On Ubuntu and anywhere the password goes through PAM, the server often advertises keyboard-interactive, not password. The command above then fails silently. Switch the method:

ssh -o PubkeyAuthentication=no \
    -o PreferredAuthentications=keyboard-interactive \
    deploy@10.0.0.5
Host prod-password
    HostName 10.0.0.5
    User deploy
    PubkeyAuthentication no
    PreferredAuthentications keyboard-interactive

Which method the server actually offers is the Authentications that can continue line in -vvv.

What to check on the server

You need another channel: hypervisor console, VNC, serial. Otherwise you are stuck in the same error you are fixing.

Logs. On Ubuntu 24.04 the unit is ssh (sshd is an alias). For rejected-key detail, raise verbosity:

# /etc/ssh/sshd_config or a drop-in under /etc/ssh/sshd_config.d/
LogLevel VERBOSE
sudo systemctl restart ssh
sudo journalctl -u ssh -e
# if rsyslog is installed:
sudo tail -f /var/log/auth.log

The log shows which key was rejected and whether MaxAuthTries was reached.

Permissions. With StrictModes yes (the default) sshd silently ignores files that are too open:

chmod 700 ~/.ssh
chmod 600 ~/.ssh/authorized_keys
# the home directory must not be group/other-writable

The public key is one line in authorized_keys; the matching private key is the one in IdentityFile.

Attempt limit. A fat agent and many keys per host can justify raising the cap (this weakens brute-force protection):

MaxAuthTries 10

Better not to raise it: shrink the client with IdentitiesOnly and one IdentityFile per Host.

Short checklist

  1. ssh -vvv — how many keys went out and which method remained.
  2. ssh-add -l — what sits in the agent; do not offer the rest.
  3. In ~/.ssh/config for the host: IdentitiesOnly yes and an explicit IdentityFile.
  4. Need a password — PubkeyAuthentication no and the method the server printed in debug (password or keyboard-interactive).
  5. From the server console: journalctl -u ssh, ~/.ssh / authorized_keys modes, LogLevel VERBOSE if needed.

A “too many” error is almost always on the client: too many keys for one login. The server only counts to six and closes the session.

63 - Trusting a custom CA: system store, browsers, and CLI

A TLS error that says the certificate is untrusted almost never means the certificate is “broken”. The trust anchor is in the wrong store.

curl, openssl, Git, Python, and the browser are different clients. Some keep their own root lists. Updating the OS bundle on Linux will not fix Chrome or Firefox by itself.

What to import

You trust the CA root that signed the server certificate, not localhost.crt / app.example.internal itself.

Typical files:

FileRole
ca.crtroot (trust anchor) — this is what you import
server.crtservice certificate signed by the CA
server.keyservice private key, server-side only

A self-signed server certificate with no separate CA can be added as an anchor on its own. For labs and internal PKI you usually keep a CA: one root, many services.

PEM (-----BEGIN CERTIFICATE-----) or DER both work. Most OS tools expect PEM. Browsers and certutil accept DER as well.

The system store

This is what OpenSSL-based clients use: curl, wget, git, many agents. Do this first, then debug the browser.

Linux

Distros rebuild /etc/ssl/certs in different ways.

Debian, Ubuntu, and derivatives — a .crt file under /usr/local/share/ca-certificates/, then rebuild the bundle:

sudo cp ca.crt /usr/local/share/ca-certificates/internal-ca.crt
sudo update-ca-certificates

RHEL, Fedora, Alma, Rocky — drop the anchor, then extract:

sudo cp ca.crt /etc/pki/ca-trust/source/anchors/internal-ca.crt
sudo update-ca-trust extract

Alpine — package ca-certificates, same update-ca-certificates, directory /usr/local/share/ca-certificates/.

Then:

curl -I https://app.example.internal
openssl s_client -connect app.example.internal:443 -servername app.example.internal </dev/null 2>/dev/null | openssl x509 -noout -issuer -subject

If curl is clean and the browser is not, the OS trust is fine. The browser is looking elsewhere.

macOS

Use the System keychain, not the login one, if every process on the machine should trust the CA:

sudo security add-trusted-cert -d -r trustRoot \
  -k /Library/Keychains/System.keychain ca.crt

For a single user, the login keychain is enough: no sudo, ~/Library/Keychains/login.keychain-db. GUI: Keychain Access → System → Certificates → import → Always Trust for SSL.

Windows

The Local Machine Root store:

certutil -addstore -f Root ca.crt

Same via certmgr.msc (current user) or certlm.msc (computer): Trusted Root Certification Authorities → import.

Note

In a container or CI job the bundle lives inside the image. Installing a CA on the host does not help curl in the container. Copy ca.crt into the image and run the same update-ca-certificates / update-ca-trust at build time.

Why the browser still complains

Chrome takes public roots from the Chrome Root Store. Extra CAs come from the OS — but not on every platform.

ClientWhere your CA has to live
curl / OpenSSLOS bundle
Chrome / Edge on Windows and macOSOS store (system install is often enough)
Chrome / Chromium on Linuxper-user NSS database, not /etc/ssl/certs
Firefoxper-profile store; OS roots only via policy, and not on Linux

“Installed the CA on the system” and “opened the site in a browser” are separate steps.

Chrome, Chromium, Edge

GUI is the same idea on every OS: Settings → Privacy and security → Security → manage certificates. Authorities tab → import ca.crt. Trust for website identification is enough.

Direct URL: chrome://settings/certificates (Edge: edge://settings/certificates).

CLI on Linux. Chromium/Chrome use an NSS Shared DB. Since M146 the default is ~/.local/share/pki/nssdb; if ~/.pki/nssdb already exists, that one wins.

certutil comes from libnss3-tools (Debian family) or nss-tools (RHEL/Fedora/Alpine).

NSSDB="${HOME}/.pki/nssdb"
[ -d "$HOME/.local/share/pki/nssdb" ] && NSSDB="${HOME}/.local/share/pki/nssdb"

mkdir -p "$NSSDB"
certutil -d "sql:${NSSDB}" -N --empty-password 2>/dev/null || true
certutil -d "sql:${NSSDB}" -A -t "C,," -n "Internal CA" -i ./ca.crt
certutil -d "sql:${NSSDB}" -L

The three -t fields are SSL, email, and code signing. C means trusted CA. Need client-certificate issuance as well — "CT,,". A self-signed server cert with no CA — "P,,".

Fully quit the browser and open it again. Background Chrome processes will not pick up the new cert.

On Windows and macOS a separate NSS import is usually unnecessary: the OS store is enough.

Firefox

Its own database per profile. Snap/Flatpak/vendor builds use different paths — importing into a “normal” profile does not reach them.

GUI: Settings → Privacy & Security → Certificates → View Certificates → Authorities → Import. Enable trust for identifying websites.

CLI — same certutil, profile directory, not ~/.pki/nssdb.

Linux: ~/.mozilla/firefox/<id>.default-release/ macOS: ~/Library/Application Support/Firefox/Profiles/<id>.default-release/ Windows: %APPDATA%\Mozilla\Firefox\Profiles\<id>.default-release\

PROFILE=$(find ~/.mozilla/firefox -maxdepth 1 -type d -name '*.default-release' | head -n 1)
certutil -d "$PROFILE" -A -t "C,," -n "Internal CA" -i ./ca.crt

Substitute the right profile path if you have more than one.

Policy: one file for every profile

For a fleet, policies.json beats hand imports.

Policy location:

OSWhere to put it
Linux/etc/firefox/policies/policies.json or distribution/policies.json in the install dir
macOSFirefox.app/Contents/Resources/distribution/policies.json
Windowsdistribution\policies.json next to firefox.exe, or GPO

ImportEnterpriseRoots makes Firefox trust roots from the OS store. It works on Windows and macOS. Mozilla does not implement it on Linux: there is no OS certificate store in the same sense.

On Linux (and as an explicit list on any OS) use Certificates.Install: a filename or an absolute path. A bare filename is searched here:

  • Linux: /usr/lib/mozilla/certificates, /usr/lib64/mozilla/certificates, ~/.mozilla/certificates
  • macOS: /Library/Application Support/Mozilla/Certificates, ~/Library/Application Support/Mozilla/Certificates
  • Windows: %LOCALAPPDATA%\Mozilla\Certificates, %APPDATA%\Mozilla\Certificates
{
  "policies": {
    "Certificates": {
      "Install": ["internal-ca.crt"]
    }
  }
}

PEM and DER both work. Restart Firefox after changing the policy.

security.enterprise_roots.enabled in about:config is the same as ImportEnterpriseRoots: Windows and macOS, not Linux. On Linux, Certificates.Install or the p11-kit-trust PKCS#11 module is the equivalent.

CLIs and runtimes that still ignore the OS

Even after the system CA is in place, parts of the stack ship their own bundle.

StackWhat it usesWhat to set
Python (certifi, some requests/httpx)its own Mozilla bundleSSL_CERT_FILE, REQUESTS_CA_BUNDLE, or truststore
Node.jsits own listNODE_EXTRA_CA_CERTS=/path/to/ca.crt
JavaJDK cacertskeytool -importcert -alias internal-ca -file ca.crt -keystore "$JAVA_HOME/lib/security/cacerts"
Gitoften OS OpenSSL/SchannelOS CA; otherwise http.sslCAInfo
export SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
export NODE_EXTRA_CA_CERTS=/etc/ssl/certs/internal-ca.pem

macOS and Windows use different paths for the system bundle; for Node it is simpler to point at ca.crt itself.

Short checklist

  1. Import the CA, not the service leaf.
  2. Put it in the OS store and verify with curl / openssl s_client.
  3. Chrome on Linux needs NSS as well; on Windows/macOS step 2 is often enough.
  4. Firefox needs a profile import or policies.json; ImportEnterpriseRoots is Windows/macOS only.
  5. In a container, a JDK, and Node, check that runtime’s bundle, not only the OS.

There is no single command that covers all of this. There is a predictable order: OS → browser → runtime.

64 - ssh-connection-manager: a TUI for hosts in ~/.ssh/config

When ~/.ssh/config holds dozens of stands, bastions, and jump hosts, memorizing aliases stops being fun. ssh-connection-manager is a TUI on top of ordinary OpenSSH: a host list, a filter, a connect via the system ssh, and a way to append a new block to the config.

The CLI command is ssh-connect. Repository: gitlab.com/unsorted-projects/ssh-connection-manager.

Why not another SSH client

The client already exists: the system ssh. What is missing is navigation over the config file.

The tool:

  • reads ~/.ssh/config or another file (-c);
  • lists concrete Host entries, not wildcards such as Host *;
  • follows Include (up to 16 levels);
  • on Enter suspends the TUI and runs ssh -F <config> <alias>;
  • after the session ends, shows the list again.

It is Python 3.11+ and Textual. There is no custom SSH stack: keys, ProxyJump, and the agent stay with OpenSSH.

What the table shows

Column names keep the alias and the address apart:

FieldMeaning
HostNamealias from the Host directive
ConnectPointaddress from HostName in the config (IP or DNS)
Useruser, if set
Portport, or 22 by default

The details line shows IdentityFile and ProxyJump when present. If a host has no User (and none comes from Host *), the TUI asks for a username before connecting.

The filter (/) matches alias and ConnectPoint.

Keys

KeyAction
Enterconnect
aadd a host to the config
/filter
qquit

Adding a host appends a block; existing sections are not rewritten. If the file does not exist yet, it is created with mode 0600. Wildcard characters are rejected in the alias: only explicit names appear in the list.

Example of what gets written:

Host prod
    HostName 10.0.0.5
    User deploy
    Port 22
    IdentityFile ~/.ssh/id_ed25519

Install

On Debian/Ubuntu do not install into the system Python (externally-managed-environment). Use a venv or pipx.

git clone git@gitlab.com:unsorted-projects/ssh-connection-manager.git
cd ssh-connection-manager
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

Then ssh-connect is available in the activated venv. To run it from any directory:

mkdir -p ~/.local/bin
ln -sf "$(pwd)/.venv/bin/ssh-connect" ~/.local/bin/ssh-connect

Or pipx install -e . — pipx isolates the environment and puts the binary in ~/.local/bin.

An interactive TTY is required: without a terminal, Textual cannot hand the screen over to ssh.

ssh-connect
ssh-connect -c /path/to/other/config
Note

The tool does not rewrite other config blocks and ignores Match. Only named Host entries without *?[ appear in the list.

Why this fits the job

One config file stays the source of truth. The TUI does not duplicate inventory in YAML and does not store passwords: only what already lives in OpenSSH. For a Lead DevOps that is the usual loop: bastion, prod, jump — pick a row and you are in the session.