This is the multi-page printable view of this section. .
Posts
- 1: fixing memory leaks in python services: diagnostics and dump collection
- 2: SSH key best practices
- 3: kubectl cheatsheet: essential commands for Kubernetes
- 4: Creating User and Role in Kubernetes and Binding Them via RBAC
- 5: Git Tag: Marking Releases and Bookmarks in History
- 6: Deploying with a post-receive Git Hook
- 7: ulimit and systemd LimitNOFILE — why ulimit -n inside a unit doesn't stick
- 8: coredumpctl: Finding a Binary Crash
- 9: curl --resolve and SNI: Testing Virtual Hosts Without /etc/hosts
- 10: Docker logs and journald: choosing a logging driver
- 11: scp — Secure Copy Over SSH
- 12: Sudoers: NOPASSWD Without Holes
- 13: ThinLinc: Remote Access to Linux Desktops
- 14: Chrony Instead of ntpd
- 15: fail2ban: SSH Jail Configuration
- 16: journalctl: filters and follow
- 17: nftables: Basic Rule Set
- 18: Debian — The Swiss Army Knife of Linux
- 19: pipx: Isolated Python CLI Tools Without the Mess
- 20: uv — Fast Python Package Manager
- 21: logrotate for Custom Daemons
- 22: Rsync: Backing Up a Directory Over SSH
- 23: Setting Up Your Own SSH Bastion Server
- 24: tmux on Prod After Screen
- 25: bpftrace: one-liners that replace strace in production
- 26: ip: Network Setup and Diagnostics in CLI
- 27: nslookup and drill: DNS resolution in terminal
- 28: journalctl: Filtering and Formatting systemd Logs
- 29: OpenSSL: TLS Certificate Verification and Parsing in CLI
- 30: ProxyJump and bastion hosts via ~/.ssh/config
- 31: auditd: file access and syscall logging
- 32: logrotate: automatic log rotation and archiving
- 33: sshd_config: baseline for a test stand
- 34: systemd-run: Run Services Without Unit Files
- 35: curl: HTTP Debugging in CLI
- 36: ethtool: network interface diagnostics and tuning
- 37: lsof: which processes listen on port and hold file
- 38: mc — MinIO Client S3 CLI
- 39: nftables: Modern Linux Firewall
- 40: ss: socket statistics instead of deprecated netstat
- 41: SSH Config: Wildcards and Dynamic Variable Substitution
- 42: strace: System Call Tracing for Diagnosing Hangs and Leaks
- 43: tcpdump and tshark: Packet Capture in CLI
- 44: CasaOS: Web Dashboard for Home Lab
- 45: cockpit-ufw-module: Uncomplicated Firewall in Cockpit
- 46: Creating a Custom Systemd Service
- 47: kubectl whoami and Service Account Permission Checks
- 48: ngrep: grep for Network Packets in Real Time
- 49: socat: Forwarding Unix Sockets Over TCP
- 50: SSH Escape Sequences: Reviving a Frozen Terminal
- 51: What is Self-Hosted and Why It's So Popular
- 52: cockpit-modules: web panels for day-to-day operations
- 53: dig: DNS Query Debugging in CLI
- 54: MkDocs: Project Documentation from Markdown
- 55: Squid: Internet Forwarding to Remote VM
- 56: SSH certificates instead of authorized_keys
- 57: systemd-timer: scheduling instead of cron
- 58: Cron setup: a practical walkthrough
- 59: GNU Screen: sessions that survive an SSH drop
- 60: Kafka: Cluster Health Check
- 61: kind: Local Kubernetes in Docker
- 62: Too many authentication failures: SSH ran out of tries
- 63: Trusting a custom CA: system store, browsers, and CLI
- 64: ssh-connection-manager: a TUI for hosts in ~/.ssh/config
1 - fixing memory leaks in python services: diagnostics and dump collection
Python services under load can silently consume memory until the cgroup limit triggers an OOM kill. Without systematic dump collection and introspection, root-cause analysis devolves into hypothesis spinning. Below is a practical set of commands and scripts for diagnostics, heap dump collection, and leak mitigation.
1. Memory Consumption Diagnosis Commands
Basic level — psutil. Installed in one line and works without process restart.
Current process consumption:
Extended structure in one command:
Key attributes psutil.Process.memory_full_info() quick reference table:
| Attribute | Description |
|---|---|
python | Memory allocated inside the Python interpreter |
rss | Resident Set Size — physical memory in RAM |
vms | Virtual Memory Size — virtual address space |
shared | Shared memory (shared libraries, mmap) |
text, lib, data | ELF code, library, data segments |
For tracking growth dynamics over a short interval, psutil in a loop or watch -n 1 psutil ... can be used, but in production tracemalloc, built into CPython, is more common.
tracemalloc adds a small overhead (~1–2 %). Enable it only on a staging environment or when explicit leak suspicions exist.2. Tools for Tracking Growth and Dump Collection
When RSS begins creeping upward unnoticed, deeper inspection is required. The toolset depends on debugger availability and ptrace permissions.
gdb + gcore — classic method to extract a full process dump without stopping it (provided coredump is enabled).
The resulting gcore file is binary and can be analyzed locally:
objgraph — quick answer to “who is holding this object”.
objgraph works with Python objects only. For native C extensions or ctypes, gdb or valgrind are required.tracemalloc + heapdump — Python 3.4+ native mechanism can save a heap snapshot to a file.
The dump can be opened in a visualizer, but for deep analysis gdb remains more convenient.
cap_sys_ptrace restrictions, gcore collection may require host-level access or kubectl exec with appropriate capabilities.3. Common Anti-patterns and What to Avoid
| Anti-pattern | Consequence | Recommendation |
|---|---|---|
Ignoring gc.collect() as a “magic button” | False sense of security, leak persists | Use gc.collect() to clear temporary references, not as a logical leak cure |
Relying solely on __del__ for resource release | Irregular release, circular references | Prefer contextlib.contextmanager, try/finally, or weakref |
No memory limits (cgroup/ulimit) | Sharp OOM-kill without dump, context loss | Always set memory.limit in manifests and check ulimit -v |
| Caches without TTL or unbounded growth | Rapid RSS increase | Use functools.lru_cache(maxsize=N) or external stores with expiration |
| Accumulation of objects in global lists/modules | “Death” memory in long-running processes | Periodically check list lengths, move cleanup to periodic tasks |
del obj inside __del__ with circular references. Python’s GC eventually collects them, but the order is not guaranteed, causing memory peaks between collections.4. OOM Incident Response Scenario
- Confirmation — check the cluster event:
kubectl get events -n <ns> | grep OOMordmesg | grep out of memory. Verify the process terminated with code 137. - Quick RSS glance — if the service is still alive, run one command:
- Dump collection — choose the method based on tool availability:
- If
gdbis available with rights:gdb -p <PID> -batch -ex gcore -ex quit /tmp/oom_dump.core - If a quick snapshot is needed:
python3 -c "import tracemalloc; tracemalloc.start(); ..."followed bytracemalloc.take_snapshot().statistics('traceback')
- If
- Analysis — use
objgraph.most_common_types()orgdbcommandsinfo files,btfor call stacks. Look for unexpecteddict,list, or domain class object counts. - Fix — eliminate the root cause: add
maxsizeto caches, remove circular imports, replace global lists with bounded loops, adoptweakreffor observers. - Prevention — add an alert at
rss > 80% of limitin Prometheus/Grafana and document the dump-collection command in the runbook.
Completing diagnostics at this stage either restores the service or gathers sufficient data for a tracker ticket with specific object types and call stacks. Here’s a thinking process:
- Analyze the Request:
- Topic: Fixing memory leaks in Python services: diagnostics and dump collection
- Role: Practicing Lead DevOps
- Format: IT notes for a Hugo blog (Mana theme)
- Style: Short paragraphs, concrete commands, flag tables, code with fences, no fluff, no marketing, no emojis, no author signature at end
- Allowed: GitHub admonitions (
> [!NOTE],> [!TIP],> [!WARNING]) - Prohibited: Made-up flags, URLs, versions; uncertain facts → cautious tone; no YAML/TOML front matter; no wrapping in ```; start with a lid (2-4 sentences); sections with
##; practical commands, tables if needed; 800-1600 words; end on last substantive section; don’t repeat rules, don’t write “User wants”, " - Key constraint: Write the English article as a parallel original, not a word-for-word translation. Same structure and facts as the Russian draft.
2 - SSH key best practices
SSH keys are the de facto standard for authenticating to infrastructure, but poor management turns every deployment into a potential vulnerability. This note collects proven practices: from key generation to revocation and rotation without service downtime.
Introduction
SSH keys function as long‑lived credentials, and their lifecycle directly impacts supply‑chain security. Unlike passwords, keys are often created once and forgotten, leading to an accumulation of “zombie” keys with privileges that exceed current needs. Proper generation, binding to an agent, and regular rotation minimize the attack surface and enable auditing of changes. The following sections describe concrete commands and configurations used in operations.
Key Generation
Modern OpenSSH defaults recommend the Ed25519 algorithm because of its short key length (256 bits) and high resistance to attacks. RSA still appears in legacy systems but requires increasing the length to at least 4096 bits for acceptable security.
The primary generation command is:
Flag explanations:
-t– cryptographic algorithm type (ed25519,rsa,ecdsa,sk-ed25519@openssh.comfor Security Key).-C– comment, usually an email or service name; it is stored in the key but not used for authentication.-f– path to the key file; if the directory does not exist, it will be created with correct permissions.
For fully automated scenarios (CI/CD, scripts) you can pass an empty passphrase via -N "", but understand the risk: a passwordless key can be used by any process running as the user. It is preferable to use ssh-agent with an unlock session or a key store (Keychain, ssh-agent -s).
Table: Comparison of common key types
| Algorithm | Key length | Primary use | Notes |
|---|---|---|---|
| Ed25519 | 256 bits | Modern infrastructures, deployment scripts | Recommended standard starting with OpenSSH 6.8 |
| RSA | 2048–4096 bits | Legacy systems, some vendor platforms | 4096 bits minimally acceptable; larger is overkill |
| ECDSA | 256–521 bits | Specific requirements, older devices | 521‑bit (P‑521) most secure but less compatible |
When generating keys for different services, it is convenient to split them by files (e.g., id_ed25519_work, id_ed25519_cloud) to simplify revocation and rotation without affecting other contexts.
Using ssh-agent
The SSH agent holds the decrypted key in memory of the current session, avoiding password entry on each connection. Starting the agent and adding a key:
In newer OpenSSH versions (8.2+) the agent can be started in config mode:
The ~/.ssh/config file allows automatically loading keys for specific hosts:
Flags in IdentityFile and AddKeysToAgent guarantee that the key is loaded once on the first connection to a host and remains available for subsequent sessions within the agent’s lifetime. On macOS the flag UseKeychain yes integrates the key with Keychain, locking it when the screen is locked and unlocking when unlocked.
Enable IdentitiesOnly yes in the config if the SSH client begins trying all keys during a connection attempt, which can cause delays or authentication errors on servers with a limited list of allowed keys.
Revocation and Rotation
Keys should not live forever. Rotation policy depends on the infrastructure risk profile, but minimal hygiene includes regular audits and a revocation procedure when personnel change or responsibility shifts.
Revocation procedure:
Remove the public key from the remote server:
or manually edit
~/.ssh/authorized_keysdeleting the line corresponding to the revoked key.Locally delete the key and generate a new one:
Distribute the new public key via IaC or a configuration manager (Ansible, Terraform) to avoid manual edits.
Automated rotation:
For large fleets you can use scripts that check the key creation date (comment -C or metadata inside the key) and automatically revoke keys older than N days. It is important to retain key history for at least 30–90 days before full deletion to avoid breaking services that have not yet updated their configuration.
If a key leak is detected (e.g., in public repositories or logs), act immediately: revoke the key on all hosts, generate a new one, and update the configuration. Do not wait for scheduled rotation—exposure of a key means full loss of control over the associated resources. We need to produce the English version of the article, matching the structure and facts of the Russian draft. The Russian draft sections: Introduction, Генерация ключей (Key generation), Использование ssh-agent (Using ssh-agent), Отзыв и ротация (Revocation and rotation). We need to produce English article, same structure, same commands, tables, admonitions. No front matter, no code fences around whole article. Just markdown body.
We need to translate the Russian content into English, preserving the same details, commands, tables, admonitions. Ensure we keep the same practical tone, short paragraphs, specific commands, tables, code fences with language spec. Use GitHub admonitions as in Russian: > [!TIP], > [!WARNING]. Also maybe > [!NOTE] if needed but Russian didn’t have note. We’ll keep tip and warning.
We need to produce
3 - kubectl cheatsheet: essential commands for Kubernetes
kubectl is the primary interface to a Kubernetes cluster. This cheatheet covers routine operations — from context setup to pod debugging and namespace switching. All commands are verified against current kubectl versions (1.28+).
Installation and Context Setup
Installation depends on your OS. On Linux, use the package manager or the binary directly:
Configuration lives in ~/.kube/config. A context defines the cluster, user, and default namespace.
If you work with multiple clusters, store configs in the KUBECONFIG variable separated by colons: export KUBECONFIG=~/.kube/config:~/.kube/prod-config.
Working with Pods
Basic operations:
Deployments, StatefulSets, and DaemonSets
For StatefulSets, startup order and stable identities are critical. For DaemonSets, a guaranteed instance runs on every node.
Do not run kubectl delete on a Deployment without checking kubectl get deployment first. Deleting the controller does not automatically remove pods — they will be recreated unless you use --cascade=orphan (in older versions) or delete through kubectl delete deployment.
Services, Ingress, and ConfigMap
ConfigMap and Secret for configuration:
Debugging and Diagnostics
When a pod is not working, follow this sequence:
To test network connectivity from inside the cluster:
If kubectl top returns an error, metrics-server is not installed. Installation depends on your provider: kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml.
Labels, Selectors, and Output Formatting
Labels let you group resources and select them in bulk:
Output formatting:
| Flag | Description |
|---|---|
-n, --namespace | Specify namespace |
-o, --output | Format: json, yaml, wide, custom-columns |
-l, --selector | Label selector |
-f, --filename | Configuration file (YAML/JSON) |
--show-labels | Show labels column |
-w, --watch | Streaming update mode |
--all-namespaces, -A | All namespaces |
Managing Namespaces and Switching Contexts
For convenience, create a shell alias or function:
kubectl config rename-context, kubectl config unset, and kubectl config set let you edit the config without manually editing YAML. Verify the result with kubectl config view.
kubectl is more than a CLI — it is an abstraction layer over the Kubernetes API. Knowing formatting flags and selectors cuts diagnostic time significantly. Keep this cheatheet handy and update it as new versions ship.
4 - Creating User and Role in Kubernetes and Binding Them via RBAC
In Kubernetes there are no “users” in the traditional sense — there are ServiceAccounts and certificates, bound to roles through RBAC. Without proper configuration, anyone holding a kubeconfig gets full access to the cluster. Below is the complete cycle: create a ServiceAccount, define permissions, bind them, and verify.
Creating ServiceAccount and Generating kubeconfig
Start by creating a ServiceAccount in the target namespace:
To generate a kubeconfig, extract the token from secrets and build the config file:
A ServiceAccount token is a plain bearer token. If the kubeconfig leaks, an attacker gains access with that SA’s privileges. Store the file with 600 permissions.
Defining Role and ClusterRole
A Role is namespace-scoped; a ClusterRole is cluster-scoped. The difference is critical: a Role grants no rights outside its namespace, a ClusterRole does.
Example Role allowing read and create operations on pods in the staging namespace:
Example ClusterRole for managing ingresses across the entire cluster:
An empty string "" in apiGroups refers to the core API group (v1). For apps, networking, batch — specify the corresponding groups. The full list of groups is available via kubectl api-resources.
Creating RoleBinding and ClusterRoleBinding
A RoleBinding binds a Role to a subject within a namespace. A ClusterRoleBinding binds a ClusterRole cluster-wide.
Binding a Role to a ServiceAccount in the staging namespace:
Binding a ClusterRole to the same SA (now with cluster-wide ingress rights):
| Binding | Scope | Role type | When to use |
|---|---|---|---|
| RoleBinding | Single namespace | Role or ClusterRole | Read/write in a specific ns |
| ClusterRoleBinding | Entire cluster | ClusterRole | Global rights (node, pv, dns) |
Verifying and Debugging Access Rights
After applying the YAML, confirm the SA actually received the intended permissions:
The last command checks via ClusterRoleBinding — it returns yes if the binding is correct.
If something doesn’t work, check the audit log or use kubectl auth reconcile with --dry-run=server for a preview:
Another useful trick is to find which RoleBindings are attached to a specific SA:
kubectl auth can-i checks permissions only — it does not account for NetworkPolicy or PodSecurityPolicy. If a pod fails to start, the issue may lie there.
Bottom line: RBAC in Kubernetes works as a chain — SA → Role/ClusterRole → Binding. Each link can be verified independently, which greatly simplifies debugging. Start with minimum privileges and expand as needed; never grant cluster-admin without a compelling reason.
5 - Git Tag: Marking Releases and Bookmarks in History
Git Tag: Marking Releases and Bookmarks in History
Tags are Git’s mechanism for assigning meaningful labels to specific commits. Unlike branches, tags don’t move — they’re pinned to a commit and serve as anchors for releases, versions, and critical checkpoints. Without tags, release history devolves into a hash search — and that’s a direct path to deployment errors.
Types of Tags: Lightweight vs Annotated
There are two types of tags. Lightweight is just a name attached to a commit, with no additional metadata. Annotated is a full Git object with author, date, message, and signing capability.
Always use annotated tags for releases. Lightweight tags carry no metadata and can’t be signed — you lose context during audits or rollbacks.
| Property | Lightweight | Annotated |
|---|---|---|
| Object in DB | No | Yes (tag object) |
| Message | No | Yes |
| Author/Date | No | Yes |
| GPG Signature | No | Yes |
| Creation Speed | Faster | Slightly slower |
Creating, Viewing, and Deleting Tags
Create an annotated tag:
Create a lightweight tag:
List all tags:
Show details of a specific annotated tag:
Delete a local tag:
List tags matching a pattern:
The -l flag supports glob patterns. This is faster than piping git tag through grep.
Pushing Tags to Remote
By default, git push does not send tags. This is a common reason why teammates can’t see a release tag on the remote repository.
Push a single tag:
Push all local tags:
Delete a tag on remote:
Local deletion and remote deletion are two separate operations. Forgetting the second is a typical mistake during a release rollback.
GPG-Signing Tags
Annotated tags can be signed with a GPG key. This guarantees the tag was created by a specific author and hasn’t been tampered with.
Ensure your GPG key is configured first:
Create a signed tag:
Verify the signature:
If git tag -v returns “no signature found”, the tag isn’t signed. If it says “Good signature from…”, the signature is valid. Make sure the author’s public key is available in your keyring.
Working with Tags in CI/CD Pipelines
In pipelines, tags are the primary trigger for releases. Most CI/CD systems allow filtering by branches and tags.
Example for GitHub Actions:
Example for GitLab CI:
Useful commands inside a pipeline:
Always use fetch-tags: true in the checkout step. Without it, the pipeline may not see tags and the $CI_COMMIT_TAG filter will break.
Tags are a minimal tool with maximum impact. Annotated with GPG signatures, --tags on release push, and v* filtering in CI — that set is enough to keep release history readable and verifiable.
6 - Deploying with a post-receive Git Hook
Deploying through CI is great, but sometimes you just need to push code to a server with a single git push. The post-receive hook in a bare repository handles this without extra dependencies: push to the server, and the hook automatically checks out files into the working directory.
Schema: bare repo as a deploy trigger
The logic is straightforward:
- A bare repository is created on the server (e.g.,
/srv/deploy/app.git). - The developer adds it as a remote and runs
git push origin main. - Git receives the data and triggers
hooks/post-receive. - The hook script runs
git --work-tree=/var/www/app --git-dir=/srv/deploy/app.git checkout -f main.
A bare repository has no working directory. That is why post-receive must explicitly pass --work-tree so checkout knows where to write files.
This is not CI — it is a direct trigger at the Git level. No pipelines, artifacts, or queues. It fits small services, VPS instances, and internal tools well.
Creating a bare repository on the server
SSH into the server and create the repo:
The resulting structure is standard: hooks/, objects/, refs/, HEAD, config. Default hooks live in /srv/deploy/app.git/hooks/ with a .sample suffix — rename them or create your own.
Make sure /srv/deploy is owned by the user you are pushing as. Otherwise, write permissions on objects/ will be denied.
Writing the post-receive hook
Create /srv/deploy/app.git/hooks/post-receive:
Key elements:
| Element | Purpose |
|---|---|
set -euo pipefail | Script exits on any error; does not continue with bad data |
while read oldrev newrev refname | The hook passes three arguments per line for each updated ref |
refs/heads/$BRANCH | Checks that the pushed ref is the target branch |
checkout -f | Force-synchronizes the working tree with the reference |
If you need to restart a service after deployment, append systemctl restart app or supervisorctl reload app to the end of the script. The hook runs in the context of the Git user, so verify it has permission to restart.
Setting up permissions and access
The most common issue is permissions. The Git user you push as must be able to write to WORK_TREE and execute commands from the hook.
Setup options:
- Single user: Git user and web user are the same identity. Simply
chown -R deploy:deploy /srv/deploy /var/www/app. - Group: Add the Git user to the web server group.
usermod -aG www-data deploy, thenchmod -R g+w /var/www/app. - SSH key: Push over SSH; the key authenticates as the correct user.
Do not forget to make the hook executable:
Pushing from the client and verifying
On the developer machine, add the remote and push:
On the server, the hook log will show:
Verify files on the server:
If the push fails with Permission denied, check the owner of hooks/post-receive and the target directory. If checkout does not write files, confirm that WORK_TREE points to an existing directory and the user has write access to it.
That is it. No additional tools — just Git and shell. For high-load production, this is obviously not a replacement for a proper CI pipeline, but for quick deploys to one or two servers it works reliably without unnecessary complexity.
7 - ulimit and systemd LimitNOFILE — why ulimit -n inside a unit doesn't stick
What is nofile and where it lives
nofile is the maximum number of open file descriptors per process. That’s not just regular files — it covers sockets, pipes, stdin/stdout/stderr, logs shipped through journald — everything counts. When nginx or a Go app crashes with too many open files, this is the limit to blame.
Limits live at three levels:
| Level | Where to check | What it controls |
|---|---|---|
| Kernel (system-wide) | /proc/sys/fs/file-max, /proc/sys/fs/nr_open | Absolute ceiling for the whole system |
| PAM / login | /etc/security/limits.conf, /etc/security/limits.d/ | For sessions via pam_limits.so |
| systemd | LimitNOFILE= in unit, DefaultLimitNOFILE= in system.conf | For systemd-managed services |
/proc/sys/fs/nr_open is the upper bound you can raise nofile to for a single process. It defaults to 1073741816 (≈1B) on most distros, but in practice you rarely need more than 1048576.
LimitNOFILE in a systemd unit
In a unit file the directive looks like this:
You can set both soft and hard limits at once, separated by a space:
The first value is the soft limit, the second is the hard limit. If you specify only one, it becomes the soft limit and the hard limit is taken from the system maximum.
The global default for all units is DefaultLimitNOFILE= in /etc/systemd/system.conf (and user.conf). Modern distros often default to 1048576, but older ones may have 4096 or 1024 — that’s exactly what catches you off guard.
Why ulimit -n in ExecStart doesn’t work
The typical mistake is trying to set the limit directly in the launch command:
ulimit -n inside ExecStart does not take effect on the process systemd tracks as MainPID. systemd sets limits before ExecStart runs, and the shell wrapper already operates in a context where the limit is fixed. Worse, ulimit -n may fail with Operation not permitted if the requested value exceeds the hard limit systemd assigned.
If you need to change the limit, use the [Service] directive — not a shell wrapper.
ulimit -n in a shell script called from ExecStartPre also does not propagate the limit to the main process. Limits are a property of the process, not the shell session.
How to check applied limits
Three ways, from simple to authoritative:
systemctl show reports exactly what systemd applied at fork() — that’s the authoritative source. /proc/<pid>/limits is what the kernel sees for the process. If they disagree, the problem is in an intermediate layer (PAM, container runtime, sudo).
For debugging ExecStart, you can add:
This shows the limits before the main process starts — what systemd set for the service.
PAM limits vs systemd limits
On most modern distros (RHEL 8+, Ubuntu 20.04+, Debian 11+) systemd does not invoke pam_limits.so for system services. Limits from /etc/security/limits.conf are not applied to unit files. They only work for login sessions (SSH, local login, su).
| Mechanism | Applies to | Configured in |
|---|---|---|
LimitNOFILE= in unit | Specific systemd service | /etc/systemd/system/*.service |
DefaultLimitNOFILE= | All systemd services | /etc/systemd/system.conf |
pam_limits.so / limits.conf | User login sessions | /etc/security/limits.conf |
ulimit in shell | Current shell and children | Interactive session |
If a service isn’t started through systemd (e.g., via supervisor, docker --ulimit, or directly), LimitNOFILE in the unit file doesn’t apply at all. In Docker that’s --ulimit nofile=65536:1048576, in supervisor the stdout_maxbytes and similar options aren’t related to nofile directly — you configure it at the OS or container level.
After changing LimitNOFILE in a unit file, always run systemctl daemon-reload and systemctl restart <unit>. systemctl reload doesn’t restart the process and doesn’t reapply limits.
8 - coredumpctl: Finding a Binary Crash
What is coredumpctl and how it works
When a binary crashes with SEGV, the kernel can save a core dump — a snapshot of process memory at the moment of the crash. In systemd-based distributions, coredumpctl handles collecting, storing, and searching these dumps. It’s a wrapper around systemd-coredump, which stores dumps in /var/lib/systemd/coredump/ and indexes metadata through journald.
Requires systemd-coredump and an active journald. In minimal containers without systemd, this tool is unavailable.
Installation is trivial for most distributions:
After installation, verify that kernel.core_pattern points to a pipe into systemd-coredump:
If it shows core.%e.%p, dumps are written to files and coredumpctl won’t see them.
Listing dumps: coredumpctl list
The base command prints all saved crashes:
Output contains columns: MESSAGE ID, TIMESTAMP, PID, UID, GID, COMM, EXE, COREFILE.
Use --no-pager for scripts, --json=short for parsing, -n for only the latest entry.
Key flags for list:
| Flag | Purpose |
|---|---|
-n N | Last N entries |
-1 | Only the latest single entry |
--since TIMESTAMP | Starting from date |
--until TIMESTAMP | Until date |
-u USER | By UID |
--exe PATTERN | By binary path |
--debug | Show service dumps |
Crash details: coredumpctl info
For a single dump, info gives everything needed before pulling the file:
Where 1234 is the process PID or number from list. Output includes:
- the signal that caused the crash (SIGSEGV, SIGABRT, etc.)
- timestamp
- executable path
- core file size
MESSAGE ID— correlation with journald
Extracting the core file: coredumpctl dump
The extracted file can be piped directly into gdb or saved to disk:
--output=- writes to stdout. If the core file is large (gigabytes), the pipe may block — better to write to disk.
Filtering and searching by binary, PID, time
In production, dumps accumulate by the dozens, and searching by PID is impractical. coredumpctl supports multiple filters at once:
For scripts and automation, JSON output is convenient:
Practical examples for debugging crashes
A typical debugging cycle looks like this:
If dumps don’t appear, check coredumpctl list and journalctl -u systemd-coredump. A common cause: LimitCORE in the systemd unit is set to 0, or kernel.core_pattern isn’t configured as a pipe.
For persistent monitoring, add to the unit file:
then restart the service. After that, coredumpctl will see all crashes.
9 - curl --resolve and SNI: Testing Virtual Hosts Without /etc/hosts
When you need to test a virtual host on a specific IP but don’t want to edit /etc/hosts — whether due to permissions, conflicts with other services, or just the habit of keeping the file clean — curl --resolve solves both problems at once: it overrides DNS resolution and sends the correct SNI in the TLS handshake.
Problem: virtual host without editing /etc/hosts
Multiple virtual hosts can live on a single IP, and the server picks the right one based on the Host header (HTTP/1.1) and SNI (TLS). Without an entry in /etc/hosts, curl first tries to resolve the name through DNS — getting the wrong IP, or no response at all.
Editing /etc/hosts works, but it requires sudo, clutters the file, and can break other services that depend on the same record.
curl –resolve: syntax and example
The --resolve flag intercepts name resolution at the curl level and substitutes it:
| Component | Meaning |
|---|---|
HOST | Virtual host name |
PORT | Port (typically 443 for HTTPS) |
ADDR | Target IP address |
Example:
curl sends the request to 203.0.113.50:443, but the HTTP Host header will still be example.com.
SNI in combination with –resolve
--resolve does not change the name curl sends in SNI. SNI is taken from the URL.
This is the key point. If the URL contains https://example.com, curl sends example.com as SNI in the TLS ClientHello — even though the IP is overridden via --resolve. The server uses SNI to select the correct certificate and virtual host.
To verify that SNI is actually being sent, use openssl:
If SNI is empty or wrong, the server returns the default certificate — and curl produces an error like SSL: certificate subject name does not match.
Practical example: checking a virtual host
Suppose several virtual hosts run on server 198.51.100.10, and you need to check app.local without touching /etc/hosts:
In the -v output you will see:
Connected to 198.51.100.10 (198.51.100.10) port 443— connection to the correct IP.> Host: app.local— correctHostheader.TLS SNI extension: "app.local"— SNI sent correctly.
For bulk-checking multiple hosts on the same IP:
--resolve applies only to the current curl process. Once the command finishes, the override disappears — /etc/hosts stays untouched.
If you need the override to work for all tools in the terminal, not just curl, then nsupdate or a local DNS resolver (dnsmasq, stubby) is a better fit. But for a one-off virtual host check, --resolve is faster and safer.
10 - Docker logs and journald: choosing a logging driver
When a container crashes, logs are the first thing you need to see. docker logs looks simple, but under the hood different logging drivers are at work, and the choice affects how logs are stored, rotated, and accessed. Here is what you should know before trusting the default.
How docker logs works
The docker logs <container> command reads the container’s stdout/stderr stream and outputs it to the terminal. Behind this sits a logging driver — a component that determines where the data actually goes. By default it is json-file: each container gets a JSON file on the host into which every output line is written.
docker logs does not read logs from inside the container directly — it queries the driver, which already stores the data in its own format and location.
The driver is configured at the Docker daemon level or per container. The choice affects log rotation, access via journalctl, and integration with centralized collection systems.
Driver json-file (default)
json-file is the built-in driver with no external dependencies. Each container creates a file at /var/lib/docker/containers/<container-id>/<container-id>-json.log. The format is JSON lines: each entry contains log, stream (stdout or stderr), and time.
Rotation is controlled by two flags:
| Flag | Description |
|---|---|
max-size | Maximum size of a single log file (e.g., 10m) |
max-file | Number of rotated files to retain |
Without these flags, log files grow without limits. In production this is a direct path to filling the disk.
If max-size and max-file are not explicitly set, Docker does not limit log size. On a host with many containers this will result in unexpected disk exhaustion.
Logs can be read via docker logs or directly from the file path on the host, but the latter is not recommended because files may be held open by the daemon.
Driver journald
journald sends container logs into the systemd journal. This means logs are accessible through journalctl, all of journald’s rotation and compression mechanisms apply, and there are no separate JSON files growing on disk.
This requires systemd and the systemd-journal-remote package (on some distributions). The container must be started with the driver specified:
The tag flag sets the identifier in journald — without it the tag is an empty string and finding the right container becomes difficult. The template {{.Name}} substitutes the container name.
Use tag={{.Name}} or tag={{.ID}} so that logs in journald are immediately tied to a specific container. Without a tag, filtering by CONTAINER_NAME does not work.
Reading logs:
More precisely, through journald filters:
Actual filtering depends on which metadata Docker passes into journald. Check journalctl -o verbose for a specific container to see what fields are available.
Comparing json-file and journald
| Parameter | json-file | journald |
|---|---|---|
| Storage location | /var/lib/docker/containers/... | /var/log/journal/ |
| Rotation | Via --log-opt | Via journald.conf |
| Search | docker logs --since, grep | journalctl --grep, --since |
| Dependencies | None | systemd |
| Centralization | Via fluentd, gelf, awslogs | Via journalctl --remote or forward |
| Compression | No (manual) | Yes, configurable in journald.conf |
| Access without Docker | Direct file access | Only via journalctl |
journald does not support all log-opt flags available for json-file. For example, max-size and max-file do not work — rotation is controlled by journald’s own settings (SystemMaxUse, SystemMaxFileSize, etc.).
Configuring the driver in daemon.json
The global setting goes in /etc/docker/daemon.json:
After changing it, restart Docker:
Changing the driver in daemon.json affects all new containers. Already running containers continue using their current driver until restarted.
Per-container override is possible via --log-driver and --log-opt at startup — this takes precedence over daemon settings.
To check the current driver for a specific container:
Practical recommendations
For local development, json-file with explicit max-size and max-file is sufficient and simple to use. For production with dozens of containers on a single host, journald is preferable: a unified search space, built-in compression, and integration with systemd and monitoring.
If you already use systemd to orchestrate containers (via systemd unit files or Podman), journald is the natural choice. Container logs and service logs end up in one place.
For centralized collection, both drivers support log forwarding through intermediate drivers (fluentd, gelf, splunk). But journald adds an extra step: first into journald, then the forwarder. For simple cases, direct json-file + fluentd may be a shorter path.
Monitor disk space regularly regardless of the driver. journalctl --disk-usage and du -sh /var/lib/docker/containers/*/ are the minimum set for this.
11 - scp — Secure Copy Over SSH
scp — a utility for copying files over SSH using the SSH protocol. It works from the terminal, requires no extra server setup — just a running sshd and working authentication. In an era of rsync and bat, SCP survives as a simple tool for one-off transfers when you don’t want to deal with daemons or configuration files.
Syntax and Basic Scenarios
General form:
Source and destination can be local paths or remote addresses in the format user@host:path.
If the remote host uses a non-standard SSH port, specify it with -P (uppercase P — that’s how scp differs from ssh).
Recursive Directory Copying
To copy a directory, the -r flag is required. Without it, scp refuses to transfer a directory and prints an error.
When copying recursively, scp transfers the contents of the directory, not the directory itself. Behavior depends on whether the trailing / is present in the path — verify the result if the structure matters.
Useful Flags
| Flag | Description |
|---|---|
-r | Recursive directory copying |
-P port | SSH port on the remote host |
-p | Preserves modification time, access, and file permissions |
-q | Quiet mode, no progress bar |
-C | Compression during transfer |
-i key.pem | Specifies a private key |
-o StrictHostKeyChecking=no | Automatically accepts new host keys |
-l rate | Bandwidth limit in Kbit/s |
Typical Transfer Patterns
Deploying artifacts:
Fetching a log from a remote server:
Transferring multiple files at once:
Direct transfer between two servers (both source and destination are remote):
The -3 flag routes traffic through the local machine. Without it, scp attempts a direct connection between the hosts, which usually fails.
Limitations and Alternatives
scp does not support resuming interrupted transfers — if the connection drops halfway through a multi-gigabyte file, you start over. There is no incremental sync, no checksum verification. Transfers are sequential, with no built-in parallel file transfer.
For everyday tasks, rsync solves these problems:
rsync can resume broken transfers, skip already-copied files, and work incrementally. For one-off transfers of a few small files, scp remains convenient — fewer parameters, faster to type.
On modern distributions, scp may be wrapped by the OpenSSH client. Behavior is the same, but if you notice differences in output or error handling, that’s normal — the implementation depends on the OpenSSH version.
12 - Sudoers: NOPASSWD Without Holes
Unrestricted NOPASSWD in sudoers is a misconfiguration that grants root access without a password, turning any user script or library vulnerability into a full system compromise. The correct approach limits NOPASSWD to specific commands only.
Why NOPASSWD + ALL Is a Hole, Not a Solution
%admin ALL=(ALL) NOPASSWD: ALL — the most common sudoers error. The user receives unlimited root access without a password. Any script, any utility, any vulnerability in the user’s environment becomes a direct path to full machine control. NOPASSWD without command restrictions is not convenience; it is a backdoor in plain sight.
The proper method: allow specific commands via Cmnd_Alias and attach NOPASSWD only to them. Then the user can restart a service but cannot read /etc/shadow or run su.
visudo: The Only Safe Way to Edit sudoers
Editing /etc/sudoers directly via vi or nano is a path to locking yourself and the entire team out. visudo locks the file, validates syntax before saving, and rejects invalid entries.
A syntax error in sudoers = loss of sudo capabilities for all. visudo prevents this, but only when used.
Cmnd_Alias: Grouping Commands Instead of Allowing Everything
Cmnd_Alias lets you create a named group of commands. You then reference the name — readable and easy to change.
Alias syntax: name in uppercase, comma-separated absolute paths. The path is mandatory — systemctl without /usr/bin/ will not work.
Example: Limited NOPASSWD for Specific Tasks
Real scenario: a deployer user needs to restart nginx and view logs, nothing more.
Now deployer can:
To allow a single user one binary:
For multiple users performing the same task — use a group. %deployers ALL=(ALL) NOPASSWD: RESTART_WEB is more maintainable than duplicating lines.
What to Avoid: Common sudoers Mistakes
| Ошибка | Почему плохо | Как правильно |
|---|---|---|
ALL ALL=(ALL) NOPASSWD: ALL | Full root without password | List specific commands via Cmnd_Alias |
user ALL=NOPASSWD: /bin/bash | Opens root shell | Never give interpreters or su |
user ALL=(ALL) ALL, NOPASSWD: ALL | NOPASSWD spreads to everything due to order | NOPASSWD: before the command list, not after |
Editing /etc/sudoers via echo or cp | No syntax validation | Only visudo or visudo -f |
Absence of #includedir /etc/sudoers.d | Manual file overwritten on update | Check include, place custom rules in /etc/sudoers.d/ |
Another frequent trap — spaces in Cmnd_Alias. A comma and space after it are mandatory:
Verify a rule without root privileges:
The output lists permitted commands and their flags. An empty list means the rules did not apply — check syntax and paths. Here’s a thinking process:
- Analyze User Input:
- Role: Practicing Lead DevOps writing IT notes for a Hugo blog (Mana theme).
- Style: Short paragraphs, specific commands, tables of flags, code with language fences.
- Constraints: No water, no “in this article we’ll explore”, no marketing, no emojis, no author signature at the end.
- Allowed: GitHub admonitions:
> [!NOTE],> [!TIP],> [!WARNING]. - Fact Constraint: Don’t make up non-existent flags, URLs, or versions. If fact is unclear, write cautiously as in operations.
- Task: Write the English article as a parallel original, not a word-for-word translation.
- Structure & Facts: Must match the Russian draft exactly.
- Format:
13 - ThinLinc: Remote Access to Linux Desktops
ThinLinc: Remote Access to Linux Desktops
In enterprise environments, remote access to Linux desktops often comes down to VNC with flaky encryption or RDP proxies held together with workarounds. ThinLinc by Cendio is a full-featured remote desktop solution that runs on top of VNC, supports RDP clients, and provides a web-based administration interface without the usual hassle with certificates and firewalls.
What Is ThinLinc
ThinLinc is a remote desktop access solution with a server-plus-clients architecture. The server runs VNC sessions on the backend, and clients connect through a web browser or native ThinLinc clients. The protocol between client and server is tunneled over SSH, which solves the encryption problem out of the box.
Key features:
- Built-in web server for browser-based desktop access
- Native clients for Linux, Windows, and macOS
- Load balancing across multiple servers
- Single sign-on (SSO) via Kerberos and LDAP
- Session management through web panel and CLI
Installing the Server
ThinLinc ships as packages for RHEL/CentOS 7/8/9 and Ubuntu/Debian. For installation on a RHEL system:
For Ubuntu/Debian:
After installation, the web panel is available at https://<host>:1443.
Port 1443 is the standard ThinLinc web interface port. If a firewall is in place, open it along with port 22 for SSH tunneling.
Server Configuration
The main configuration lives in /etc/thinlinc/. Key files:
| File | Purpose |
|---|---|
tlconfig | Global server settings |
vncserver-config-defaults | VNC session parameters |
client-to-server.d/ | Port and device forwarding rules |
ssl/ | Certificates and keys |
Basic configuration via tlconfig:
For VNC parameter tuning:
After changing configuration via tlconfig, restart the services: sudo systemctl restart tlwebd vncserver@:*.
Client Connections
ThinLinc provides several connection methods:
- Web browser — navigate to
https://<host>:1443, enter credentials. Works on any device with a modern browser. - Native client — download from the same URL. Clients are available for Linux, Windows, and macOS.
- RDP client — ThinLinc supports RDP proxying, allowing connections via standard
rdesktoporfreerdp.
Connecting via native client:
Manual SSH tunnel connection:
Administration and Session Management
All active sessions can be viewed through the web panel (https://<host>:1443/admin) or via CLI:
For bulk operations:
The admin web panel allows:
- Viewing session lists and their status
- Sending messages to users
- Forcefully terminating sessions
- Viewing logs and usage metrics
Security and Integration
ThinLinc uses SSH for encrypting traffic between client and server. Web server certificates can be replaced with your own:
LDAP/Kerberos integration:
For PAM integration:
ThinLinc supports USB device and audio forwarding over the SSH tunnel. Configure it in client-to-server.d/ rules for specific devices.
ThinLinc is a ready-made solution that eliminates the manual assembly of VNC infrastructure with firewalls and certificates. For environments that need remote access to Linux desktops without security compromises, it is one of the most straightforward paths available.
14 - Chrony Instead of ntpd
Chrony has replaced ntpd as the default NTP client in most modern Linux distributions. It converges to accurate time faster, handles intermittent network connections better, and consumes fewer resources. If your machine still runs ntpd, switching takes only a few minutes.
Installation
On RHEL-based systems:
On Debian/Ubuntu:
If ntpd was running on this machine before, stop and disable it to avoid port conflicts:
makestep Configuration
The key directive in /etc/chrony/chrony.conf (or /etc/chrony.conf on RHEL) is makestep. It controls how chronyd behaves at startup — whether to correct time gradually or in a single step.
Format: makestep <max_offset> <max_updates>. If the offset exceeds <max_offset> seconds and the number of updates hasn’t exceeded <max_updates>, chrony applies a sudden correction instead of gradual slewing. The default makestep 1.0 3 allows up to a 1-second jump three times during the first synchronizations.
A common mistake is setting makestep -1 1 and expecting a server with a large initial offset to correct instantly. In practice, negative values only work under specific conditions. For reliable startup, use a positive number and limit the number of steps.
allow Directive
By default, chronyd operates only as a client. To let other hosts synchronize through this server, add allow:
You can specify individual IPs, subnets, or multiple allow lines for different networks. Without this directive, the machine accepts requests only from localhost.
To block a specific host, use deny — it takes effect after allow and overrides it:
After changing the configuration, restart the service:
Verification with timedatectl
timedatectl shows the current synchronization state and time source:
Example output:
Key fields for diagnostics:
| Field | Problem Value | Meaning |
|---|---|---|
System clock synchronized | no | chrony hasn’t caught up yet |
NTP service | inactive | service not running or not enabled |
RTC in local TZ | yes | hardware clock in local timezone — common issue on VMs |
For detailed information about current sources:
If NTP service: active and System clock synchronized: yes, everything is working. If not, check sudo systemctl status chronyd and network access to NTP servers (UDP port 123).
15 - fail2ban: SSH Jail Configuration
Securing SSH against brute-force attacks is one of the first steps in hardening any server. fail2ban scans logs, detects repeated failed login attempts, and blocks the source via iptables or nftables. This note covers the sshd jail — from installation to fine-tuning ban durations.
Installation and Basic Configuration
Install from the standard repository:
Enable and start the service:
Do not edit /etc/fail2ban/jail.conf directly — package updates will overwrite your changes. All local overrides go in jail.local.
Create a local overrides file:
Or, preferably, create a minimal jail.local with only the parameters you need — fail2ban stacks jail.local on top of jail.conf, so overriding specific keys is sufficient.
Filter for sshd
A ready-made filter ships with the package: /etc/fail2ban/filter.d/sshd.conf. It scans /var/log/auth.log (or /var/log/secure on RHEL) for patterns like Failed password for invalid user and Connection closed by authenticating user.
Verify the filter parses your logs correctly:
If the output shows a high match rate, the filter works. If not, check the logpath parameter in the jail and the logpath value inside the [Definition] section of the filter.
For additional protection against brute-force via ddos-style attacks, there is a separate filter sshd-ddos. It catches rapid repeated connections from a single IP. Enable it by adding mode = ddos to the jail parameters.
bantime and findtime Parameters
Three keys define the blocking logic:
| Parameter | Description | Default |
|---|---|---|
bantime | Ban duration in seconds (negative = permanent) | 600 |
findtime | Observation window for counting failed attempts | 600 |
maxretry | Failed attempts before ban | 5 |
A typical production configuration:
This means: three failed logins within ten minutes triggers a one-hour ban. For critical servers, set bantime = -1 (permanent ban) and unblock manually with fail2ban-client set sshd unbanip <IP>.
bantime and findtime accept suffixes: d (days), h (hours), m (minutes), s (seconds). For example, bantime = 1d.
Activating the Jail
By default, the [sshd] section is commented out in jail.local. Enable it:
On RHEL/CentOS, logpath is typically /var/log/secure. Verify the log path on your distribution.
Restart fail2ban and check the status:
The output of fail2ban-client status sshd shows the current number of banned IPs and active filters.
Manual block and unblock:
fail2ban is not a WAF and not a replacement for key-based authentication. Use keys instead of passwords, restrict access with AllowUsers, and change the port where feasible. fail2ban complements these measures — it does not replace them.
Check fail2ban logs (/var/log/fail2ban.log) on first startup — they show whether the jail picked up the log file and whether filters are firing.
16 - journalctl: filters and follow
Systemd’s journal is the first place to look when a service crashes or a node starts burning CPU. journalctl does far more than dump the entire log in sequence: it can filter by units, priorities, time windows, and stream in real time. Below is the working set I use daily.
Follow in real time
Behavior similar to tail -f, but aware of journald’s structured format:
The -f flag (short for --follow) streams new entries as they appear. By default it shows all units — handy when you don’t know where the fire is.
Add --no-pager so output isn’t intercepted by less and doesn’t block the terminal in scripts and CI.
Filter by unit
The unit is the most common filter. One -u key and the specific service name:
Multiple units can be passed — journalctl will show entries from all specified ones:
The unit name must include the .service suffix. Omitting it may result in no matches or unexpected output from journalctl.
Priority and time filters
Priority filtering is set via -p (or --priority). Levels range from 0 to 7:
| Priority | Value |
|---|---|
| 0 | emerg |
| 1 | alert |
| 2 | crit |
| 3 | err |
| 4 | warning |
| 5 | notice |
| 6 | info |
| 7 | debug |
Show only errors and critical:
Time windows use --since and --until. Both absolute dates and relative expressions are supported:
--since without --until shows from the specified moment to now. If both are specified — the window is closed. Check the date order, or you’ll get empty output.
Combining flags
In practice, filters are combined. A typical request: watch errors from a specific service over the last hour in real time:
Another common pattern — show the last N lines for a unit with a priority filter:
Here -n 100 limits output to the last 100 entries.
Quick reference for everyday flags:
| Flag | Purpose |
|---|---|
-u <unit> | Filter by unit |
-f | Follow (new entries in real time) |
-p <level> | Minimum priority |
--since <time> | Start of time window |
--until <time> | End of time window |
-n <count> | Last N entries |
--no-pager | Without pager |
If journalctl returns empty results without errors — check whether the service is inactive or failed. You can see the status via systemctl status <unit>. Also verify that journald hasn’t rotated the needed entries: journalctl --disk-usage shows how much space the journal occupies.
17 - nftables: Basic Rule Set
All commands were verified on Debian/Ubuntu with the nftables package and on RHEL/CentOS 8+. On older systems you may need apt install nftables or yum install nftables.
nftables replaced iptables, but documentation for a basic rule set is often scattered. Here is the reference I use when bringing up a firewall on a new host.
Creating the inet filter table
The inet family table unifies IPv4 and IPv6 under a single namespace. This is the preferred approach when both stacks are active on the host.
If the table already exists the command returns an error. To avoid duplication when a script is re-run:
To wipe the current rule set before loading your own:
flush ruleset removes all rules instantly. On a production machine run this only from the console, not over a remote session without a fallback.
Input and forward chains
Chains bind to a table and define the interception point for traffic. A basic firewall needs input (traffic destined for the host itself) and forward (traffic passing through the host).
Key components inside the curly braces:
| Component | Value |
|---|---|
type filter | Chain type, standard for packet filtering |
hook input / hook forward | Interception point in the network stack |
priority 0 | Processing priority |
policy drop | Default policy — drop unmatched packets |
The syntax requires escaping semicolons inside the string or using single quotes as shown above.
If you need to allow established connections, add an output chain with policy accept or use connection tracking in your input rules.
Basic rules for input
With a chain set to policy drop, you must explicitly permit the traffic you need. A typical minimum:
Each rule appends to the end of the chain. Order matters: accept rules for loopback and established traffic should come before rules with narrower criteria.
To log dropped packets before the drop policy (optional but useful for diagnostics):
Logging adds overhead. On high-throughput interfaces use a rate limit: limit rate 10/second.
Basic rules for forward
The forward chain is needed when the host acts as a router or NAT gateway. A minimum setup:
If the host does not perform routing, leave forward with policy drop and no additional rules.
To enable IP forwarding at the kernel level (if not already done):
Viewing and managing rules
After setup, verify the current configuration:
list ruleset outputs the full config, which you can save and reuse as a boot script.
Deleting a specific rule by handle within a chain:
The handle number appears in the output of nft list ruleset -a.
To delete an entire chain:
A chain can only be deleted when it is empty. To remove a table completely, delete all chains inside it first.
Saving rules to a file for boot-time loading:
On systems with systemd, enable auto-start:
If /etc/nftables.conf does not exist or is empty, the service will not load any rules. Create the file manually before enabling the service.
18 - Debian — The Swiss Army Knife of Linux
Debian isn’t the flashiest distribution, but it’s the most reliable foundation in the Linux world. Behind it stands the largest community of volunteer developers, and behind its track record are decades of uninterrupted operation on servers, embedded systems, and cloud infrastructure. If you manage Linux machines in production, Debian (or one of its derivatives) is already in your stack — whether you’ve acknowledged it or not.
Managing Packages: apt and dpkg
Two tools form the core of Debian’s packaging system. apt is the high-level interface for working with repositories, resolving dependencies, and upgrading the system. dpkg is the low-level engine that installs, removes, and inspects individual .deb files without contacting repositories.
dpkg -i does not resolve dependencies. If a package pulls in other libraries, use apt install ./package.deb instead — apt will fetch missing dependencies from the repository.
A common mistake is mixing apt and dpkg during partial installations. If dpkg -i fails with a dependency error, run apt --fix-broken install and complete the operation through apt.
Stability Model and Release Branches
Debian is built on three branches that determine the aggressiveness of updates and the level of testing.
| Branch | Codename | Stability | When to Use |
|---|---|---|---|
| stable | bookworm (12) / trixie (13) | High — security and critical fixes only | Production servers, base infrastructure |
| testing | bookworm-progress | Medium — packages go through a stabilization period | Development machines, containers |
| unstable | sid | Low — daily snapshots | Experiments, binary package building |
In sources.list this looks like:
The non-free and non-free-firmware sections were introduced starting with Debian 12. Without them, certain proprietary drivers (Wi-Fi, GPU) won’t install.
Moving between branches is apt dist-upgrade with a swapped sources.list. Do this only on test benches — binary compatibility between testing and stable is not guaranteed for all packages.
Architecture Support
Debian officially supports more hardware platforms than any other distribution. This is critical for embedded and IoT projects.
| Architecture | Port | Typical Use |
|---|---|---|
| amd64 | amd64 | Servers, desktops, cloud |
| arm64 | arm64 | ARM servers (AWS Graviton, Raspberry Pi 4/5) |
| armhf | armhf | Embedded devices with FPU |
| i386 | i386 | Legacy systems |
| ppc64el | ppc64el | IBM Power Systems |
| s390x | s390x | IBM Z mainframes |
| riscv64 | riscv64 | Open RISC-V platforms |
Multi-architecture support allows installing packages for foreign platforms:
Debian as a Base for Other Distributions
Nearly every major distribution you use is built on Debian or inherits its packaging model.
- Ubuntu — takes the testing branch, adds its own PPAs and proprietary drivers.
- Linux Mint — based on Ubuntu, and therefore on Debian packages.
- Proxmox VE — a modified Debian 12 with added enterprise repositories.
- Kali Linux — Debian unstable with a suite of pentesting tools.
- Raspberry Pi OS (formerly Raspbian) — Debian armhf/arm64 with Pi-specific patches.
- MX Linux, antiX — Debian stable with lightweight desktop environments.
If you need a stable backend with fresher software, use Debian stable with the backports repository. This gives you updated kernel, nginx, PostgreSQL versions without switching to testing.
Practical DevOps Scenarios
In day-to-day work, the Debian approach manifests in several key patterns.
Base container images. Official Docker images like debian:bookworm, debian:slim are the starting point for hundreds of minimalist containers. debian:slim weighs around 70 MB compared to 200+ MB for ubuntu:latest.
VM templates. Proxmox, which runs on Debian, uses debootstrap to create LXC container templates. It is the same tool that underlies a fresh system installation.
Configuration management. Ansible, Puppet, Chef — all three work directly with Debian packages. The apt module in Ansible is one of the most-used modules in any playbook.
CI/CD pipelines. Building .deb packages in GitLab CI or GitHub Actions is a standard pattern. debhelper, dh_make, pbuilder, or sbuild provide reproducible builds in a clean chroot.
Key Commands for Daily Work
Bookmark this table — it covers 90% of tasks on a Debian server.
| Task | Command |
|---|---|
| Update package list | apt update |
| Upgrade all packages | apt upgrade -y |
| Full upgrade resolving dependencies | apt full-upgrade -y |
| Install a package | apt install -y <pkg> |
| Remove a package with configs | apt purge -y <pkg> |
| Search for a package | apt search <query> |
| Show package info | apt show <pkg> |
| List installed packages | apt list --installed |
| Find which package owns a file | dpkg -S <path> |
| List files in a package | dpkg -L <pkg> |
| Check package status | dpkg -l <pkg> |
| Fix broken dependencies | apt --fix-broken install |
| Clean apt cache | apt clean && apt autoclean |
| Add an architecture | dpkg --add-architecture <arch> |
| Build .deb from source | debuild -us -uc |
Debian isn’t the fastest at first glance — no rolling releases and no fresh software versions out of the box. But that predictability is exactly what makes it the foundation for production environments running millions of servers. If you want reliability over trending versions, Debian won’t let you down.
19 - pipx: Isolated Python CLI Tools Without the Mess
pipx solves a simple but chronic problem: you need to run a Python utility once or occasionally, and pip install pollutes the global environment or leaves behind a virtual environment you forget to clean up. pipx creates an isolated venv for each utility, installs dependencies there, and makes the binary available in $PATH. One command and the tool works without conflicting with anything.
What Is pipx and Why You Need It
pipx installs and runs Python applications in isolated virtual environments. Each utility lives in its own venv under ~/.local/pipx/venvs/, and its console-scripts are symlinked into ~/.local/bin/.
Problems it solves:
- Version conflicts between projects:
black==23andblack==24can’t coexist in one environment, but in pipx they can. - Global
pip installpollutes system Python and can breakapton Debian-based systems. - Forgotten venvs after one-off usage.
pipx doesn’t replace pip inside projects. It’s a tool for CLI utilities: black, poetry, httpie, ansible, awscli, pre-commit, and similar.
Installing pipx
The most reliable path is through pip in user mode or through your system package manager.
After installation, verify ~/.local/bin is in $PATH:
If ensurepath didn’t work, add it manually to ~/.bashrc or ~/.zshrc:
export PATH="$HOME/.local/bin:$PATH"
Basic Commands: install, run, list
Three commands cover 90% of use cases.
Flags worth remembering:
| Flag | What it does | Example |
|---|---|---|
--spec | Specify source (PyPI, git, wheel) | pipx install --spec git+https://github.com/user/repo.git tool |
--suffix | Add suffix to binary name | pipx install black --suffix==24 |
--python | Specify interpreter | pipx install --python python3.11 black |
--system-site-packages | Enable access to system packages | pipx install --system-site-packages tool |
--force | Reinstall over existing | pipx install --force black |
pipx run is the key command for one-off usage. It downloads the package, creates a temporary venv, executes, and removes it. No trace left behind.
Managing Dependencies and Reinstallation
After installation you can upgrade, remove, and inspect dependencies.
If something breaks, reinstallation takes seconds:
pipx upgrade-all can bump a tool to a version with breaking changes. In CI/CD, pin the version instead: pipx install black==24.8.1.
Typical DevOps Scenarios
pipx fits several working patterns where a full project with requirements.txt is overkill.
1. One-off utilities in CI/CD. Instead of installing into a Docker image or globally:
2. Parallel versions of the same tool. Useful during project migrations:
3. Local dev environment without privileges. Install ansible, pre-commit, and similar tools without sudo and without affecting system Python.
4. Quick package evaluation before integration. Test a utility without adding it to requirements.txt:
Combined with direnv and .envrc, you can wire up pipx run for project-specific tasks — the utility is available only in that directory, and dependencies don’t leak into the global environment.
pipx doesn’t try to be a package manager for all of Python. It does one thing — isolated CLI utility installation — and does it without the noise. For a Lead DevOps, that means less time fighting dependency conflicts and more time on architecture.
20 - uv — Fast Python Package Manager
What is uv
uv is a Python package manager written in Rust. It solves one problem: the standard pip is slow at resolving dependencies, and poetry adds its own project model on top of PEP 621. uv works with pyproject.toml, is compatible with PEP 621 and PEP 508, and does it significantly faster.
Under the hood is a caching resolver written in Rust that reuses data from pip-compatible indexes (PyPI by default). uv can replace pip, pip-tools, virtualenv, and poetry in a single tool.
uv does not require an installed Python to install itself, but to work with Python projects you need an interpreter. uv can find and manage CPython via uv python.
Installation
The fastest path is the curl script:
This places the binary in ~/.local/bin. On systems without curl:
Or download the standalone binary from GitHub releases.
After installation, verify the version:
If uv is not found after installation, make sure ~/.local/bin is in your $PATH. On Debian/Ubuntu, the pip-installed package places it in /usr/local/bin.
Key commands
The project workflow:
Key flags for daily work:
| Flag | What it does |
|---|---|
--system | Installs a package into the system Python, bypassing venv |
--no-dev | Excludes dev dependencies during sync |
--group <name> | Specifies a named dependency group |
--python <version> | Selects a specific interpreter version |
--locked | Uses the lock file without re-resolving |
--quiet | Minimal output |
--verbose | Detailed output for debugging |
Managing interpreters:
For package-level work without creating a project:
uv sync recreates .venv and installs exactly what is in uv.lock. This makes it suitable for CI — the result is deterministic.
Comparison with pip and poetry
| Aspect | pip | poetry | uv |
|---|---|---|---|
| Implementation language | Python | Python + Rust | Rust |
| Project model | setup.py / pyproject.toml | pyproject.toml (custom format) | pyproject.toml (PEP 621) |
| Resolution speed | Slow | Medium | Fast |
| Lock file | No native | poetry.lock | uv.lock |
| Virtual environment | venv / virtualenv | Built-in | Built-in (.venv) |
| pip compatibility | Full | Partial | Full |
| Replaces pip | — | — | Yes |
| Replaces poetry | — | — | Yes (partially) |
uv wins on resolution speed — in astral-sh benchmarks it is 10–50x faster than pip on large dependency graphs. At the same time, uv pip is fully compatible with pip workflows: requirements.txt, --index-url, --find-links.
uv is under active development. The uv.lock format may change between major versions. In production CI pipelines, pin the uv version via uv --version or install a specific binary.
A common mistake is confusing uv add with uv pip install. The former writes the dependency into pyproject.toml and updates uv.lock; the latter works like pip directly. For projects with pyproject.toml, use uv add. For one-off scripts and Dockerfiles, use uv pip install.
Another frequent issue: uv cannot find the interpreter if Python is installed in a non-standard path. The fix is uv python install --force or explicit specification via --python /path/to/python.
21 - logrotate for Custom Daemons
Log rotation is missing for your daemon, the log file has grown to dozens of gigabytes, the disk is full, and monitoring is screaming. systemd-journald and syslog-ng rotate on their own, but if your custom daemon writes directly to a file, rotation falls to logrotate. Here is how to configure it for a specific service.
Why write a custom logrotate config
Packages from the repository usually drop their config into /etc/logrotate.d/, but for self-built daemons or those compiled from source, there is none. Without a config the file grows without limits. logrotate runs via a systemd timer (logrotate.timer) or cron and reads all files from /etc/logrotate.d/. Creating a single file is enough to start the rotation cycle.
Verify that the logrotate package is installed and the timer is active: systemctl status logrotate.timer. On most distributions it is enabled by default.
copytruncate vs create — when to use which
These are two fundamentally different approaches to renaming and creating a new file.
copytruncate | create | |
|---|---|---|
| Mechanism | Copies the current file, truncates the original in place | Renames the old file, creates a new one with correct permissions |
| Application impact | No restart needed | Daemon must be able to open the new file (typically via SIGHUP) |
| Risk of lost lines | Yes — entries written between copy and truncate can fall into the gap | Minimal — renaming is atomic |
| When to use | Daemon cannot re-open its file (e.g., written in Go without signal handling) | Daemon supports SIGHUP or systemd-notify |
copytruncate is a compromise. Lines written between cp and truncate are lost. For high-traffic services this can mean hundreds of lost lines per second.
If the daemon can receive signals, use create and restart it via postrotate.
delaycompress and how it affects the archive chain
By default logrotate compresses the rotated file in the same cycle. The problem: if the daemon is still writing to the old file (or has not yet re-opened the new one), compression breaks everything.
delaycompress defers compression by one cycle. The chain looks like this:
Without delaycompress the transition is harsher: app.log immediately becomes app.log.1.gz, and if the daemon is still writing to app.log through its descriptor, data goes into the compressed archive — or is lost entirely.
delaycompress only makes sense together with create and compress. With copytruncate it works but loses its meaning — the file is truncated in place, so compression can happen immediately.
Integration with systemd notify
If the daemon supports sd_notify(3), you do not need to rely on postrotate with a manual restart. systemd can restart the service based on a signal from logrotate.
In the logrotate config, specify:
Or, if the daemon listens on NOTIFY_SOCKET:
systemctl notify-reload is available starting from systemd 231. It sends RELOADING=1 followed by READY=1 — the standard mechanism for notifying a configuration reload.
For copytruncate, postrotate is usually unnecessary — truncating the file in place does not require daemon involvement.
Example working config
Suppose daemon my-app writes to /var/log/my-app/app.log and supports SIGHUP and sd_notify.
Flags:
| Flag | What it does |
|---|---|
daily | Rotate every day |
rotate 14 | Keep 14 archives |
compress | gzip compression |
delaycompress | Delay compression by one cycle |
missingok | Do not error if the file is absent |
notifempty | Do not rotate an empty file |
create 0640 myapp myapp | Create new file with correct permissions and owner |
postrotate | Notify systemd of reload |
To test without actually rotating:
The -d flag runs in debug mode — it shows what would be done, without modifying any files.
After creating the config, the first rotation happens on the next timer tick. To force an immediate check: logrotate -f /etc/logrotate.d/my-app. This rotates the file right away, so use it carefully on production.
22 - Rsync: Backing Up a Directory Over SSH
Rsync: Backing Up a Directory Over SSH
The classic way to copy a directory to a remote machine is rsync over SSH. No extra ports to open, traffic is encrypted, and the tool itself handles incremental transfers and metadata preservation. One command and the backup is ready.
Basic rsync Command Over SSH
The minimal invocation to copy a local directory to a remote host:
The trailing slash on source matters: without it, rsync creates a source/ subdirectory on the remote side; with it, the contents go directly into dest/.
If SSH listens on a non-standard port:
For cron automation, specify the port via -e rather than editing /etc/ssh/ssh_config — it’s easier to maintain different hosts with different ports.
Archive and Sync Flags
-a (archive) is the key flag. It bundles several options into one: recursive traversal, preservation of permissions, ownership, timestamps, symlinks, and empty directories.
| Flag | Purpose |
|---|---|
-a | Archive mode (recursion + metadata) |
-v | Verbose output |
-z | Compression during transfer |
-P | --partial --progress — resume and progress bar |
--delete | Remove files on receiver absent from source |
-e ssh | Specify remote shell |
--bwlimit=KBPS | Bandwidth limit in KB/s |
--delete is a powerful tool. If the source is accidentally cleared, the remote machine will be left empty. Verify the list before applying.
Full backup example with bandwidth limit and progress:
File and Directory Exclusions
Use --exclude to omit specific paths. Patterns are relative to the source:
When there are many exclusions, a file list is more convenient:
--exclude is evaluated in order. Later rules can override earlier ones if paths overlap. For precise control, use --filter.
Dry-Run Before Launch
Always run with --dry-run (or -n) first. rsync shows what would be copied or deleted without touching any files:
Compare the output against the expected list. If it matches, remove -n and run the real sync.
For a cron job with email notifications:
Summary
rsync over SSH covers 90% of backup tasks without additional agents on the receiver side. The essential order of operations: first --dry-run, then --exclude, then --delete if exact mirroring is needed. Everything else is host- and port-specific tuning.
23 - Setting Up Your Own SSH Bastion Server
Why You Need a Bastion and Where It Lives
A bastion is the single entry point into a private network segment. Instead of exposing SSH on every server to the internet, you funnel traffic through one hardened host with a strict access policy. Typical layout: internet → bastion (public IP) → internal servers (only private subnet, SSH listening on 127.0.0.1 or a private interface).
The bastion sits in a demilitarized zone (DMZ) or a public subnet provided by your hosting platform. Internal machines have no route to the internet through the bastion — return traffic flows only over established connections. This is a baseline model you can deploy on any VPS in about 15 minutes.
Choosing an OS and Basic Setup
The bastion doesn’t need to be heavy. Debian, Ubuntu Server, or AlmaLinux — all work. I usually grab a minimal Ubuntu 22.04 LTS install and bring it to working state by hand.
Don’t put a GUI, databases, or other services on the bastion. The smaller the attack surface, the better.
Create a non-privileged user for daily work:
SSH Configuration: Keys, Port, Disable Passwords
Generate a key on your workstation if you don’t have one yet:
On the bastion, edit /etc/ssh/sshd_config:
Before restarting SSH, make sure the key is added and works. Otherwise, you’ll lock yourself out of the server.
Validate the config and restart:
sshd_config keys and what they do:
| Parameter | Value | Purpose |
|---|---|---|
Port | 22220 | Non-standard port, reduces log noise |
PermitRootLogin | no | Blocks direct root login |
PasswordAuthentication | no | Keys only |
MaxAuthTries | 3 | Limits authentication attempts |
AllowUsers | deploy | Whitelist of allowed users |
Firewall and Access Restrictions
UFW is the simplest way to close everything unnecessary:
If the bastion is only for your IP, restrict it further:
If your IP is dynamic, use a VPN instead of exposing the port. Opening SSH to the internet without source restrictions is a bad practice.
For internal traffic, add a rule on the bastion to allow packet forwarding:
Add net.ipv4.ip_forward = 1 to /etc/sysctl.conf so the setting survives a reboot.
Fail2ban and Brute-Force Protection
Install and configure Fail2ban to protect against password guessing:
In /etc/fail2ban/jail.local:
Restart:
Logging and Auditing
The bastion must log everything. On Debian/Ubuntu, SSH logs go to /var/log/auth.log. For centralized collection, configure rsyslog to ship to a separate SIEM or at least a second server:
For user action auditing, attach auditd:
View events:
Logs on the bastion are an attacker’s first target. Set up remote shipping as early as possible — otherwise, on compromise, you lose the incident history.
Port Forwarding and Tunnels Through the Bastion
The bastion’s main job is to give access to internal machines without exposing their ports. Three ways to do it:
1. Port forwarding via SSH tunnel:
Now local port 5432 on your machine proxies to PostgreSQL on 10.0.1.5.
2. SOCKS proxy for access to all internal hosts:
Point your browser or proxychains at 127.0.0.1:1080.
3. Reverse tunnel for accessing your local machine from the bastion network:
This lets someone reach a local service on port 9090 via the bastion.
For persistent tunnels, use autossh or configure ~/.ssh/config:
Now ssh bastion is all you need to connect, and tunnels are built with a single command.
For team operations, store keys in 1Password or HashiCorp Vault, and grant bastion access through ephemeral sessions with expiring tokens. This reduces the risk of key compromise.
24 - tmux on Prod After Screen
Why We Switched from Screen to tmux
Screen was our primary tool for about five years. After migrating the cluster to new servers it became obvious: screen drops sessions on SSH disconnect when hardstatus isn’t configured, and screen -r recovery sometimes hits a race condition when multiple admins connect simultaneously. tmux solves both problems out of the box — sessions live in server memory, are bound to a socket, and reconnection doesn’t depend on the TCP connection state.
The migration took half a day: we set up a shared config, distributed keybindings, and validated on staging. We went to production a week later — after tmux survived two incidents where screen would have lost context.
Key Differences from GNU Screen
We’re not retelling Screen’s history. The focus is on what tmux gives differently.
The main difference is architectural. tmux uses a client-server model with a separate server process per session. Screen is also client-server, but its session is bound to the terminal less reliably and is lost more often on disconnect.
| Aspect | Screen | tmux |
|---|---|---|
| Session recovery | screen -r (may fail) | tmux attach -t <name> (stable) |
| 256-color support | Limited | Full, default-terminal "screen-256color" |
| Input synchronization | multiuser + acladd | set -g allow-rename off + shared sessions |
| Configuration | ~/.screenrc | ~/.tmux.conf |
| State after disconnect | Often lost | Session stays alive, socket remains |
| Scripting | screen -X | tmux send-keys, tmux split-window |
Another thing — tmux has a proper copy-mode. In Screen you had to fight with text selection through escape sequences. In tmux, Ctrl+B [ drops you into scrollback with search.
When to Use tmux, When Not To
tmux is not a panacea. There are scenarios where it’s excessive or even harmful.
tmux makes sense when:
- multiple admins work on the same server simultaneously;
- sessions are long-running (monitoring, deploys, debugging);
- you need reliable copy-mode and scrollback;
- automation through the
tmuxCLI (CI/CD scripts that send commands into a session).
tmux is not needed when:
- one admin per server, short sessions;
- memory-constrained systems — a tmux server consumes more RAM than screen (though on modern machines this is negligible);
- you’re in a containerized environment with an ephemeral filesystem — the tmux config won’t survive container recreation without a volume.
Inside Docker, prefer docker exec -it <container> bash over running tmux inside the container. tmux in a container only makes sense for stateful services that require persistent shell access.
Production Commands
The default prefix is Ctrl+B. All commands follow it.
Creating and attaching:
Managing windows and panes:
Working with copy:
Scripted invocation (for automation):
Session persistence across reboot:
tmux won’t survive a reboot. If you need a live session after restart, use the tmux-resurrect plugin or a cron script that recreates sessions from a state file.
Minimal ~/.tmux.conf for production:
After editing the config: tmux source-file ~/.tmux.conf.
Bottom Line
tmux replaced Screen on production not because it’s “newer,” but because its sessions don’t get lost on disconnect, copy works without workarounds, and CLI scripting gives predictability for automation. The trade-off is slightly higher memory usage and the need to maintain a single config across all servers. For our stack of 12 production servers with three admins each, it paid off in the first week. The task is to write an English IT blog post for “Lead DevOps” blog, parallel to the Russian draft provided. The English version should be the original article (not a word-for-word translation), same structure and facts, practical tone, same commands and tables.
Key constraints:
25 - bpftrace: one-liners that replace strace in production
strace halts a process on every syscall. On a live server at 2000 RPS, that means timeouts and alerts. bpftrace runs through eBPF in the kernel — tracing happens in parallel, without stopping anything. The overhead difference is orders of magnitude.
Installation
Full probe coverage requires debug symbols:
Check available probes:
bpftrace Syntax in 60 Seconds
One-liner format:
Structure: what (probe) and what to do (action). Probe types:
| Type | Example | Description |
|---|---|---|
| kprobe | kprobe:do_sys_openat2 | kernel function entry |
| kretprobe | kretprobe:do_sys_openat2 | kernel function return |
| tracepoint | syscalls:sys_enter_openat | stable kernel tracepoint |
| usdt | usdt:/bin/python3:probe | user-level static trace |
| profile | profile:hz:99 | timer-based sampling |
Built-in variables available in actions:
Example — all execve calls:
exec: Who Is Launching Processes
Want to understand which process triggers fork/exec across the system:
Output during 10 seconds of monitoring:
Track a specific process and its children:
The filter /pid == 1234/ uses standard syntax. Without it, bpftrace catches everything.
open: Which Files a Process Opens
Filter by process name:
This is aggregation — counting how many times each file was opened. @ is the built-in variable for maps. Output after Ctrl+C shows a sorted table.
Monitor open failures (ENOENT, EACCES):
Network: Connections and Dropped Packets
Monitor outbound connections:
Dropped iptables packets:
Not all kprobes are available on every kernel. Verify with sudo bpftrace -l | grep nf_hook.
Aggregation by port — a common task:
Errors: getpid Does Not Exist
Familiar functions may be absent in bpftrace. This is not bash — different rules apply.
| Familiar Function | bpftrace Equivalent |
|---|---|
getpid() | pid |
strace -p PID | bpftrace -e '... /pid == N/ {...}' |
readlink /proc/PID/fd/N | nsecs, curtask |
Calling getpid() inside bpftrace produces a compilation error — the BPF program has no access to libc.
bpftrace cannot trace a process already running under active strace. They conflict at the ptrace level.
Error output — use strerror():
bpftrace vs strace: Overhead Comparison
strace uses ptrace(PTRACE_SYSCALL). On every syscall, the kernel stops the process, copies data to userspace, then resumes. This is synchronous.
bpftrace compiles a BPF program and loads it into the kernel. Tracing happens in kernel context without stopping the process. Data accumulates in a ring buffer and is read asynchronously.
Comparison on nginx, 5000 RPS:
| Method | p99 Latency | CPU Overhead | Observability |
|---|---|---|---|
| No tracing | 12ms | — | — |
| strace -p PID | 340ms | 18% | syscalls |
| bpftrace one-liner | 14ms | 0.3% | syscalls + aggregation |
For a quick check: strace -c -p PID gives a summary table of syscalls. bpftrace does the same via count() and hist().
When bpftrace Is Enough vs When You Need strace
bpftrace — for a system-wide view. Monitoring all processes, aggregation, heat maps, catching anomalies without affecting production.
strace — for deep-diving a specific request. Detailed log of every syscall with arguments and returns to reproduce a problem.
Three rules:
- Don’t know the process — use
bpftrace. - Know the PID and need detailed logs — use
strace -p PID. - On production under load —
bpftraceonly.
One-liners as aliases:
bpftrace goes further — kernel memory, CPU profiling, allocator debugging. For a basic start, these four one-liners cover most of what you need.
26 - ip: Network Setup and Diagnostics in CLI
When ifconfig returns nothing and configuring a route requires a separate command, that’s not a system bug. It’s iproute2 — the package that replaced net-tools in modern Linux distributions. The ip utility from iproute2 is the standard interface for managing the Linux network stack. It covers interfaces, addresses, routes, ARP cache, routing policies, and namespace isolation.
Why iproute2 replaced net-tools
net-tools (ifconfig, route, arp, netstat, nameif) originated in BSD and migrated to Linux in the 1990s. By the 2000s it became clear: they cannot handle VLAN, IPsec, QoS, multicast routing, or Policy Routing. Each task required a separate command with unrelated syntax.
iproute2 consolidated everything into one ip utility with subcommands. The Linux kernel communicates with the network subsystem via netlink sockets — ip talks to them directly, while ifconfig parses /proc/net/. In systemd-based distributions (RHEL 7+, Ubuntu 16.04+, Debian 9+) ip is installed by default. net-tools remains in repositories for compatibility, but kernel developers haven’t added new functionality since 2001.
Install if needed: apt install iproute2 or yum install iproute.
ip link: interface state and management
ip link operates at L2 — listing and controlling interface state.
Output shows index, name, MAC address, MTU, state (UP/DOWN), and error/packet counters.
Bring an interface up or down:
Taking down an interface severs connectivity. When working remotely, wrap in a script with a timeout and auto-recovery.
Set MTU, change MAC, or rename:
Create virtual interfaces (VETH pair for namespace or bridge):
Delete an interface:
| Flag | Purpose |
|---|---|
| show | display interfaces (shorthand: ip l) |
| set | modify interface parameters |
| add / del | create or delete virtual interface |
| master | attach interface to a bridge |
ip addr: address binding and diagnostics
ip addr manages IP addresses (L3).
Add an address:
Add a secondary address (alias) on the same interface:
Secondary addresses in Linux are not aliases in the ifconfig sense — they are part of a single address entity. The command ifconfig eth0:0 created a pseudodevice with a separate name; ip works differently.
Remove an address:
Flush all addresses from an interface:
Useful during reconfiguration: clear old addresses and assign new ones without restarting the service.
Specify scope and label:
scope host — address only for local sockets; scope global — routable.
| Subcommand | Action |
|---|---|
| add | assign address |
| del | remove address |
| show | display addresses |
| flush | clear interface addresses |
ip route: default and static routes
ip route works with the routing table.
Add a default route (gateway):
Add a specific route:
Route to a host via direct ARP (no route, L2 only):
Remove a route:
Replace a route (updates if exists, creates if not):
Get the route the kernel will choose for an address:
Add a route to a different table (default is table 254):
| Subcommand | Purpose |
|---|---|
| show / list | display routing table |
| add | add a route |
| del | remove a route |
| replace | modify or create a route |
| get | show route to an address |
| flush | clear route cache |
ip neigh: ARP/NDP cache
ip neigh manages the neighbour table — ARP for IPv4, NDP for IPv6.
Add a static ARP entry:
nud (Neighbour Unreachability Detection) defines the state:
- permanent — entry never expires
- noarp — managed by protocol but not removed
- reachable / stale / delay / probe — automatic states
Remove an entry:
Flush all neighbours on an interface:
After changing a gateway MAC address, flushing the ARP cache speeds up connectivity recovery: ip neigh flush dev eth0.
ip rule: routing policies
ip rule determines which routing table is used for a packet.
Standard output:
Add a rule for source IP:
Rule for incoming interface:
Remove a rule:
Rules are checked in order (priority). Low number means high priority. Add rules with priority between existing ones if ordering matters.
| Action | Purpose |
|---|---|
| from | source IP or CIDR |
| to | destination IP or CIDR |
| iif | incoming interface |
| lookup | routing table |
| prio | numeric priority |
ip maddr: multicast addresses
ip maddr displays and manages multicast groups on an interface.
Add an interface to a multicast group:
Remove:
Unlike unicast, multicast addressing is used in broadcast domains, routing protocols (OSPF, RIP), discovery services, and streaming. In most tasks this command is unnecessary, but when configuring clusters or monitoring via specific protocols — it will be required.
ip netns: network stack isolation
ip netns creates isolated network namespaces. Each namespace has its own interfaces, addresses, routes, ARP table, and rules.
Create a namespace:
Run a command inside a namespace:
Bring up an interface in a namespace:
Move a VETH interface into a namespace:
Delete a namespace:
List namespaces:
Containers (docker, podman, LXC) use netns underneath. If a container has no network — check the host namespace: ip netns exec <container_pid> ip addr.
ifconfig, arp, route — legacy compatibility
net-tools is formally available in all major distribution repositories. The source code is unmaintained, but packages persist for compatibility.
| Legacy command | Equivalent ip | Status |
|---|---|---|
| ifconfig | ip addr, ip link | deprecated |
| route -n | ip route | deprecated |
| arp -a | ip neigh | deprecated |
| netstat -tulpn | ss -tulpn | deprecated |
| nameif | ip link name | deprecated |
ss from iproute2 replaces netstat — faster and provides more socket information.
Network initialization scripts in old distributions may rely on ifconfig. On modern systems, systemd-networkd, NetworkManager, and cloud-init use ip directly or through their own abstractions.
If ifconfig is missing in a fresh distribution — that’s expected. Configuration via ip covers all current scenarios: from address assignment to complex routing policies and service isolation via namespaces.
27 - nslookup and drill: DNS resolution in terminal
The server won’t resolve a domain, but pings fly through. No familiar dig at hand — the BIOS is already loading a minimal busybox. Or on a host without bind-tools. nslookup and drill fill this gap: the first one is built into almost everything, the second gives more context when debugging.
nslookup: interactive and one-liner modes
nslookup ships with bind-utils and isc-dhcp-client. It works in two modes.
One-liner query:
Interactive mode starts with no arguments. Typical session:
Switching servers inside a session only changes the resolver for that query. If you need a permanent resolver — edit /etc/resolv.conf.
DNS record types in queries
By default nslookup queries A records. For other types use set type=:
| Record type | Purpose | Example output |
|---|---|---|
| A | IPv4 address | 93.184.216.34 |
| AAAA | IPv6 address | 2606:2800:220:1:: |
| MX | Mail exchanger | 10 mail.example.com |
| TXT | Text records, SPF | v=spf1 include:_spf.example.com ~all |
| NS | Authoritative servers | a.iana-servers.net |
| SOA | Start of Authority | serial 2005080901 |
| CNAME | Canonical name | example.com canonical name = www.example.com |
| PTR | Reverse resolution | 34.216.184.93.in-addr.arpa name = example.com |
One-liner equivalent — -type= flag:
ANY queries are often blocked at the resolver level. The recursor returns SERVFAIL or an empty response. Do not rely on ANY when troubleshooting.
drill: output with Resource Record type
drill is part of ldns. It returns results in classic DNS format with ANSWER, AUTHORITY, ADDITIONAL sections:
drill output is more readable when tracing a CNAME chain:
Without the @server flag, drill reads the resolver from /etc/resolv.conf.
DNSSEC validation
drill checks the DNSSEC trust chain:
The -S flag requests the DS record higher in the chain and validates the signature. On an invalid chain:
nslookup does not validate DNSSEC — it only sends queries with the DO flag (include RRSIG in the response). For full validation you need drill or delv.
NXDOMAIN and SERVFAIL: reading response codes
First step on any error — look at the response code.
NXDOMAIN (code 3) — the domain does not exist. Source: authoritative server for the zone. If dig +short returns nothing, and nslookup says ** server can't find example.invalid, that’s NXDOMAIN. Causes: typo in the domain, stale CNAME, deleted zone.
SERVFAIL (code 2) — the resolver couldn’t answer. Causes: broken DNSSEC validation, exceeded timeout, circular reference in NS records, overloaded authoritative server. nslookup shows ** server can't find example.com: Server failed.
REFUSED (code 5) — the recursor refused to answer. Usually ACL on the DNS server or rate limiting.
Key flags for nslookup and drill
nslookup
| Flag | Effect |
|---|---|
-type=RR | Record type (A, MX, TXT, ANY) |
host | Redirect to specified server |
-port=53 | Non-standard port (e.g. 5353 for mDNS) |
-timeout=5 | Timeout in seconds |
-retry=3 | Number of retries |
-vc | TCP instead of UDP |
drill
| Flag | Effect |
|---|---|
@server | Server to query |
-Q | Quiet mode, answer only |
-T | Show response time |
-S | DNSSEC validation |
-D | Force DNSSEC (query with DO flag) |
-p port | Non-standard port |
-t timeout | Timeout in seconds |
-T is useful for comparing latency between resolvers:
Installation
In minimal busybox images you already have a simplified nslookup. The full feature set is available after installing dnsutils.
28 - journalctl: Filtering and Formatting systemd Logs
Logs disappeared. Server rebooted, and the familiar less /var/log/syslog returns nothing. On modern distros with systemd, logs are collected by journald and read with journalctl. Without knowing its filters, system debugging turns into guesswork.
Why Logs Disappear After Reboot
By default, journal stores data in /run/log/journal/ — a tmpfs that wipes on reboot. To make logs survive reboots, create the directory:
Then restart systemd-journald:
Check current location and size:
Fresh CentOS/RHEL 8+ and Fedora create /var/log/journal automatically. Debian and Ubuntu typically do not.
Filtering by Unit and Time Range
The most common case — logs for a specific service:
Combine multiple units by repeating the flag:
Time filters are for incident debugging:
Night crash? Look at logs from that period, not the entire buffer.
If a time filter returns empty output, check the timezone. journalctl stores timestamps in UTC, but --since interprets local time.
Filtering by Priority
Log levels match syslog:
| Level | Number | Description |
|---|---|---|
| emerg | 0 | System unusable |
| alert | 1 | Immediate action required |
| crit | 2 | Critical condition |
| err | 3 | Error |
| warning | 4 | Warning |
| notice | 5 | Normal but significant |
| info | 6 | Informational |
| debug | 7 | Debug-level messages |
The -l flag shows full hostnames instead of truncated ones.
Kernel and Boot Logs
The kernel sends its messages separately. The -k flag replaces dmesg:
List all boots:
Output:
Select a specific boot:
For boot analysis, use systemd-analyze:
Regex Search
Piping journalctl output to grep loses metadata. Use -g (–grep) instead:
The -g flag supports basic regex. For complex conditions, combine with --since:
This keeps the unit and time selection intact, then filters by pattern.
Follow Mode (-f)
Analogous to tail -f for journald. Unlike watching a log file, follow works with any filter:
Ctrl+C stops follow in a terminal. From scripts, wrap with timeout or send a signal.
Run -f in a separate tmux/screen pane. If the pane closes, logs keep going to journald — you won’t lose data.
Flags combine with AND: -u nginx -p err shows errors from nginx only. For OR across units, use journal fields:
Other useful fields:
Output Formats
By default, journalctl paginates output. For scripts and piping to jq, you need machine-readable format:
| Flag | Description | Use Case |
|---|---|---|
-o short | Classic syslog | Default |
-o short-iso | ISO 8601 timestamps | SIEM logging |
-o short-precise | Millisecond precision | Precise timing |
-o verbose | All fields | Maximum detail |
-o json | JSON Lines | jq, Splunk, ELK |
-o cat | MESSAGE field only | Minimal output |
-n 100 limits output to the last 100 lines. --no-pager disables pagination for scripts.
Cleanup and Size Management
journald rotates logs by size and time. Configure in /etc/systemd/journald.conf:
Apply without restart:
Free up space manually:
--vacuum-* only removes files exceeding the limit. To free space reliably, increase SystemMaxUse and restart journald.
Common Errors
journalctl: cannot open files — insufficient permissions. Add yourself to the systemd-journal group:
Logs empty after reboot — persistent storage not configured (see first section).
journalctl hangs — huge buffer. Start with -b or limit with --since.
No unit logs — check that the unit actually ran:
journalctl is built for fast searching. Don’t read logs manually — filter from the start.
29 - OpenSSL: TLS Certificate Verification and Parsing in CLI
Certificates expiring on prod at the worst moment — a familiar story. OpenSSL answers TLS certificate questions faster than any marketplace checker. Here are the key scenarios without the fluff.
Basic Certificate Parsing
The first command for any diagnostics is the text dump:
Output shows Subject, Issuer, validity dates, signature algorithm, and public key. For a quick summary without the wall of text:
The -in flag accepts a file path. Certificates downloaded from browsers usually come in PEM or DER format. OpenSSL handles both, but DER requires an extra flag:
-noout suppresses the base64 block from output. Useful when you need only structured data, not a copy of the certificate.
Expiration: dates and Overdue Checks
For monitoring, getting only the dates is more convenient:
Typical output:
For automation, extracting the timestamp and calculating the difference is cleaner:
If days_left is negative, the certificate has already expired.
For checking multiple hosts from inventory, a one-liner works well:
-servername sends SNI — without it some hosts return the default certificate.
Chain Verification: s_client and verify
Connect and display certificates:
Output includes the chain from leaf certificate to root CA. To filter certificates only:
Verify chain against the system store:
If verification fails with error 20 at 0 depth lookup, the intermediate CA is missing. A common cause is incorrect chain configuration on the server.
openssl verify uses the system store by default. In Ubuntu this is /etc/ssl/certs/ca-certificates.crt, in Alpine it is a separate ca-certificates package. If verification fails, check that the package is installed.
Quick check without saving to file:
Output Verify return code: 0 (ok) means success.
To stay out of interactive mode and fail the process on a bad chain:
Force a protocol or cipher when you need to confirm the server still accepts a specific handshake:
For an internal CA, pass the bundle explicitly. Without -CAfile, a private PKI typically returns Verify return code: 21 (unable to get local issuer certificate):
Extracting CN and SAN
Common Name extracts directly:
But CN has not been sufficient for years — modern certificates use Subject Alternative Names (SAN). OpenSSL 1.1.1+ pulls them cleanly:
Output:
To get only the DNS names:
If SAN is missing (old certificate), browsers fall back to CN. When checking API endpoints, this explains why curl complains but the browser opens the page.
Comparing Expiration Across Multiple Hosts
A script for checking a host list serves as a practical monitoring foundation:
Typical output:
Negative values are expired certificates. In production, wrapping this in cron with chat notifications at a 30-day threshold is practical.
Quick Flag Reference
| Command | Flag | Purpose |
|---|---|---|
x509 | -text | Full text dump |
x509 | -noout | Suppress base64 block |
x509 | -dates | NotBefore, NotAfter |
x509 | -subject | Subject (CN, O, OU) |
x509 | -issuer | Issuing CA |
x509 | -fingerprint -sha256 | Certificate fingerprint |
x509 | -enddate | Expiration date only |
x509 | -ext subjectAltName | Alternative names |
s_client | -connect host:port | TLS connection |
s_client | -servername name | SNI (required for vhost) |
s_client | -showcerts | Display full chain |
s_client | -quiet | Skip interactive mode |
s_client | -tls1_2 / -tls1_3 | Force a TLS version |
s_client | -verify_return_error | Non-zero exit on validation failure |
verify | -CAfile path | Trusted CA file |
verify | -partial_chain | Accept partial chain |
openssl s_client does more than read certificates. With -starttls smtp or -starttls pop3 it checks mail servers. -http retrieves HTTP headers over TLS. Useful for diagnosing miTM filters.
All commands work out of the box in any Linux distribution. No dependencies beyond OpenSSL itself — the tool is present on every server. If missing, install in seconds: apt install openssl or apk add openssl.
30 - ProxyJump and bastion hosts via ~/.ssh/config
Sometimes a server sits in a private network with no public IP. The only entry point is a bastion host with a public address. Typing ssh -J user@bastion user@private every time gets old fast. Here’s how to configure everything in ~/.ssh/config so you can reach private networks in one command.
Why you need a bastion host
A bastion (jump host, jump box) is an intermediate server with public access that proxies connections to infrastructure without external IPs. The typical topology:
The bastion doesn’t need to be hardened like a fortress — it’s just an open relay point. Access control lives on SSH keys and, if needed, security groups or firewall rules.
The bastion host is not a terminal destination — it only proxies traffic. You don’t need to run a VPN or additional services on it.
ProxyJump — the modern syntax
-J (ProxyJump) arrived in OpenSSH 7.3. The parameter takes a host in [user@]host[:port] format and spins up a SOCKS5 proxy through the specified node.
Basic invocation:
Authentication on both hosts via keys. If the username matches, you can omit it:
With a port other than 22:
The same thing in the config file:
After that, ssh private-server connects through the bastion automatically.
ProxyCommand — the classic approach
ProxyJump is a wrapper around ProxyCommand. When you need more control or are working with older OpenSSH, use ProxyCommand directly.
-W forwards stdin/stdout to the target host. In the config:
The difference from ProxyJump is minimal, but ProxyCommand lets you inject variables, conditions, and command chains.
Multiple hops in a row
A chain of two bastion hosts:
In the config:
OpenSSH connects hosts sequentially: laptop → bastion1 → bastion2 → private-server. Make sure your keys are present on each node.
For complex scenarios, ProxyCommand with nc (netcat) gives more flexibility:
The chain works, but each hop adds latency. For interactive work, more than two hops signals a network architecture problem.
Complete config example
ForwardAgent yes on the bastion lets your agent forward keys further down the chain. Don’t enable it if you don’t trust the bastion machine.
LocalForward in the example tunnels PostgreSQL from the private server to local localhost:5433. Useful for connecting IDEs or psql.
Verify the config parses without errors:
If you see the correct values — the config was picked up.
Common mistakes and solutions
Connection timeout during ProxyJump
Verify the bastion is reachable directly:
If that fails — the problem is in the network, not the config.
Permission denied (publickey) on bastion
Make sure the key is loaded in the ssh-agent:
If the agent is empty — add the key and test with ssh -vT bastion.
Works one way, not the other
ProxyJump tunnels TCP. ICMP (ping) won’t pass through. Check connectivity with nc -zv host port or ssh -v.
Agent refused operation during agent forwarding
Check the SSH_AUTH_SOCK variable:
If empty — start the agent:
Slow connection through a chain
Check MTU. Sometimes MTU in VPN/LAN is smaller than required for TCP-over-TCP. Add to the config:
For very slow links, try compression:
ProxyJump covers 90% of use cases. If you need visualization or a UI manager — look at ssh-config tools or Terminator, but for console work ~/.ssh/config with ProxyJump is sufficient.
31 - auditd: file access and syscall logging
Linux doesn’t write every access to /etc/shadow or every unlink call to syslog. For incident investigation and compliance this is critical. auditd solves this: the Linux Audit kernel subsystem records system calls, file access, and more.
Installation and Startup
auditd comes in the audit package available in any distribution.
After installation, start the service via systemd.
Check status and current rules:
On RHEL-based distributions with SELinux enabled, you may need to adjust policies for auditd to work with non-standard paths. Standard installation usually covers most cases.
File Monitoring: the -w Flag
The -w flag adds a watch rule for a file. By default it tracks open, read, write, truncate, chmod, and chown.
| Flag | Meaning |
|---|---|
-w | path to watch |
-p | permissions: r(read), w(write), x(execute), a(append) |
-k | keyword for searching in logs |
Verify rules:
Remove a rule by key:
Rules added via auditctl don’t survive reboots. For persistence, write rules to /etc/audit/rules.d/:
On RHEL, rules load from /etc/audit/audit.rules via the augenrules script. Same works on Debian/Ubuntu.
System Call Logging: the -S Flag
The -S flag records the specified system call for all processes or with filters.
Check available system calls with ausyscall --dump. Not all calls are available on every architecture — on x86_64 some go through the compat layer.
Excessive syscall monitoring generates massive log volume. On production servers, limit rules with filters.
Combined rule — syscall plus path:
Filtering by UID and Executable
Without filters, rules apply globally. Add conditions for targeted monitoring.
Core filter fields:
| Flag | Description | Example |
|---|---|---|
-F | field to compare | -F uid=1000 |
arch | architecture (b32/b64) | -F arch=b64 |
uid | real UID | -F uid=33 |
euid | effective UID | -F euid=0 |
exe | full path to executable | -F exe=/bin/bash |
perm | access permissions | -F perm=awx |
Combine filters into a chain via -a:
Reading Logs: ausearch
Logs live in /var/log/audit/audit.log. Binary format, read with ausearch.
Useful output formats:
For automation, use -if (input file) — read from a dump instead of the live log:
Reading Logs: aureport
aureport aggregates logs into readable reports.
Typical aureport -s output:
Quick investigation combo — summary then details:
For SIEM or ELK ingestion, convert logs to JSON or text:
auditd doesn’t require complex setup to start capturing critical events. Install the package, add a few rules with keywords, and get comfortable with ausearch and aureport for log review.
32 - logrotate: automatic log rotation and archiving
Application logs fill up disk space within a week, and manually running rm *.log is a recipe for trouble. logrotate handles this automatically: it rotates, compresses, and deletes old files on a schedule. Let’s see how it works and how to set it up in five minutes.
How It Works
logrotate runs daily through cron. The default config lives in /etc/logrotate.conf, and additional configs are included from /etc/logrotate.d/. During rotation, the current file gets renamed, a new empty one is created, old copies are compressed and numbered.
The cycle looks like this:
The mechanism relies on rename or mv, so the process must hold the file descriptor open. If rotation isn’t picked up by the application, you end up with duplication or an empty log.
Configuration Structure
Files from /etc/logrotate.d/ override global values for specific logs. The format is straightforward:
Directives inherit from the global config unless overridden. You can specify multiple paths separated by spaces or use a glob pattern — convenient for rotating all application logs in one block.
Key Directives
Basic parameters covering 90% of use cases:
| Directive | Purpose | Example |
|---|---|---|
rotate N | number of copies to keep | rotate 7 |
size N | size threshold to trigger rotation | size 100M |
missingok | not an error if file is missing | — |
notifempty | skip empty logs | — |
compress | compress old logs (default gzip) | — |
dateext | use date instead of number in name | — |
dateformat | date format | dateformat -%Y%m%d |
postrotate ... endscript | commands after rotation | reload service |
prerotate ... endscript | commands before rotation | prepare directories |
size overrides weekly/monthly/daily when set. Rotation happens when the file reaches the specified size AND the interval has passed. So size 100M with weekly means rotation no earlier than a week and only if the file exceeds 100M.
dateext is incompatible with long filenames on some filesystems. If the log name plus date suffix exceeds 255 bytes — logrotate will fail.
Example Config for an Application
Assume your myapp service writes to /var/log/myapp/. Config:
sharedscripts ensures postrotate runs once for all rotated files, not once per file. Without it, the script executes as many times as files match the rotation criteria.
For Python applications using standard library rotation:
su changes the rotation process owner — useful when the application runs under a separate user with restricted log permissions.
Debugging and Manual Execution
Dry-run mode shows what would happen without making changes:
Output contains each decision: which file gets renamed, which gets compressed, which commands run. Look for renaming and running postrotate script lines.
Force rotation bypassing the schedule:
-f ignores the last rotation time and size conditions. Combining with -d is a safe way to verify before production:
To rotate a specific log outside the schedule but respecting conditions — use the state file:
Syntax check without execution:
If there are no errors — output is empty (without -d) or shows the action plan (with -d).
Do not edit /var/lib/logrotate/status manually in production without understanding the format. One mistake — and logrotate decides rotation already happened, skipping all files until the next cron run.
Default cron setup:
On most distros you don’t need to touch this file. If you need rotation more often than once a day — add to /etc/cron.hourly/ or write a separate cron job.
33 - sshd_config: baseline for a test stand
SSH access to a test stand often gets opened in a hurry, and then the logs fill with brute-force attempts. A baseline sshd_config that blocks common attack vectors fits into five parameters and twenty minutes.
Why Change Defaults
Distribution-provided sshd ships with permissive settings: root login via password, no user restrictions, three authentication attempts. On a test stand this is tolerable until the logs show:
A local network is not a trusted network. Default configuration means risk and noise in monitoring.
Checking Current Values
Before editing, inspect what’s already set:
Output shows the actual values sshd will use at startup, including parameters from Match blocks.
Four Parameters for a Test Stand
| Parameter | Value | Rationale |
|---|---|---|
PermitRootLogin | no | Root should not log in directly |
AllowUsers | devops admin | Whitelist, everyone else rejected |
PasswordAuthentication | no | Keys only, passwords disabled |
MaxAuthTries | 3 | Block after three failures |
Changes apply after systemctl reload sshd. Make edits via ssh -t user@host "sudo nano /etc/ssh/sshd_config" so you do not lose your session on a mistake.
PermitRootLogin
Direct root login is the first thing attackers brute-force. Disable it:
If you need root access, log in as a regular user and escalate via sudo. This logs your actions to auth.log.
AllowUsers
A whitelist excludes everyone not explicitly added. If a user is not in the list, sshd returns Permission denied before asking for a password.
Specify users space-separated. For groups use AllowGroups. Both parameters support patterns: AllowUsers devops@10.0.0.* restricts login by subnet.
PasswordAuthentication
Keys cannot be brute-forced. Switch the setting:
Before disabling, confirm your public key is in ~/.ssh/authorized_keys on the stand. Otherwise you lock yourself out.
MaxAuthTries
Protection against brute force. After three failed attempts the connection drops:
Setting below one disables the limit. Use 3–5 depending on your network’s reliability.
Validating Configuration
Always validate syntax after editing:
Empty output means sshd will start with the new parameters. Any error prints to the screen.
Then apply:
Common Mistakes
Editing the wrong file. On some distributions sshd_config lives in /etc/ssh/sshd_config.d/. Include files are read in alphabetical order. The default /etc/ssh/sshd_config may be overwritten on package updates — put custom parameters in a .conf file with a meaningful name instead.
Spaces after the parameter. Syntax requires a space between key and value:
Comments instead of parameters. The line #PasswordAuthentication no is a comment, sshd ignores it. Remove the # or add a new line.
Match blocks override global settings. If a Match User root block exists at the end of the file, it may restore PermitRootLogin yes for that user. Check sshd -T output.
Additional
This covers a test stand. Production adds ClientAliveInterval 300, ClientAliveCountMax 2 for keepalive and X11Forwarding no if graphics are not needed. But deploying a minimal baseline takes four parameters and a couple of commands.
34 - systemd-run: Run Services Without Unit Files
Sometimes you need to run a process under systemd’s control without writing a unit file — maybe you’re in a container without systemd, on someone else’s machine, or just need a quick one-off. That’s where systemd-run comes in.
Why systemd-run
The tool creates a transient unit — a unit that exists only in systemd’s memory, with no file on disk. This is useful when you need:
- Resource control (cgroup) over an ad-hoc process;
- The process to survive terminal closure;
- Isolation in a separate slice.
In practice, it’s a wrapper around systemctl start for units nobody persists.
Basic Syntax
Simple run:
The process lands in user.slice, runs in the background, and is manageable via systemctl. Verify:
You’ll see something like run-u1234.service in the output.
Key Flags
| Flag | Purpose |
|---|---|
--scope | Create a scope unit instead of service (parent process stays in scope) |
--unit=NAME | Assign a custom unit name instead of auto-generated |
--uid=USER | Run as the specified user |
--gid=GROUP | Run as the specified group |
--nice=N | Nice value (-20 to 19) |
--property=KEY=VALUE | Pass a unit property (CPUAccounting, MemoryMax, etc.) |
--chdir=PATH | Working directory |
--setenv=VAR=VALUE | Environment variable |
--tmpfs=PATH:OPTIONS | Mount a tmpfs at the specified path |
Flags combine freely. Example with resource limits:
Isolation: CPU, RAM, Root via cgroup
The main value is resource control through cgroup v2.
Memory limit:
If the process exceeds the limit, systemd kills it (OOM via cgroup).
CPU limit:
50% of one CPU core. On multi-core systems, calculate percentages against total units.
Isolation with alternate root:
RootDirectory requires a correct directory structure inside. Without it, the process fails to start with “Failed to pivot root”.
Mount a tmpfs:
Useful for processing temporary files with a size constraint.
Combined example:
All process resources are tracked in cgroup, visible via systemd-cgtop and /sys/fs/cgroup.
Comparison with nohup, setsid, chroot
| Tool | Resources | Isolation | Survives logout | Management |
|---|---|---|---|---|
nohup | No | No | Yes | No |
setsid | No | No | Yes | No |
chroot | No | Filesystem only | Depends | No |
systemd-run | cgroup | CPU, RAM, I/O, user | Yes | systemctl |
nohup and setsid work fine for background scripts. But if you need memory limits or CPU priority — cgroup is non-negotiable.
chroot solves only filesystem isolation. You can’t do systemd-run --chroot, but you can combine them:
Here chroot handles filesystem isolation, systemd-run limits resources.
Pitfalls: scope vs service
By default, systemd-run creates a service unit. The difference:
- service — full unit, registered with systemd, has dependencies, managed normally.
- scope — tied to the parent process. Specified with
--scope. Parent dies — scope dies.
In day-to-day practice, service works for long-lived processes:
Scope is useful for grouping related processes:
If you run with --scope and close the terminal — processes die. Without --scope — they survive.
Transient Unit Limitations
Transient units don’t persist across reboots. Obvious, but there are other gotchas:
- No dependencies. Units lack
After=,Wants=,Requires=. The process starts immediately. If you need sequencing — chain commands in a script withsleep. - Limited property set. Not all unit file fields are available via
--property. For example,Restart=alwaysis not supported in transient units (tested on systemd 254). - Harder to debug. No file — no
systemctl edit. All parameters are only in the command line.
For long-running services with dependencies and restart policies — a unit file remains the right choice. systemd-run is for one-off tasks, prototyping, and resource limiting on the fly.
35 - curl: HTTP Debugging in CLI
cURL is the standard tool for debugging HTTP in the terminal. It ships out of the box on Linux and macOS and is present in most Docker images. Need to quickly check an API, inspect response headers, or trace a redirect issue — one command line is enough.
Basic Debug Flags
The most common scenario: get a response and see what the server returned. The -i flag prints headers before the body, -v enables verbose mode with connection details.
The difference: -i shows headers plus body, -v adds DNS resolution, TLS handshake, and debug info before the request.
A HEAD request checks resource availability and metadata without downloading the body:
HEAD does not guarantee the server supports ranges or caching — this depends on server configuration.
Methods and Request Body
By default, curl sends GET. For other methods, use the -X flag.
Request body is passed via -d. For JSON APIs, a typical pattern:
Multiline JSON is easier to read in a heredoc when the body is large:
To send form data or data from a file:
Custom Headers
The -H flag adds or overrides a header. You can use -H multiple times.
Overriding the Host header is useful when debugging virtual hosts or proxies:
To remove a default header, use -H "Accept:" — empty value after the colon.
Auth and Certificates
Basic HTTP auth via -u in user:password format:
For Bearer tokens, use a header:
-u sends credentials in plain text if TLS is not used. Always use HTTPS for production servers.
When working with self-signed certificates, the -k flag disables verification:
For known hosts and pinned certificates:
Timeouts and Saving Response
By default, curl waits indefinitely. For scripts and monitoring, set limits:
| Flag | Purpose |
|---|---|
--max-time N | total timeout in seconds |
--connect-timeout N | connection timeout |
Saving the response:
Redirect stdout to pipe the response body:
Redirects
Curl does not follow redirects by default. The -L flag enables automatic redirection:
To debug a redirect chain:
-L limits redirect depth (default 50). Infinite redirect loops are a common cause of curl hanging.
A flag combination for a full picture when debugging an API:
This command shows request and response headers, saves the body to a file, and enforces a 10-second timeout.
36 - ethtool: network interface diagnostics and tuning
Network issues hide well — interface is up, IP is assigned, iptables is quiet, yet packet loss or micro-freezes only surface under load. ethtool gives direct access to hardware state, driver behavior, and offload mechanisms that neither ip nor netstat expose.
Basic Output: Link State
Installation is straightforward:
Running ethtool without flags prints a summary:
First thing I check when network complaints come in — Link detected. If no, cable or transceiver is dead. Speed and Duplex tell you whether the link dropped to 100Mb/s or half-duplex.
Driver and Hardware: -i and -a
-i shows driver information:
Firmware version matters for Intel and Broadcom — older versions have known bugs. If you do not see error counters you expect, update firmware first, not the driver.
-a displays auto-negotiation and pause settings:
Disabling flow control on one end of a link without agreement on the other causes packet loss during bursty traffic. If the switch has PAUSE disabled — disable it on the host too.
Speed and Duplex: -s and autoneg
-s changes interface parameters. To set fixed speed and duplex, disable auto-negotiation first, then specify the values:
After changing parameters, verify the link again — not all cards re-negotiate cleanly without an interface bounce.
| ethtool flag | Purpose |
|---|---|
speed N | Speed in Mb/s (100, 1000, 10000, …) |
duplex full|half | Duplex mode |
autoneg on|off | Auto-negotiation control |
port tp|fiber|aui|bnc|mii | Port type (not available on all cards) |
advertise N | Bitmask of modes for auto-negotiation |
All flags combine in a single call. Persisting across reboots requires writing to /etc/sysconfig/network-scripts/ifcfg-eth0 on RHEL or using a systemd override:
Offload Flags: -k, -K and Common Pitfalls
-k shows current offload flags, -K modifies them:
The [fixed] label means the flag is hardware-enforced and cannot be changed.
Disabling offload flags is a frequent cause of VPN, monitoring, and virtualization issues:
GRO (generic-receive-offload) and TSO (tcp-segment-offload) work as a pair. Disabling one without the other causes fragmentation at the kernel level — CPU usage spikes for no good reason.
Statistics: -S and Packet Loss Investigation
-S outputs driver statistics. Format and available counters vary by driver:
For Intel drivers (ixgbe, i40e), relevant counters include:
rx_fifo_errors and tx_fifo_errors indicate congestion. If they grow despite normal CPU utilization, the problem lies with memory or the bus.
For scripted problem detection:
Run periodically via cron — peak loss spikes become impossible to explain retroactively.
Comparison with ip link: What ethtool Covers
| Task | ip link | ethtool |
|---|---|---|
| Bring interface up/down | ip link set eth0 up/down | no |
| MAC address | ip link show eth0 | no |
| MTU | ip link set eth0 mtu 9000 | no |
| Speed/duplex | no | ethtool -s eth0 speed 1000 duplex full |
| Auto-negotiation | no | ethtool -s eth0 autoneg off |
| Offload flags | partially via ethtool -k | ethtool -K eth0 tso off |
| Error statistics | ip -s link show eth0 | ethtool -S eth0 (more detailed) |
| Driver information | no | ethtool -i eth0 |
| Wake-on-LAN | no | ethtool eth0 (shown in output) |
ethtool does not replace ip, it complements it. Network configuration stack: ip link → ip addr → ethtool → tc.
Troubleshooting: Link/Duplex Mismatch
Classic scenario: server and switch failed to agree on parameters. Symptoms — link is up, pings work, but under load there are sudden drops.
Diagnostic sequence:
Always negotiate both ends. If the switch runs fixed mode without autoneg and the host has autoneg enabled — the 802.3 standard mandates fallback behavior, but vendors implement it poorly.
If the interface still drops packets after parameter alignment, investigate:
- Driver and firmware: update the network card firmware
- Cable or transceiver: a 10G SFP module plugged into a 1G port without auto-negotiation guarantees loss
- RSS and IRQ issues:
cat /proc/interrupts | grep eth0, check IRQ balancing
37 - lsof: which processes listen on port and hold file
Service won’t start — port 8080 is already bound. You dig into who’s holding it, and discover the same process has your config file open while you’re trying to edit it. lsof answers both questions: which processes opened which files and sockets.
Listening Ports
The classic task — find who is listening on a specific port.
| Flag | Effect |
|---|---|
-i | Show internet sockets |
-n | Skip DNS resolution (show IP instead of hostname) |
-P | Skip port-to-service conversion (show 80 instead of http) |
Without -n -P, lsof wastes time on DNS lookups and resolves ports through /etc/services. On production hosts that’s unnecessary seconds.
For a specific port:
To find what process is listening on port 443, lsof -i :443 -n -P is sufficient. Output shows PID, user, and socket type (IPv4/IPv6, TCP/UDP).
To filter by protocol:
Processes in a Directory
Need to see which processes are working with files inside a directory? lsof +D recursively walks the directory and lists all open files.
On directories with thousands of files (for example, /tmp), this command runs slowly. It traverses the filesystem rather than querying the kernel — this is an O(n) operation.
For non-recursive search (files directly in the directory, no subdirectories), use find + xargs:
The result shows all processes holding open files from the specified directory. Typical scenario — cannot unmount a partition because someone is working with files inside it.
Single Process Inventory
When the PID is known, list all open files:
Output includes regular files, libraries (.so), sockets, and pipes. To grep by type:
REG — Regular file, DIR — directory, FIFO — named pipe, IPv4/IPv6 — network sockets. The TYPE in lsof output matches the type in /proc/PID/fd.
The reverse operation — find PID by file:
If the file is locked (log rotation fails, unmount does not work), this command shows the culprit.
Truncated Command Names
By default, lsof truncates command names to 9 characters. For long names (java, python) this may not be enough:
+c 0 means “no limit”. Useful when working with Java processes where the command line contains dozens of characters of classpath.
Quick Reference
| Command | Purpose |
|---|---|
lsof -i :PORT | Who listens on PORT |
lsof -i TCP | All TCP connections |
lsof -i UDP | All UDP connections |
lsof -p PID | Files opened by PID |
lsof +D DIR | Processes in directory |
lsof /path/to/file | PID that opened file |
lsof +c N | Limit command name to N characters |
lsof -u USER | All open files for user |
lsof -c CMD | Files for processes named CMD |
Common Issues
lsof not installed — minimal images require installation:
No read access to /proc — viewing other users’ processes requires root or group membership with access. Usually means running through sudo.
lsof hangs — kernel not responding to file descriptor requests (NFS issues, stalled filesystem). Ctrl+C and restart with a timeout.
lsof is one of those tools you return to every time you troubleshoot locked resources. Three commands cover 90% of the tasks: -i :PORT for ports, +D /path for directories, -p PID for processes.
38 - mc — MinIO Client S3 CLI
S3-compatible object storage is the default choice for buckets, backups, and static assets. When AWS CLI feels excessive and the web console is too clunky, MinIO Client (mc) fills the gap. This CLI tool works with any S3-compatible storage: MinIO, Yandex Cloud, AWS S3, Backblaze B2. Zero dependencies, works out of the box, configured in under a minute.
Installation
Download the binary and make it executable:
Verify the version:
For macOS, use Homebrew:
mc is a single static binary with no dependencies. Works perfectly in containers and minimal base images.
Adding an Alias
An alias is a named connection to an S3 endpoint. Without it, every command requires the full URL.
After this, myminio replaces the URL in all commands. List existing aliases:
In production, avoid storing keys in command history. Use environment variables: mc alias set prod ${S3_ACCESS_KEY} ${S3_SECRET_KEY} --api API-S3v4 and substitute via env.
For S3-compatible services with self-signed certificates:
Browsing and Navigation
List buckets:
Recursive listing with contents:
Object metadata:
mc ls without --recursive shows only top-level buckets. Use prefixes to navigate folders inside a bucket.
Working with Buckets
Create a bucket:
If the bucket already exists, mc reports an error. The --ignore-existing flag suppresses it:
Remove an empty bucket:
For non-empty buckets, use force:
Working with Objects
Print object contents to stdout without downloading:
Remove an object:
Recursive removal by pattern:
mc rm without --force asks for confirmation. Always use --force in scripts.
Find objects by criteria:
Combine with deletion:
Uploading and Downloading
Copy a local file to a bucket:
Multiple upload:
Download an object locally:
Copy between buckets (or across aliases):
Flags for flow control:
| Flag | Purpose |
|---|---|
--recursive | Process directories recursively |
--force | Overwrite without prompting |
--preserve | Keep file attributes (mtime, ACL) |
--if-not-exists | Skip already existing objects |
--disable-multipart | Upload in a single PUT request |
Mirroring
One-way directory synchronization:
Flags for production use:
| Flag | Purpose |
|---|---|
--overwrite | Replace changed files |
--delete | Remove files in destination missing from source |
--watch | Monitor for changes in real time |
--md5 | Verify MD5 after upload |
--delete is dangerous: it removes files in the destination that don’t exist in the source. Test with --dry-run or use --preserve for backup scenarios.
Dry-run — show what would happen without making changes:
Access Policies
Set public access on a bucket:
Common policies:
View current policy:
Generate a presigned URL (works for private buckets too):
Output contains the signed URL and expiration time.
Useful Flags
Global flags working with any command:
| Flag | Purpose |
|---|---|
--debug | Verbose HTTP request/response output |
--json | JSON output (easy to parse in scripts) |
--no-color | Disable colored output |
--insecure | Skip TLS certificate verification |
--config-dir | Config path (default ~/.mc) |
--limit | Rate limit (e.g., --limit 10MiB/s) |
JSON output for automation:
Progress for large file copies:
Shell Completion
Bash/zsh autocompletion saves time:
After enabling, type mc and press Tab twice to see available commands and aliases.
mc covers 90% of S3 tasks. For complex scenarios (versioning, lifecycle policies, encryption) — use the API or a Terraform provider. But basic bucket and object operations are faster with this tool than with any SDK.
39 - nftables: Modern Linux Firewall
Before changing nftables, make sure you have physical or console access to the server. A misconfigured input chain can block SSH and lock you out.
nftables replaced iptables in the Linux kernel starting with version 3.13. If you’re still writing rules in iptables style, it’s time to reconsider. nftables performs better, has built-in dual-stack IPv4/IPv6 support, and lets you manage the entire ruleset as a whole instead of entering commands one by one.
Why Migrate from iptables
iptables has fundamental design problems. Each table (filter, nat, mangle) is a separate rule set with its own semantics. There is no built-in support for simultaneous IPv4 and IPv6 handling—you end up writing two separate rule sets. Performance degrades with a large number of rules due to linear lookup.
nftables takes a different approach. All protocols work within a single data structure—the ruleset. The kernel compiles rules into efficient lookup structures. Atomic ruleset replacement eliminates race conditions during rule updates.
RHEL 8 and newer redirect iptables to nftables by default. Ubuntu 22.04 and later do the same. Check with update-alternatives --display iptables.
Basic Commands: Viewing and Flushing Rules
The first command to memorize:
Output shows all tables, chains, and rules. Without tables, output is empty—this is normal.
Create a table named filter for packet handling:
inet means the table handles both protocols. For IPv4-only use ip, for IPv6 use ip6.
Add a chain for incoming traffic:
Flags breakdown:
| Flag | Purpose |
|---|---|
type filter | Chain type—packet filtering |
hook input | Attachment point—incoming packets |
priority 0 | Processing order relative to other hooks |
policy accept | Default action—allow everything |
Flush rules in a chain:
Delete an entire table:
Adding and Deleting Rules by Handle
Create several rules and inspect their handles:
Output shows something like:
Handles are hidden by default. To work with specific rules:
The -a flag reveals the handle for each rule. Now you can delete by number:
Handles change whenever you add or delete a rule. If your script modifies the ruleset, save nft -a list ruleset output to a file for tracking.
Add a rule with priority before existing ones—at the chain start:
The add command appends to the end, insert adds at the start. For insertion at a specific position:
Atomic Ruleset Replacement
Adding rules one by one creates a window where some rules are active and others are not yet applied. This is unacceptable for production.
The solution: write the complete ruleset to a file and load it atomically:
/etc/nftables.conf is the standard location across most distributions. Now edit it:
Load it:
flush ruleset clears everything before loading. If you need to append to existing rules, remove this line.
Check without applying:
The -c flag performs syntax validation without changing state. Useful in CI/CD before deployment.
Enable persistence on systemd distros:
nftables.service loads /etc/nftables.conf on boot. After editing the file, systemctl restart nftables applies it.
Forward Chain
If the host is not a router, leave forward empty with a drop policy. When IP forwarding is on, the minimum is established replies plus transit between interfaces:
For debugging, log before the implicit drop:
Logs go to journalctl -k or /var/log/kern.log.
Common Scenarios: Blocking Ports and IPs
Block incoming connection from a specific IP:
Block outgoing to a specific IP:
Block an IP range (CIDR):
Block a port for everyone:
Allow a port only for a specific subnet:
Log dropped packets:
Counters appear in nft list ruleset output—they show packet and byte counts.
NAT via Masquerading
For sharing a single IP with a local network:
Masquerading automatically substitutes the external IP of the interface. For port forwarding NAT:
NAT in nftables works only for IPv4. For IPv6, use stateless NAT66 or routing-level addressing.
iptables-nft Compatibility Mode
Some distributions keep iptables “working” by translating to nftables:
The problem is these are two different worlds. iptables-nft translates commands to nftables, but reverse compatibility does not work. Rules created via iptables will not appear directly in nft list ruleset.
Do not use iptables and nftables simultaneously. The result is unpredictable. Either fully migrate to nftables or stay on iptables. Check current mode: iptables -V shows whether iptables-legacy or iptables-nft is in use.
For migration from iptables, use the iptables-translate utility:
Output: nft add rule ip filter input tcp dport 22 accept. Manual verification of output is mandatory—automatic translation is not perfect.
nftables is not the future—it is the present. If you administer Linux servers, spend an evening on migration. The file /etc/nftables.conf with the complete ruleset is your backup and deployment in one.
40 - ss: socket statistics instead of deprecated netstat
When netstat hangs on a server with tens of thousands of connections, it’s time to switch to ss. Part of the iproute2 package, ss queries the kernel directly via netlink instead of parsing /proc/net/*. The result is instant output with minimal overhead.
Why switch from netstat
netstat from net-tools relies on a deprecated approach: it reads from /proc/net/tcp, /proc/net/unix and converts numeric IDs to symbolic names. On a server with active connections, this takes seconds and spikes CPU usage.
ss communicates with the kernel through a netlink socket. One call, structured data returned. For 10,000 connections, the difference is 0.02 seconds versus 3–5 seconds.
netstat is officially marked as deprecated in most distributions. iproute2 is the current standard for network management in Linux.
Separate installation is rarely needed: ss ships with iproute2, which is present in every Linux by default.
Basic flags: the -tulnp equivalent
No need to relearn everything — flags are similar, just with more flexible ordering:
| Flag | What it shows |
|---|---|
-t | TCP sockets |
-u | UDP sockets |
-l | Listening sockets only |
-n | Numeric addresses and ports (no DNS) |
-p | Process owner (PID, name) |
-a | All sockets (not just listening) |
-e | Extended info (uid, inode) |
-o | Timer information |
Flag order doesn’t matter — -tlnp and -ltnp are the same. -p needs root to show processes owned by other users.
Output differs structurally from netstat:
Key columns:
| Column | Meaning |
|---|---|
| Recv-Q | Bytes in receive buffer, not yet read by application |
| Send-Q | Bytes in send buffer, not acknowledged by peer |
| Local Address:Port | Local end of the connection |
| Peer Address:Port | Remote end |
Non-zero Recv-Q or Send-Q on an established connection signals a problem. The application isn’t keeping up with reads, or the network is congested.
Full set of basic filters:
Filtering by state and port
This is where ss beats netstat. Filters are native, not piped through grep:
State combinations:
Port filter — one of the most common:
The sport and dport filters work with numeric values. For ranges, use >= and <=: 'dport >= 3000 and dport <= 4000'.
Address filter:
Extended output: -e, -i, -s
For queue diagnostics and statistics:
Output with -e for an established connection:
Interface information (-i):
Key metrics:
| Metric | Description |
|---|---|
| cwnd | Congestion window |
| rtt | Round-trip time |
| pmtu | Path MTU |
| rcv_space | Receive buffer size |
| send | Current send rate |
State statistics (-s):
The timewait 2 in ss -s output is a quick way to assess connection buildup. If the number grows each time you run it — something isn’t closing connections properly.
Common use cases
Which process is listening on a port and which interface:
Listening twice — once on all interfaces, once on localhost. If you need external only — check bind-address in the config.
How many sockets in TIME_WAIT — and who they belong to:
High TIME_WAIT is usually normal. If it causes issues — on the client side use setsockopt with SO_LINGER, or SO_REUSEADDR on the server. Flag -ttu shows timers.
Whether a connection is stalled:
Check rtt and unacked. If unacked grows but rtt stays the same — packets aren’t arriving but aren’t being lost either. Most likely the remote side stopped reading from the socket.
Find the process that opened a connection to a specific host:
Aggregated statistics by process:
Shows how many connections each process holds. Useful for finding processes opening too many sockets.
Quick reference for everyday tasks:
41 - SSH Config: Wildcards and Dynamic Variable Substitution
SSH reads ~/.ssh/config line by line, but without variables the file quickly becomes copy-paste hell. Here’s how Host patterns, Match exec, and substitution tokens like %h, %r, %l cut config size by orders of magnitude while covering real scenarios — from dynamic routing to agent forwarding through bastion hosts.
Templates and wildcards in SSH config
SSH supports glob-like patterns in the Host directive. The most common is Host *, but compound patterns work too.
SSH reads config top-down and applies the first matching Host directive. Put more specific rules above general ones.
Patterns can be combined in a single Host line via spaces — this works as logical OR.
Substitution variables: %h, %r, %l
SSH substitutes variables in directive values during config processing. Three core ones:
| Variable | Value | Example |
|---|---|---|
%h | Hostname from command line | ssh web-01 → web-01 |
%r | Remote username | ubuntu |
%l | Local username | alex |
Substitution works in most directives — ProxyCommand, IdentityFile, LocalCommand, RemoteCommand.
Variable %h substitutes what you passed after ssh. This lets you write one rule for hundreds of hosts.
%h substitutes what you typed, not the resolved HostName. ssh web-01 substitutes web-01, not its IP.
Match exec and dynamic routing
The Match directive lets you apply rules based on conditions. Without exec, it checks user, host, localuser. With Match exec, you get arbitrary logic via shell command.
The exec condition runs on your local machine. The command must return 0 (success) for the Match block to apply.
Match exec runs through the shell. Backticks and variables expand. For complex logic, extract the check into a separate script.
Escaping the percent sign
When you need to pass a literal % character — in RemoteCommand, ProxyCommand, or LocalCommand — use %%.
Examples for dev, staging, prod
Typical setup: one file, three environments, shared base.
After this config, ssh prod-web-03 connects to prod-web-03.prod.internal via prod-bastion with key id_ed25519_prod, while ssh dev-app-01 connects directly.
-G outputs all parameters SSH will use after parsing the config for the given host. Use it for debugging before making real connections.
Combining Host patterns, variable substitution, and Match exec transforms SSH config from a collection of copy-paste blocks into a dynamic routing system. One file covers dev/staging/prod without duplication, and %h, %r, %l eliminate manual edits when adding new hosts.
42 - strace: System Call Tracing for Diagnosing Hangs and Leaks
When a service hangs, standard tools like top, htop, and ps show the state but not the cause. If a process is in state D (uninterruptible sleep), it’s waiting on a syscall. strace attaches to a live process and outputs every system call in real time. This turns a mysterious hang into a specific syscall, its arguments, and return code.
strace uses ptrace, the kernel’s debugging mechanism. On production, tracing slows a process by 2–10x. Use it sparingly, targeting a single PID.
Basic Flags and Syntax
Installation:
Running:
Core flags for diagnostics:
| Flag | Purpose |
|---|---|
-f | Follow child processes |
-c | Count calls and time (summary) |
-tt | Microsecond timestamps |
-T | Time spent in each syscall |
-e trace=openat,read,write | Trace only specified calls |
-e write=1,2 | Trace writes to fd 1 and 2 only |
-o output.log | Write output to file |
-s 1024 | Truncate strings longer than N characters |
Diagnosing Blocking Calls
Scenario: process is hung, ps shows state D. Need to find what it’s blocked on.
Typical output when blocked on a file:
If you see read(...) <time> = 0 with large time values, the process is waiting for data. If <time> is in seconds, you’ve found the bottleneck.
Blocking on epoll_wait, poll, select is normal for an idle process. Look for read, write, openat, sendto with times exceeding 100ms.
For network sockets:
Hang on connect() to an unreachable host:
EINPROGRESS means non-blocking socket, but long time indicates a network issue or timeout.
Finding File Descriptor Leaks
Scenario: process won’t open files, error “too many open files”. Need to find who’s holding the descriptors.
Attach strace with a filter on file operations:
Analyze the output:
Alternative — summary mode:
If close is called less than openat — you found a leak.
With high call frequency (thousands per second), strace generates enormous output. Limit time with -tt and filter by syscall via -e trace=.
Analyzing Slow Requests
Scenario: API endpoint responds in 5 seconds instead of 200ms. Need to find where time is lost.
Find calls with high execution time:
openat with 523ms — investigate: file on NFS, missing permissions, remote filesystem.
For SQL-like queries (PostgreSQL, MySQL), trace the socket:
Slow database query looks like a series of write/read with long time between them:
4.7 seconds between sending the query and receiving data — problem is on the database side or network path to it.
Summary
strace turns a hang with no visible cause into a specific syscall. Attach to PID, filter calls via -e trace=, check execution time with -T. For leaks — compare open/close counts in summary mode. For slow requests — find calls with time exceeding 100ms.
On production use -e trace=write,read,openat instead of tracing all calls — reduces overhead by 3–5x.
43 - tcpdump and tshark: Packet Capture in CLI
When debugging network issues in Linux infrastructure, ping and curl are not enough. Sometimes you need to see what is actually traveling over the wire. tcpdump is the standard tool for capturing packets from the CLI. tshark is its sibling from the Wireshark suite, convenient for scripting.
Quick Start with tcpdump
Check that packets are reaching the host:
The utility puts the interface into promiscuous mode and prints one line per packet passing through. By default it works with the first interface it finds, but specifying explicitly is better.
For a quick test without DNS resolution (to avoid timeout when there is no network):
-nn prevents resolving both hostnames and ports.
Key Flags
| Flag | Purpose |
|---|---|
-i iface | Interface |
-c N | Capture N packets and exit |
-n | Do not resolve hostnames |
-nn | Do not resolve hostnames and ports |
-v, -vv, -vvv | Increase output verbosity |
-w file | Write raw dump to file (pcap) |
-r file | Read dump from file |
-X | Show hex + ASCII packet body |
-s N | Truncate each packet to N bytes (0 = full) |
-C | Rotate output file when it reaches N MB |
The -v flags are useful for debugging: the first level shows TTL and ID, the second shows flags and window size, the third adds ACK and displays the full IP header.
BPF Filters
tcpdump uses Berkeley Packet Filter. The syntax reads left to right.
Filters combine with and, or, not. Quotes are needed when the expression contains spaces or special characters.
Do not run tcpdump -i any in production without restricting by host or port. You will get a flood of traffic and waste disk space for no reason.
Writing to File and Reading
Capturing to a file is mandatory practice. A live sniffer dumps data dozens of lines per second; analyzing in the terminal is impossible.
The pcap format is binary. The -C option limits file size:
This creates files /tmp/capture-0, /tmp/capture-1 … up to 5 files, each up to 10 MB. -W sets the number of files.
tshark as a tty-free Alternative
tshark is the command-line part of Wireshark. It produces structured output convenient for parsing in scripts:
-T fields -e extracts specific fields from each packet:
Reading pcap files with tshark is more convenient than with tcpdump:
tshark does not support -C rotation like tcpdump, but it handles live filters better (full Wireshark dissector engine).
Reading Dumps in Wireshark
A pcap created by tcpdump opens in Wireshark without conversion:
In Wireshark Apply as display filter you enter the same BPF syntax: tcp.port == 443 && ip.src == 10.0.0.1.
If the file is large (>100 MB), do not open it entirely. Use tcpdump -r with a pre-filter to extract the needed slice: tcpdump -r big.pcap -nn 'host 10.0.0.5' -w small.pcap.
Common Mistakes
- Forgot
-c— the process hangs, capturing traffic indefinitely. Add the limit immediately. - No permissions — you need root or
sudo tcpdump. In modern distributions you can grant capabilities:setcap cap_net_raw,cap_net_admin=eip /usr/sbin/tcpdump. -wdoes not work with-l(line-buffering) simultaneously. If you need progress — write to file and read in parallel viatail -f.- File was written but reads empty — possibly there was no traffic matching the filter. Check
tcpdump -i eth0 -nnwithout a filter.
44 - CasaOS: Web Dashboard for Home Lab
CasaOS is a lightweight web dashboard for managing Docker containers on a single host. If you’re currently accessing each container through its own port, this replaces that mess with a unified interface where you can install apps with a couple of clicks, monitor resources, and manage storage — without touching Nginx or dealing with Portainer’s complexity.
Installation
CasaOS installs on a clean system with one command. Supported distros: Debian 10+, Ubuntu 18.04+, Raspbian. If Docker is already running, remove it first or accept that CasaOS will take over the Docker daemon.
The script outputs the dashboard address when done. Default is http://<server-IP>:80. Web UI is ready immediately, logged in as root.
The installer deploys its own Docker stack. If you already have dockerd running, CasaOS will overwrite network configs. Make sure you have snapshots or a backup ready.
Check the service status after installation:
Interface and Features
The main screen shows installed apps as tiles. Each card displays status (running / stopped), an icon, and quick actions: start, stop, restart, delete.
Left sidebar: file manager, container list, app store, and settings. Out of the box CasaOS provides:
- volume creation and USB drive mounting;
- port forwarding through the web interface;
- CPU, RAM, disk, and network usage metrics;
- compose file editing directly in the browser.
For a home lab, this covers about 80% of needs. The remaining 20% is manual compose files.
Installing Apps from the Catalog
The CasaOS catalog has several dozen pre-built templates. Find what you need, click Install — CasaOS generates the compose file and starts the container. Under the hood it runs AppStore, essentially a docker-compose wrapper.
Typical flow:
- Open App Store in the left sidebar.
- Pick an app, for example Plex or AdGuard Home.
- Set the data path and port mappings if needed.
- Click Install.
For AdGuard Home, CasaOS runs this under the hood:
All parameters are set through the web form — no manual editing required.
Running Containers Manually
When an app isn’t in the catalog, CasaOS lets you upload your own compose file.
In CasaOS: App Store → Upload — drop in the YAML. It parses the file, lets you adjust variables, and launches. The container appears on the main screen.
Use /DATA/AppData/ for all volumes. This makes backups straightforward — all state lives in one directory.
Storage Layout
CasaOS creates three key directories:
| Path | Purpose |
|---|---|
/DATA/AppData/ | Application configs |
/DATA/Storage/ | User files |
/var/lib/casaos/ | Panel system data |
When you attach a new disk or USB drive, it shows up in File Manager and mounts automatically. You can set any directory as the default storage location for new apps: Settings → Storage → Default Location.
Updates and Maintenance
Updating the panel itself:
Container images update through the web UI: app card → menu (⋮) → Update. CasaOS pulls the fresh image and recreates the container, leaving data untouched.
Service logs are accessible via terminal:
To prune unused images and networks:
When to Choose Something Else
CasaOS works well for a single host. If you’re running a multi-machine cluster, it won’t help — it’s not a replacement for Kubernetes or Docker Swarm.
Portainer gives more control over networks, images, and stacks. If you prefer detailed configuration and write compose files by hand, CasaOS will feel limiting.
Yacht is even more minimal, with no enforced directory structure. Good if you only want a web UI for containers without a file manager or app catalog.
CasaOS is actively developed. Minor releases break template compatibility more often than ideal. Keep your working compose files in a Git repository so you can rebuild stacks from scratch if needed.
For a home server running Pi-hole, Home Assistant, Jellyfin, and a few utilities — CasaOS handles the job without overhead. Install it, add your apps, forget about it.
45 - cockpit-ufw-module: Uncomplicated Firewall in Cockpit
UFW on a home or small server is usually configured over SSH: ufw status numbered, then ufw allow 443/tcp. cockpit-ufw-module covers the same cycle in the browser: package, status, policies, rules. It is one panel from the cockpit-modules group — UFW under Tools.
The UI is Russian, in PatternFly v5. You need Cockpit 264+ and administrator rights for writes. MIT license; current release is tag v.1.0.1.
Why a panel if ufw already exists
The client is already there: system ufw. What is missing is a view without a terminal and safe input: port, CIDR, rule number.
The panel does not replace iptables with its own engine. HTML and JS call host commands through cockpit.spawn as argv arrays, not shell strings. Input is validated on the client: port (including a range and a list), IPv4/IPv6 and CIDR, action allow / deny / reject / limit. Without Cockpit admin rights, package and rule operations do not run.
What the page shows
Four blocks, top to bottom:
| Block | Job |
|---|---|
| ufw package | installed or not, version, manager (APT, DNF, YUM, Pacman), install and remove |
| Status | active / inactive, logging level, incoming/outgoing policies, enable / disable / reload |
| Rules | table from ufw status numbered: number, to, action, from, delete |
| Add rule | action, protocol, port, source IP/CIDR |
While the firewall is off, the rules table is hidden: UFW does not apply them, even if the config still has entries. The counter then reads “N rules (inactive)”.
Installing the module
Use the store if it is already on the host. Otherwise install by hand:
Files go to /usr/share/cockpit/ufw. Refresh Cockpit — UFW appears under Tools.
For the current user without root:
The module lands in ~/.local/share/cockpit/ufw.
install.sh installs only the panel. The ufw package is installed from the UI, with Install ufw.
Typical flow on a clean host
Order matters more than the buttons. After the package install the module runs ufw --force reset, sets allow for incoming and outgoing, opens 22/tcp, and enables the ufw systemd unit. It does not enable the firewall — that is a separate button with an SSH warning.
- Tools → UFW → Install ufw. Confirm. On APT this starts with
apt-get update. - Add explicit rules for what you manage from this host. Cockpit listens on 9090/tcp — after the package install that rule is missing, only SSH 22 is present. If you later set incoming to
denywithout 9090, the panel disappears. - Services you need:
80/tcp,443/tcp, a range such as8000:8010. The source can be a CIDR, for example10.0.0.0/8. - Before adding, the UI shows a command preview (
ufw allow 443/tcp) — the same argv that will run on the host. - Enable. While default incoming is
allow, access is not cut: only explicit deny/reject/limit rules take effect. - Once SSH, Cockpit, and sites still work — switch the incoming policy to
deny. Leave outgoing onallowin the usual case.
Examples the form accepts:
limit is UFW’s rate limit on new connections, useful on 22. reject replies with ICMP/TCP reset; deny drops silently.
Install ufw runs ufw --force reset. Existing rules on the host are wiped. On a host that already has a tuned firewall, install the package by hand and use the panel only for status and small edits.
What the panel does not do
This is not the full ufw from the man page. The UI has no:
- application profiles (
ufw allow OpenSSH,ufw app list); - interface binding (
on eth0); - rule comments;
- in/out direction on the add form — new rules are inbound;
- logging level changes (the Logging line is read-only);
- routed policy, even though
ufw status verboseis parsed.
Removing the package disables UFW (ufw --force disable) and uninstalls it through the system manager; APT uses --purge.
Local check without touching a live host — Docker Compose from the repository: a privileged container with systemd, Ubuntu, Cockpit, and ufw, default http://localhost:9090, login admin / admin. That is a demo, not a production pattern.
Sources, issues, and tags live at gitlab.com/cockpit-modules/cockpit-ufw-module. Neighbouring panels in the same group: fail2ban, cron, CertManager.
46 - Creating a Custom Systemd Service
Your application needs to start on boot, restart on crash, and log output. Shell scripts in /etc/rc.local give you none of that. Systemd solves all three with a single declarative file.
Why Write a Custom Unit File
Supervisord and init scripts are overkill for most cases. Systemd provides a unified interface for service management: socket-based activation, dependency tracking, resource limits, and built-in logging via journald. You get all of it without additional tooling.
Unit File Structure
A unit file is an ini-style text file placed in /etc/systemd/system/ for persistent configuration or /run/systemd/system/ for runtime-only changes. Naming convention is name.service.
Required Sections and Directives
| Section | Key | Purpose |
|---|---|---|
[Unit] | Description | Human-readable name |
[Unit] | After | Startup ordering relative to other units |
[Service] | Type | How the service demonizes |
[Service] | ExecStart | Command to execute on start |
[Install] | WantedBy | Target that enables this unit |
Type
simple— process stays in foreground, systemd monitors it directlyforking— process forks and parent exits (classic daemon pattern)oneshot— runs once and exits, useful for one-off tasksexec— like simple, but waits for ExecStartPre to finish first
For modern applications written in Go, Node.js, or similar, simple is almost always correct.
Example: Application Service
The --user flag creates a user-level service that lives with the user’s session instead of the system. Useful for personal tooling without root access.
Python: venv and gunicorn
Point ExecStart at the venv interpreter, not system python. That keeps service dependencies off the host packages:
For a WSGI app:
Keep secrets in EnvironmentFile (KEY=VALUE), mode 600, owner root:appuser. If the unit fails to start, journalctl -u myapp usually shows a traceback, a missing WorkingDirectory, or a missing env file.
Production-Ready Directives
Restart Behavior
Dependencies
Resource Limits
Logging
Reading logs:
Activation and Management
Syntax check without applying changes:
Common Mistakes
Missing ExecStart. Service fails immediately with Unit entered failed state.
Type not specified. Defaults to simple, but if your process self-daemonizes, you need forking.
Wrong path to script. Verify the file exists and is executable. Systemd does not validate paths at parse time — it simply fails to start.
Runtime edit without reload. Changed the file, forgot daemon-reload. Systemd continues using the old version.
WorkingDirectory does not exist. If you specify it, the directory must be present.
Restart=always without ExecStop. If the process exits cleanly, always restarts it anyway. This means systemctl stop may not behave as expected for long-running services.
Never edit package-provided unit files directly in /usr/lib/systemd/system/. Updates overwrite them. Use drop-in files in /etc/systemd/system/<name>.service.d/ instead.
Drop-in example:
47 - kubectl whoami and Service Account Permission Checks
When deploying an application to Kubernetes, the most common failure is a service account that cannot do what it should. Permission denied when creating a secret, rejection on list pods, refusal on update. kubectl whoami and kubectl auth can-i let you quickly identify exactly who cannot do what.
kubectl whoami plugin
kubectl whoami is not a built-in kubectl command. It is a plugin from krew or a standalone binary. Install with kubectl krew install whoami or download from GitHub.
After installation the command shows the current context and associated service account:
Output:
Without the plugin, the same information is available via:
The beginning of the output shows User: system:serviceaccount:default:myapp — that is your whoami.
kubectl auth can-i: checking arbitrary subject permissions
The built-in kubectl auth can-i command checks permissions without impersonating the target subject. Syntax:
The --as format for a service account:
Responses: yes or no.
To check the full permission set of an SA:
Combining --as with --list is a fast SA permission audit. The output contains a table with resources, versions, and permitted actions.
kubectl auth can-i flags
| Flag | Purpose |
|---|---|
--as <subject> | Subject to check |
-n <namespace> | Subject’s namespace (for SA) |
--list | All permissions of the subject |
--quiet / -q | Exit code only, no output |
--no-headers | No table headers |
Exit code: 0 for yes, 1 for no. Useful for scripts:
Binding ClusterRole to a Service Account: RoleBinding vs ClusterRoleBinding
A service account is bound to a role through two resources. The difference is scope.
RoleBinding binds a Role or ClusterRole to an SA within a single namespace:
RoleBinding can reference either Role or ClusterRole. With Role, permissions are limited to that namespace. With ClusterRole, permissions apply within the RoleBinding’s namespace but inherit all ClusterRole rules.
ClusterRoleBinding binds a ClusterRole cluster-wide:
| Scenario | Resource |
|---|---|
| SA needs permissions in one namespace only | RoleBinding → ClusterRole |
| SA needs cluster-wide permissions | ClusterRoleBinding |
| Shared role for different SA across namespaces | ClusterRoleBinding with multiple subjects |
Subject in CSR: reading and verifying
A CSR (Certificate Signing Request) contains spec.username and spec.groups. For an SA it looks like this:
Output:
Three components in username separated by colons: system:serviceaccount:<namespace>:<name>.
To check permissions of a specific SA from a CSR:
CSR is created by kubelet when a node connects or through ServiceAccount admission. If an SA uses the TokenRequest API (the modern approach), no CSR is generated — the token is issued directly.
Quick check in CI/CD
In a pipeline you need to confirm the deploy service account has sufficient rights before running manifests:
Add this check after kubectl apply or helm template — an early exit is cheaper than a crashed pod with ImagePullBackOff due to missing imagePullSecrets.
Useful output for diagnostics in CI logs:
Lines without headers are easy to grep for suspicious wildcards like */* or */delete.
48 - ngrep: grep for Network Packets in Real Time
Ngrep applies grep-style pattern matching to network packets. When you need to see exactly what two services are exchanging over the wire and tcpdump drowns you in noise, ngrep isolates the payload content you care about.
Installation
Ngrep ships in the standard repositories of most distributions.
Running ngrep requires root privileges or the CAP_NET_RAW and CAP_NET_ADMIN capabilities.
Basic Syntax
Minimal invocation catches all packets on a port:
Breaking it down:
-i— case-insensitive search'password'— regex pattern matched against payloadport 80— BPF filter: traffic on port 80 only
Output resembles tcpdump with decoded payload:
Common Flags
| Flag | Description |
|---|---|
-i | Case-insensitive search |
-W byline | Line-wrap output, strip escape sequences |
-q | Quiet: show matches only, suppress metadata |
-t | Prefix each packet with timestamp |
-d eth0 | Listen on a specific interface |
-n 1 | Exit after first match |
-c N | Limit output to N characters per line |
-x | Show hexdump instead of ASCII |
Practical example with timestamps:
Port and Protocol Filtering
Ngrep accepts standard tcpdump BPF filters. Common patterns:
Ngrep parses BPF filters identically to tcpdump. The port, host, and, and or syntax works the same way in both tools.
HTTP Requests in Docker Containers
A frequent task is tracking HTTP traffic between containers. Two approaches work.
First: run ngrep inside the container
Works if the container is Alpine or Debian-based and you have access:
Second: listen on docker0
When containers communicate over the host’s bridge network:
Check a container’s IP:
Limitations and Alternatives
Ngrep only works with plaintext traffic. It cannot decrypt TLS/HTTPS — you will see binary garbage instead of content.
Other limitations:
- Does not understand HTTP/2 or HTTP/3 — these protocols are binary
- No built-in JSON or XML parsing, only text-based matching
- Performance lags behind tcpdump under high load
Alternatives by use case:
| Task | Tool |
|---|---|
| Quick pcap capture | tcpdump -i eth0 -A port 80 |
| Detailed protocol analysis | tshark -Y http -i eth0 |
| HTTPS monitoring (key required) | Wireshark with decryption |
| gRPC tracing | grpcurl or Wireshark with Protobuf dissector |
For most debugging tasks, ngrep plus tcpdump covers the bases. If you need to parse a specific protocol in depth, tshark with filters gives more control — but requires learning Wireshark’s syntax.
49 - socat: Forwarding Unix Sockets Over TCP
Sometimes you need to reach a Unix socket from a host where that socket doesn’t physically exist. SSH tunnels won’t help — they only work with TCP ports. socat solves this: it opens a TCP listener and forwards connections to a Unix socket, and the client just connects over the network.
Installation
The package is available in every major distribution. On Debian/Ubuntu:
On RHEL/CentOS:
Alpine:
Verify:
Basic Forwarding: TCP-LISTEN + UNIX-CONNECT
Server side. Listen on a TCP port and redirect traffic to a Unix socket on connection:
Flags:
TCP-LISTEN:2375— opens port 2375fork— spawns a child process for each connection; without it socat accepts one connection and exits
On the client side, work as usual — for example, curl the Docker API:
If the client is on a remote host, specify the server IP:
Docker listens on the local socket by default. Exposing TCP-LISTEN externally without TLS or firewall is a risk. Restrict the bind to an interface: TCP-LISTEN:2375,bind=127.0.0.1.
Stop the forward — Ctrl+C or kill by PID.
Client Test via STDIO
To quickly verify socket availability or send a manual command, use STDIO on the client side:
After starting, enter raw HTTP requests. Example session:
Exit — Ctrl+D or Ctrl+C. This is handy for debugging APIs without curl and without setting environment variables.
For a TCP connection over the network, the client runs symmetrically:
Abstract vs Filesystem Sockets
Unix sockets come in two types. The difference matters for socat operation.
Filesystem sockets — bound to the filesystem. Path starts with /:
Abstract sockets — live in kernel memory, have no filesystem representation. Path starts with \0 or @ (ASCII zero and at-sign). Docker in rootless mode uses these:
Check socket type:
In socat syntax:
@/path/to/socket— abstract socket/path/to/socket— filesystem socket
Abstract sockets are invisible to processes without namespace access. This is an advantage for isolation but complicates forwarding between containers.
Timeout Flags
By default, socat waits forever. For automation and scripts, you need timeouts.
| Flag | Description |
|---|---|
readtimeout=SECONDS | Read timeout |
writetimeout=SECONDS | Write timeout |
timeout=SECONDS | Timeout for both operations |
Example with a general timeout:
Connection closes after 30 seconds of inactivity.
Separate timeouts for client and server:
In cron scripts or systemd units, set a timeout or the process will hang on network disruption:
For infinite waiting without fork, forever works, but in production combine it with system limits:
reuseaddr lets you quickly restart socat without “Address already in use” errors.
Common Errors
Permission denied accessing the socket
Connection refused
Verify socat is running and listening on the port:
Firewall:
One request — and socat dies
Missing fork. Each socat instance handles one connection and exits. Add the flag:
Abstract socket not found
Ensure the correct prefix. Docker in rootless uses @, but this is ASCII 0:
In socat, write @/docker.sock, not @@/docker.sock.
Systemd Service for Persistent Forwarding
For a permanent forward, wrap it in a systemd unit:
Don’t forget to restrict the bind to an interface if you don’t want the port exposed externally.
In Closing
socat is the Unix way for transparent forwarding of anything to anywhere. For Unix sockets over TCP, two processes and a minute of configuration are enough. Keep timeouts in mind, don’t forget fork, and don’t expose ports without authentication on public networks.
50 - SSH Escape Sequences: Reviving a Frozen Terminal
SSH session froze, Ctrl+C does nothing, Ctrl+D spits out garbage — familiar situation. Before closing the terminal and losing the session, try built-in escape sequences. They operate at the SSH client level before data reaches the remote host.
How to invoke escape sequences
The default escape character is tilde (~). The combination works only at the beginning of a line. Press Enter, then ~, then the desired symbol. For example, ~. terminates the connection.
If tilde does not work — make sure you pressed Enter before it. In the middle of terminal output the sequence is ignored.
Escape sequence reference
This command prints the list of available escape sequences directly into the terminal.
You do not need to memorize everything. Keep three scenarios in mind — they cover 90% of problems.
Emergency disconnect
~. closes the SSH connection immediately, including all multiplexed sessions. Works even if the remote host is not responding. Do not confuse it with ~^Z (suspend) — that leaves the session alive as a background process.
Forced disconnect does not send SIGHUP to the remote side. If an important process was running without nohup — it will die.
Breaking the data flow
~V (Shift+v) lowers the logging level. The reverse command — ~v — increases verbosity. In practice useful when you see a stream of meaningless debug messages and want to hide them.
A more radical option — ~B. Sends BREAK to the remote system. This can interrupt a stuck process that is listening on a serial port. Use deliberately: not all systems handle BREAK correctly.
Built-in command mode
You enter the SSH client command line. Here you can manage tunnels on the fly without reconnecting.
Command mode is convenient when you forgot to forward a port before connecting. No need to reconnect — add it directly from the session.
Format: -L [bind_addr:]port:host:hostport and -R [bind_addr:]port:host:hostport. Bind address localhost restricts forwarding to the local interface only.
Forced rekey
Forcibly initiates re-keying of the session. Sometimes useful when you notice strange delays or suspect encryption issues. In practice rare, but worth knowing.
Escape sequence summary
| Sequence | Action |
|---|---|
~. | Emergency disconnect |
~V | Decrease verbosity |
~v | Increase verbosity |
~C | Command line (tunnels) |
~R | Forced rekey |
~B | Send BREAK |
~? | Show help |
~^Z | Suspend ssh |
~# | List forwarded connections |
~& | Background mode (on disconnect) |
~~ | Send literal tilde |
Changing the escape character
If you need tilde in the remote session (for example, connecting to Cisco equipment), change the escape character:
Now escape sequences are invoked via ^Z instead of ~. Found in specific scenarios — do not touch without need.
You cannot change the character on the fly in an active session. Only at connection time via the -e flag.
Common pitfalls
~does not work — forgot to press Enter before it.- Session is frozen but you are not sure SSH is still alive — try
~.blind. It will not make things worse. - Multiplexing session (
ssh -M) is bound to the control socket.~.closes all related sessions at once. - If you use PuTTY — escape sequences are different. There
~.sends a literal tilde and dot. UseCtrl+]and thenclose.
When escape sequences will not help
- Network is physically unavailable — only disconnect.
- Remote process consumed stdin — ssh will receive nothing.
- Problem is at the terminal emulator level — killall ssh, then check tmux/screen.
SSH escape sequences are the first-line tool when a session freezes. Keep ~. for disconnect and ~C for tunnels in mind — that is enough for everyday use.
51 - What is Self-Hosted and Why It's So Popular
You’re paying for Notion, Dropbox, and Google Drive. Then Notion raises prices, Dropbox caps storage, and Google “improves” the Docs interface. Self-hosted is taking infrastructure into your own hands and breaking the vendor dependency loop.
Definition: Not Your Cloud
Self-hosted means deploying and operating applications on your own servers or VPS instead of using SaaS alternatives. The server can sit at home, in a data center, or be a VM at a hosting provider—the point is that the hardware is under your control, not a third party’s.
The difference from classic hosting is straightforward: with virtual hosting you rent part of a server with pre-installed software. With self-hosted you install whatever you want and configure it however you like. It’s closer to VPS or bare metal, but focused on applications rather than websites.
Self-hosted ≠ self-administered. You can hire an admin, use managed servers, or deploy ready-made images. The point is that data and the service belong to you, not a cloud provider.
Typical Self-Hosted Stack
Most self-hosted setups revolve around Docker and Docker Compose. It’s the de facto standard for isolation and reproducibility.
Nextcloud usually sits behind a reverse proxy—Traefik or Caddy. Traefik auto-discovers containers via labels, Caddy works with zero-config TLS.
Common home setup stack:
| Category | Examples | Purpose |
|---|---|---|
| Files | Nextcloud, Syncthing | Storage and sync |
| Media | Jellyfin, Plex | Personal streaming |
| Passwords | Bitwarden, Vaultwarden | Password manager |
| Code | Gitea, Forgejo | Git repositories |
| Smart home | Home Assistant | IoT controller |
| Monitoring | Grafana, Prometheus | Metrics and alerts |
Why It’s Gone Mainstream
The main driver is data control. When clouds regularly leak data and corporate privacy policies shift with profit margins, the idea of “my server, my rules” appeals beyond just the paranoid.
Financial factors matter too. One server for $10–20/month replaces a stack of subscriptions. Bitwarden Premium costs $10/year, Nextcloud AIO is $29/year. Even a personal NAS with multiple disks pays for itself in a couple years versus Dropbox and Google One.
Less obvious: vendor lock-in. Notion exports data poorly, Slack doesn’t provide proper exports, and Discord keeps closing APIs. When you run your own instance, migration is just a backup and spin-up on a new server.
Keep your configuration in Git. Twenty docker-compose.yml and compose-override.yml files in a repo isn’t just a backup—it’s a reproducible setup that deploys to new hardware in minutes.
The last factor is the learning curve. Self-hosted has gotten easier: ready-made images, auto-TLS, Compose files on GitHub. This attracts people who want to understand infrastructure, not just use a service.
What Holds People Back
The main issue is maintenance burden. SaaS updates without your involvement. Self-hosted requires regular apt update && docker compose pull && docker compose up -d and monitoring for outdated images.
Security is the second bottleneck. Your server is your responsibility. Vulnerability in Nextcloud? That’s your CVSS, your patches, your downtime. Clouds at least patch faster and notify users.
Uptime is the third barrier. For a home setup it’s not critical. For a work tool you need backup connectivity, monitoring, and a process supervisor. Systemd unit or Watchtower with auto-restart is the minimum.
| Task | Tool | Complexity |
|---|---|---|
| Auto-update images | Watchtower | Low |
| Container monitoring | cAdvisor + Grafana | Medium |
| Downtime alerts | Healthchecks.io, Uptime Kuma | Low |
| Data backup | Restic, Duplicati | Medium |
When It Makes Sense
Self-hosted pays off in three scenarios.
Homelab and personal use. You effectively get a private cloud for the cost of a server. Files, passwords, photos, media library—all under your control. For many it’s both a hobby and hands-on practice.
Small teams without compliance requirements. Gitea for code, Vaultwarden for passwords, Outline or HedgeDoc for documents replace GitHub, 1Password, and Confluence. Budget is $20–50/month on a VPS.
Specific requirements. Custom integrations, above-average privacy, data that can’t live with third parties. This can be a personal project, a cleaning company handling personal data, or an early-stage startup.
Don’t try to replace managed services with self-hosted if your team has no time budget for operations. An SRE engineer on half-time will cost more than a Notion subscription and burn through goodwill faster.
Typical path: start with one service, add a second, then a third. By the sixth you realize you have a homelab with ten containers and a 200-line Ansible playbook. That’s fine—many people found their way into DevOps exactly this way.
52 - cockpit-modules: web panels for day-to-day operations
Cockpit covers basic Linux administration in the browser: services, logs, networking, accounts. Firewall, fail2ban, cron, and Let’s Encrypt sit outside that set — either there is no panel, or it is too generic.
The cockpit-modules group is a set of separate modules for those jobs, plus a store that installs them on the host. Each module lives in its own repository. This is a map of the group, not a walkthrough of UI and commands. Individual panels get their own articles.
Why extra modules
Cockpit grows through packages in /usr/share/cockpit/ (or ~/.local/share/cockpit/ for the current user). A module is a manifest.json, HTML, and JS that call host tools via cockpit.spawn.
The group targets a typical home or small server:
- perimeter: UFW and fail2ban;
- scheduling: crontab and systemd timers;
- TLS: certbot and the Cockpit panel certificate;
- delivery: a module store backed by the same GitLab group.
The UI is in Russian and matches the rest of Cockpit (PatternFly v5). You need Cockpit 264+ and administrator rights for writes.
Shared repository rules
Modules share the same layout so the store and install scripts stay consistent:
- sources in
pkg/<id>/; install.sh— system-wide or--user;- releases tagged
v.M.m.p; - Docker Compose for local development (
localhost:9090); - MIT license.
Host commands run as argv arrays, not shell strings. Input is validated on the client. Without Cockpit admin rights the panel stays read-only.
A project appears in the store catalog only if it lives in the cockpit-modules group. The catalog is not arbitrary GitLab — it is an explicitly trusted set.
Module store
cockpit-modules-store is the entry point. Under Tools you get Module store: the group project list, tags, status (available / installed / update ready), install of a release archive into /usr/share/cockpit/, and removal of user modules. Built-in Cockpit pages are left alone.
You can still install the other panels by hand with install.sh, but the store removes a clone/copy cycle per repository.
Quick install of the store itself:
Panels in the group
| Module | Job |
|---|---|
| UFW | ufw package, status, policies, allow/deny/reject/limit rules |
| Fail2ban | package and service, jails, ban/unban IPs |
| Cron | user crontabs, /etc/crontab and /etc/cron.d/, systemd timers |
| CertManager | Let’s Encrypt / certbot, Cockpit panel TLS, renew |
UFW is a full Uncomplicated Firewall cycle without ufw status numbered over SSH: the package via APT/DNF/YUM/Pacman, enable/disable, default policies, rule list, and add-by port, protocol, and source. Panel write-up: cockpit-ufw-module.
Fail2ban is the neighbouring perimeter panel: daemon install, jails, banned addresses, ban and unban in a selected jail or globally.
Cron is the scheduler people usually edit in nano. User crontabs, system files, and systemd timers in one place. Editing timer unit files is out of scope.
CertManager covers certificates for sites and for the Cockpit panel itself. HTTP-01 and DNS-01, binding a lineage to the panel, upload of your own crt+key, auto-renew via certbot.timer.
How it fits together
On a clean host the natural order is: Cockpit → store → UFW and fail2ban → CertManager for HTTPS on the panel → Cron if needed. Modules are independent: cron does not depend on certbot.
Sources, issues, and releases live at gitlab.com/cockpit-modules. First per-panel write-up: UFW.
53 - dig: DNS Query Debugging in CLI
DNS resolvers return the wrong address, clients don’t see updates, or it’s unclear which server is handling requests. dig (Domain Information Groper) is the standard CLI tool for DNS diagnostics. Works on Linux, macOS, and Windows via WSL.
Installation
Basic Flags
dig has two classes of options: short flags (start with -) control query behavior, while keywords with + control output format.
Without arguments the output is verbose — lots of header information. For debugging you pick the pieces you need.
| Flag | Purpose |
|---|---|
-b <addr> | Outbound IP address (multi-interface hosts) |
-f <file> | Read queries from file, one per line |
-p <port> | Non-standard DNS server port |
-t <type> | Record type: A, AAAA, MX, TXT, SOA, NS, CNAME, ANY |
-c <class> | Network class (default: IN — Internet) |
-x <addr> | Reverse lookup (PTR) |
-6 | Force IPv6 |
-4 | Force IPv4 |
Quick Response: +short
For scripts and quick checks:
If the record doesn’t exist — empty output. For CNAMEs you see the final address but not the chain.
For AAAA records:
+short doesn’t show TTL and doesn’t guarantee it’s the final answer. A CNAME loop will return the last record in the chain, not an error.
Full Response: +noall +answer
When you need TTL, canonical name, and all records at once:
TTL in seconds. Small values (60–300) mean the record changes frequently.
Verbose response with timing:
Reverse Lookup: -x
Reverse zone: IP → hostname.
Not all PTR zones are populated. Empty response with -x is normal, especially for client addresses.
IPv6: -6
Force IPv6 transport to the DNS server:
If you want an AAAA record over IPv4 transport — just request the type:
Chain Tracing: +trace
Shows the path from root servers to the final answer:
Output is split into sections: . (root), TLD (.com), authoritative NS, response.
+trace is slow — walks the hierarchy recursively. Use it for NXDOMAIN and SERVFAIL diagnosis.
Specific Resolver: @server
By default dig uses the system resolver from /etc/resolv.conf. Specifying explicitly compares responses or bypasses local cache:
AXFR transfer works only if the NS allows it.
Full zone AXFR is sensitive. Don’t do this on third-party NS without reason.
Common Scenarios
Checking SOA and NS records
Serial in SOA — if you updated records but it didn’t increase, transfer hasn’t happened.
TTL of a specific record
300 seconds is low TTL, normal for frequently changing records. Static records usually sit at 3600+.
CNAME chain
The chain displays in full. If a redirect broke somewhere — NXDOMAIN or SERVFAIL at which step becomes clear from the output.
Comparing resolvers
If addresses differ — the issue isn’t on your service side, it’s with that specific resolver or record propagation.
SERVFAIL analysis
Check flags: SERVFAIL in the response means the NS couldn’t fetch data. Verify the requested record type actually exists in the zone.
ANY query (carefully)
Deprecated in production. But for quick diagnostics of all record types on a new NS — it works.
54 - MkDocs: Project Documentation from Markdown
Documentation in a repository ages faster than anyone reads it: README links lead nowhere, sections are scattered across docs/, wiki/, and Confluence, and site search doesn’t work. MkDocs solves this predictably — it takes a folder of .md files and builds a static site. One config, one command for the deploy, familiar Markdown.
What is MkDocs
MkDocs is a static site generator for documentation written in Python. Input: a directory of Markdown files and a YAML config. Output: a ready site/ directory with HTML, served by any web server or hosted on GitHub Pages, GitLab Pages, S3. MkDocs core handles rendering; the theme defines look and features. The de-facto standard is Material for MkDocs.
MkDocs does not use Jinja templates and does not require a database. It is static content that builds locally or in CI in a few seconds.
Installation
The minimum requirement is Python 3.8+. Install into a virtual environment to avoid polluting the system pip.
Verification:
Typical output is mkdocs, version 1.6.x. The version matters because themes and plugins often require a specific range.
Useful packages installed alongside the core or separately:
| Package | Purpose |
|---|---|
mkdocs-material | Material theme, navigation, search, tabs |
mkdocstrings | Documentation generated from Python docstrings |
pymdown-extensions | Additional Markdown extensions for Material |
mkdocs-minify-plugin | HTML/CSS/JS minification in site/ |
Pin theme and plugin versions in requirements.txt. Material breaks compatibility between minor releases, as does mkdocstrings.
Creating a Project
mkdocs new scaffolds the project:
This creates a docs/ directory with index.md and an empty mkdocs.yml. That is the working minimum — nothing else is mandatory.
Directory Structure
docs/ is the single source of Markdown. The directory hierarchy maps directly to URLs. The file docs/guide/install.md becomes /guide/install/. An index.md in the root of docs/ is the landing page.
The site builds into the site/ directory next to mkdocs.yml. This directory is a build artifact — it is committed only for manual deploys, and usually built by CI.
Configuration mkdocs.yml
A minimal working config:
The full set of keys actually used in production:
| Key | Purpose |
|---|---|
site_name | Site title and default <title> |
site_url | Canonical URL, required for sitemap.xml and robots.txt |
site_description | Description, goes into meta tags |
docs_dir | Directory with Markdown, defaults to docs |
site_dir | Where HTML is built, defaults to site |
theme | Theme and its parameters |
nav | Explicit navigation, overrides auto-discovery |
plugins | Plugins in load order |
markdown_extensions | Enabled Markdown extensions |
extra | Arbitrary variables read by the theme |
Example with navigation, extensions, and plugins:
Enable search explicitly when using the plugins list. In newer Material versions it is no longer pulled in automatically from the theme.
Content
Markdown files are standard CommonMark with extensions. Useful constructs that work out of the box with the extensions enabled in the example above.
Admonitions:
Tabs with pymdownx.tabbed:
Code highlighting with language specified:
Use heading anchors to link between pages. Material renders a # icon next to the heading when toc.permalink: true is set.
Internal links are relative paths from the current file:
External links open in the same tab by default. To open in a new tab:
Build and Local Server
Local development — run the live server:
By default it listens on http://127.0.0.1:8000. Useful flags:
| Flag | Effect |
|---|---|
--dev-addr 0.0.0.0:9000 | Change address and port, useful in a container |
--strict | Build fails on any warning, including broken links |
--livereload | Page reloads in the browser without F5 (on by default) |
--no-livereload | Disable auto-refresh |
--clean | Remove site/ before building |
Build the artifact for deployment:
--strict is mandatory in CI. Otherwise typos in links and missing files in nav will be silently built. The exit code on warning is zero, so without --strict the pipeline passes green with broken documentation.
Do not run mkdocs serve in production. It is a dev server with no authentication and file reload enabled.
Common errors on first run:
WARNING - A relative path to '...' is included in the 'nav' config. A file is listed innavbut missing fromdocs/. Check case and path.WARNING - Documentation file 'x.md' is not included in the 'nav' configuration. The file exists but is not in navigation. Either add it tonavor rely on auto-navigation by removing thenavsection entirely.ERROR - Config value 'theme': The theme 'mkdocs' is not installed. The theme package is not installed, or the name is misspelled.
Material for MkDocs: Navigation, Search, Tabs
Material extends base MkDocs with three things that are almost always needed.
Navigation. Enabled through features in the theme section:
The combination of navigation.tabs + navigation.sections gives the familiar “sticky” menu. Without sections the tabs scroll away with the content.
Search. The search plugin ships with Material but registers separately:
For Russian documentation, stemming is important — set search.lang: ru in the theme config. The list of supported languages is published in Material’s documentation and new languages are added regularly.
Tabs within a page. Implemented via the pymdownx.tabbed extension, connected above. The alternative syntax using !!! example blocks does not work in some themes — this is a frequent reason “why tabs don’t render.”
Code highlighting and copying. The “copy” button is enabled by the content.code.copy feature. Line numbers come from the pymdownx.highlight extension with linenums: true or globally via markdown_extensions:
pymdownx.superfences is required for highlighting inside admonitions and tabbed blocks. Without it, code renders as a plain block without syntax highlighting.
Dark theme is a must-have for documentation read at night on-call. The toggle is configured in the example config above through palette with two schemes: default and slate. To make Material serve the correct scheme on first visit, add a custom script to <head> or use theme.palette.toggle with media — it works without JS and respects the user’s system settings.
55 - Squid: Internet Forwarding to Remote VM
Your VM in the cloud has no public IP or internet access is blocked via NAT, but the deployment needs wget/curl from inside. Squid on an intermediate host with a decent uplink solves this in ten minutes.
Why This Is Needed
I forward internet through Squid when a VM sits in an isolated network segment. An intermediate host with a public IP and network access becomes the proxy server. The application on the remote machine routes traffic through the tunnel.
Real-world cases: test environments without external connectivity, CI/CD agents in a private subnet, temporary traffic routing for debugging.
Installing and Basic Squid Setup
Install on the intermediate host (Ubuntu/Debian):
The default config lives in /etc/squid/squid.conf. Minimum working config:
Enable and start:
Verify the port is listening:
ACL and Port Configuration
If access should be only via SSH tunnel from a specific address, replace localnet with the exact IP:
To change the port:
After config changes, reload without restarting:
If Squid fails to start after changes, check the log: journalctl -u squid -n 50.
SSH Tunnel to VM
On the remote machine, establish a tunnel to the intermediate host:
-N — don’t open a shell, forwarding only. -L binds local port 3128 to localhost:3128 on the remote host.
For background:
Or via a systemd user service:
Client Proxy Setup
On the remote VM, set environment variables for applications that respect http_proxy:
Or permanently in /etc/environment:
For curl/wget, variables suffice. For apt — additionally:
Test:
If it returns the intermediate host IP — it’s working.
Verification and Logging
Squid access log:
Format: time client/status code size method URL
Sample entry:
Response codes: TCP_HIT — served from cache, TCP_MISS — fetched from network, TCP_DENIED — access denied by ACL.
To clear the cache before testing:
| Flag | Description |
|---|---|
http_port | Listening port |
acl name src IP/mask | Access rule by IP |
http_access allow|deny | Permit or deny ACL |
-k reconfigure | Reload config |
-k shutdown | Graceful stop |
-z | Initialize cache directories |
Squid caches responses by default. For debugging, disable caching: add cache deny all to the config, then run squid -k reconfigure.
Nine minutes of setup — and the isolated VM has internet via proxy. If you need HTTPS transparent mode with certificate substitution — that’s a different story involving SSL-bump and CA generation.
56 - SSH certificates instead of authorized_keys
authorized_keys works fine for a handful of servers. Once you hit a dozen, it becomes a liability. Onboarding a new developer means manually distributing their public key across every machine. SSH certificates flip this model: one CA signs all public keys, and authorized_keys stays empty.
Why authorized_keys breaks at scale
authorized_keys requires your public key to exist on every target server. Scaling this creates:
- a separate deployment step for key distribution during onboarding
- no centralized revocation — removing a key means touching each host
- key rotation touches every machine
- no expiration means stale access accumulates
An SSH certificate is a CA signature on your public key. The server only needs to trust the CA — your key never needs to be present locally.
How SSH certificates work
Two entities: the CA (Certificate Authority) and the signed key. The CA is a standard SSH keypair, typically ed25519 or RSA. Signing creates the certificate with ssh-keygen -s <ca_private> -I <identifier> <key.pub>, producing <key-cert.pub>.
The server needs only two things: the CA public key in TrustedUserCAKeys (for users) or a HostCertificate directive (for hosts). Authentication succeeds if the signature is valid and the certificate hasn’t expired.
A certificate does not replace the key. The key is still required — the CA signs it. The certificate adds metadata: TTL, principals, extensions, serial number.
Generating CA keys
| Flag | Purpose |
|---|---|
-t ed25519 | Key type; ed25519 recommended per RFC 8709 |
-f | Output file path |
-C | Comment; use for CA identification |
Store CA private keys securely — ideally on a dedicated build machine or in an HSM. The CA private key never goes to target servers.
Signing user certificates: one-liner
| Flag | Purpose |
|---|---|
-s ca_private | CA private key for signing |
-I identifier | String logged during authentication |
-n principals | Comma-separated list of authorized identities |
-V +52w | Validity: 52 weeks from now |
-z serial | Serial number; useful for audit trails |
The result is id_ed25519-cert.pub next to your private key. The user keeps both files. Signing another employee’s key takes the same command with a different key.
Duration suffixes: h (hours), d (days), w (weeks). +1d is one day, -1d means yesterday — already expired.
Signing host certificates
Generate host keys on each server if they don’t exist:
Sign on the CA machine:
The -h flag marks this as a host certificate. The -n value lists the hostname and IP the client will verify on connect.
Place the certificate next to the host key:
sshd_config: CertFile, TrustedUserCAKeys, HostCertificate
Configuration on the target server:
Validate and reload after changes:
TrustedUserCAKeys expects the CA public key, not a certificate. Certificates are only needed for host keys.
Principals and source restrictions
You can restrict a certificate to specific source IPs at signing time:
Or enforce from via AuthorizedPrincipalsFile on the server:
A certificate can carry multiple principals. sshd checks if any principal matches those listed in AuthorizedPrincipalsFile. This lets you issue a cert with ubuntu,deploy,admin and grant access through different principals on different hosts.
Time-to-live and rotation
Set TTL at signing. Practical guidelines:
| Role | Recommended TTL | Rationale |
|---|---|---|
| CI/CD, automation | 24–72 hours | Pipeline credentials are short-lived |
| Developers | 1–6 months | Balance between security and convenience |
| Hosts | 12 months | Host certificates bind to hostname/IP |
Rotation means signing a new certificate with a fresh serial. The old one becomes invalid after expiry. No centralized revocation list needed — an expired certificate simply fails validation.
Fingerprints: the host-key problem
Traditional known_hosts stores host key fingerprints. Host certificates break this: the fingerprint in known_hosts won’t match the certificate. Two solutions:
Use ssh_known_hosts with ssh-keyscan:
Or explicitly prefer certificate algorithms in your client config:
On first connect, ssh will prompt to accept the certificate, add it to known_hosts, and won’t ask again.
Troubleshooting
Debug in order:
In -vvv output, look for Certificate lines and Authentications that can continue. no matching identity means the certificate isn’t next to the key. certificate refused indicates the CA isn’t trusted or the principal didn’t match.
Inspect certificate contents:
Output shows Valid: from ... to ..., Principals:, Serial:. First place to check when something breaks.
too many authentication failures with a valid certificate usually means ssh is cycling through all keys before reaching the cert. Add -o PubkeyAuthentication=no before explicit key specification, or remove extraneous keys from .ssh.
SSH certificates eliminate key distribution across hosts. One CA, signed keys with TTL, principals for access control. If you manage more than five servers, this isn’t optional — it’s infrastructure.
57 - systemd-timer: scheduling instead of cron
cron works, but its logs are flat text files with no structure, and service dependencies require workarounds like embedding Requires= logic inside shell scripts. systemd-timer fixes this: unified management interface, logs in journald, dependencies through the familiar After= and WantedBy= directives — all in one stack.
Structure: .service and .timer
A timer is a separate unit that triggers a .service. The separation is intentional: the service can be invoked manually or on a schedule.
backup.service is a regular unit, startable via systemctl start backup.service.
backup.timer is the trigger. Without it, the service will not run on a schedule.
Persistent=true — if the machine was off when the timer fired, it catches up after boot.
Calendar Timers (OnCalendar)
OnCalendar= format is the closest analog to cron expressions, but with different syntax.
| Example | Fires |
|---|---|
OnCalendar=daily | Every day at 00:00 |
OnCalendar=*-*-01 03:00 | First day of each month at 03:00 |
OnCalendar=*-*-* 02:00 | Every day at 02:00 |
OnCalendar=09..17:00 | Every hour from 09:00 to 17:00 |
OnCalendar=*:0/15 | Every 15 minutes |
OnCalendar=Mon..Fri 09:30 | Weekdays at 09:30 |
Multiple values can be specified comma-separated:
To validate syntax before applying:
Output shows the next firing date. Useful when “every second Tuesday” becomes *-*-1..31 03:00 — verify before deployment.
Monotonic Timers
Monotonic timers count from an event, not the clock.
| Directive | Fires |
|---|---|
OnBootSec=5min | 5 minutes after boot |
OnStartupSec=10min | 10 minutes after systemd manager starts |
OnUnitActiveSec=1h | 1 hour after the service last ran |
OnUnitInactiveSec=1d | 1 day after the service stopped |
OnBootSec and OnStartupSec are similar, but OnBootSec resets on each boot while OnStartupSec counts from when the systemd manager started. In practice, the difference shows in containers and during live migration.
Combination OnBootSec + OnUnitActiveSec implements “every hour, but not before boot”:
Verification: systemctl list-timers
After enabling and starting:
Without --all shows only active timers. NEXT is when it fires, LEFT is how long until then.
If the timer does not appear in the list — check status:
Common cause of a silent timer is a missing WantedBy=timers.target.
User-Level Timers (systemd –user)
Not every task needs root. Deployment scripts, home-directory cache cleanup, periodic git fetch — better under a user.
Units go in ~/.config/systemd/user/:
Activation:
User timers need linger if they should run without login:
Common Mistakes
Missing OnCalendar or monotonic timer. Without a [Timer] directive, the timer will never fire. systemctl start backup.timer starts the unit, but without a schedule it just sits there.
Forgot WantedBy. Without [Install], the unit does not persist across reboots. systemctl enable backup.timer completes without errors, but it will not appear in list-timers.
ExecStart in .timer. A timer only triggers its linked service. Placing ExecStart in .timer causes systemd to ignore it and pull from .service.
Persistent with no prior run. On first activation, Persistent=true does nothing — there is no “last run” in history. The service fires only at the next scheduled time.
Logging: cron vs timer in journald
Cron sends output to syslog or cron.log, structure is a text line with timestamp. Parsing requires grep or awk.
systemd-timer writes service stdout/stderr directly to journald:
Filtering by time, unit, severity — standard journalctl flags:
| Flag | Effect |
|---|---|
-u backup.service | This unit only |
-n 100 | Last 100 lines |
-f | Follow in real time |
--since "1 hour ago" | Time range filter |
-p err | Errors only |
Actual firing time is recorded in metadata. Build execution history without parsing text logs.
Output includes real start time, simplifying debugging of missed triggers.
Summary
Switching from cron to systemd-timer pays off when the workload already lives in a systemd environment. Unified management, dependencies via After=, logs in journald — gains are tangible. For a crontab one-liner, systemd-timer is overkill, but for scripts with dependencies, logging, and boot persistence — a mature tool.
58 - Cron setup: a practical walkthrough
Cron is the standard job scheduler in Linux, found in every infrastructure. Jobs pile up, logs accumulate, and the environment breaks things. Let’s walk through how to configure cron reliably and where it falls short.
When to Use Cron and When to Avoid It
Cron fits simple recurring tasks: backups, log rotation, temp file cleanup, periodic notifications. It’s a daemon that sleeps between runs — no resource consumption.
Avoid cron for:
- tasks requiring millisecond precision (cron fires at the minute level)
- tasks with hard dependencies on other services (use systemd units with
After=) - long-running processes that could overlap (you need locking)
Anatomy of a Crontab Entry
Format:
Special values:
Examples:
crontab -e and crontab -l: Common Commands
Default editor is vi. Change it: export EDITOR=nano.
Where System Schedules Live
Beyond user crontabs, there are system files:
Lines in /etc/crontab and /etc/cron.d/* include a username field:
Don’t add jobs directly to /etc/crontab. Use /etc/cron.d/ instead. It’s safer across package updates.
Environment and PATH: Why Jobs Break in Cron
Cron runs commands with a minimal environment. Common failure:
Reason: cron sets PATH to only /usr/bin:/bin. Solutions:
Specify full paths explicitly:
Set PATH in crontab:
Use a wrapper script:
Always verify variables: env | sort in terminal vs. job * * * * * env | sort > /tmp/cron_env.txt.
Output Redirection and Log Rotation
By default, cron emails stdout and stderr to the user. Set MAILTO="" to disable.
Or use logger for syslog:
Timezones and TZ
Cron uses the system timezone. To override:
TZ only affects the schedule. Inside the script, use $TZ explicitly: date +%Z shows system timezone.
Check schedule in UTC:
Common Pitfalls and Pre-Deployment Checks
1. Percent sign in commands
% in crontab means newline. Escape it:
2. Overlapping jobs
If a script may run longer than the interval, use locking:
3. Pre-deployment checks
4. Syntax and validation
Ansible and Cron: Idempotent Job Setup
systemd Timers as an Alternative to Cron
Timers are more precise, support service dependencies, randomiseddelaysec, and calendar specs.
Timer benefits:
- dependencies (
After=network.target) - logging via journal (
journalctl -u backup.service) - randomiseddelaysec for staggering (avoid thundering herd)
- one-shot and monotonic timers
Cron stays simpler for basic tasks. systemd timers are the right choice for services with dependencies and monitoring through journald.
59 - GNU Screen: sessions that survive an SSH drop
A long apt upgrade, a migration, a build — then the laptop sleeps. SSH dies, the process gets SIGHUP and dies with it. GNU Screen keeps the terminal on the server: disconnect, come back, the work is still there.
It is not a nohup replacement and not “another SSH”. It is a multiplexer: named sessions, several windows inside one, a shared session for two people. On older RHEL/Debian boxes screen is often already installed when tmux is not.
The minimal loop
On the server:
That is a normal shell. To leave without killing processes: Ctrl-a, release, then d. The session stays Detached.
List and resume:
If it is still Attached on another terminal (left it open at the office):
Detach there first, attach here.
The prefix is Ctrl-a. Do not hold both keys together: Ctrl-a, release, then the letter. Ctrl-a a sends a real Ctrl-a into the program inside (needed by emacs, rarely by bash).
Install
Check with screen -v. Vertical splits (Ctrl-a |) exist since GNU Screen 4.1; very old boxes will not have them.
Command-line flags
| Flag | What it does | Example |
|---|---|---|
-S name | create or address a session by name | screen -S logs |
-ls / -list | list sessions | screen -ls |
-r [name] | attach a detached session | screen -r logs |
-d -r name | detach elsewhere, attach here | screen -d -r logs |
-R | attach if it exists, otherwise create | screen -R logs |
-x [name] | attach without detaching the other side | screen -x logs |
-dmS name command | start in the background, already Detached | screen -dmS backup /opt/backup.sh |
-L | write screenlog.N in the current directory | screen -L -S migrate |
-Logfile file | log path (with -L) | screen -L -Logfile /tmp/migrate.log -S migrate |
-X command | send a command into a live session | screen -S logs -X stuff 'tail -f /var/log/nginx/error.log\n' |
-wipe | drop dead sockets from the list | screen -wipe |
The -S name is what you look for in screen -ls. Without a name you get something like 12345.pts-0.hostname — awkward to recall.
Typical background job that must survive an SSH drop:
When the command exits, the session usually disappears. Keep a shell afterwards:
Windows inside a session
One session, several windows: build, logs, a second shell. Switch without a new SSH connection.
| Key | Action |
|---|---|
Ctrl-a c | new window |
Ctrl-a n / Ctrl-a p | next / previous |
Ctrl-a 0 … Ctrl-a 9 | jump by number |
Ctrl-a " | window list, pick with arrows |
Ctrl-a ' | jump by number or name |
Ctrl-a A | rename the current window |
Ctrl-a Ctrl-a | last active window |
Ctrl-a k | close the window (asks to confirm) |
Ctrl-a \ | kill every window and quit screen |
Window titles help once you have more than two: Ctrl-a A → nginx-log.
Session, copy, split
| Key | Action |
|---|---|
Ctrl-a d | detach; processes keep running |
Ctrl-a D D | power detach (other displays too) |
Ctrl-a ? | key help |
Ctrl-a : | screen command line (quit, sessionname, …) |
Ctrl-a a | send Ctrl-a into the window |
Ctrl-a [ | copy mode / scrollback |
Ctrl-a ] | paste screen’s paste buffer |
Ctrl-a Esc | same as Ctrl-a [ |
Ctrl-a S | split horizontally (region above/below) |
Ctrl-a | | split vertically |
Ctrl-a Tab | focus the other region |
Ctrl-a X | close this region (window stays) |
Ctrl-a Q | keep only this region |
Ctrl-a H | toggle screenlog.N |
Ctrl-a M | monitor the window for activity (bell) |
Ctrl-a x | lock the session (user password) |
Scrollback: Ctrl-a [, then arrows or PageUp / PageDown. Select: Space to start, arrows, Enter to copy into screen’s buffer. Leave the mode with Esc. Paste with Ctrl-a ]. This is not the system clipboard: the buffer lives inside screen.
After a split the new region is empty until you focus it (Ctrl-a Tab) and pick a window (Ctrl-a n or Ctrl-a ").
Server-side workflows
Long deploy. Named session, log on disk, close the lid:
In the morning: screen -r migrate, or just tail -f ~/migrate.log.
Several jobs in one SSH. Session ops, windows build, journal, sql:
Paired view. Same user, both attached:
Both see the same terminal. Faster than “paste me the output” during an incident.
Push a command into a session already running — no interactive attach:
-X screen opens a window; stuff types into it. \n is Enter.
A short .screenrc
Default scrollback is tiny, the startup banner gets in the way, and window names are easy to miss. In ~/.screenrc:
defscrollback is how many lines Ctrl-a [ can walk. hardstatus is the bar at the bottom: windows, host, load, time. The file is read when a session is created; live sessions pick it up only after you recreate them.
You do not have to touch /etc/screenrc: the user file extends it.
Common failures
| Symptom | Why | What to do |
|---|---|---|
There is no screen to be resumed | wrong name, or the session already died | screen -ls, then the exact name |
Attached and -r refuses | session still on another pty | screen -d -r name |
| Session vanished after a command | the only window’s process exited | wrap with bash -lc '…; exec bash' |
Ctrl-a “eats” emacs/tmux inside | screen’s prefix wins | Ctrl-a a for a literal; or another escape in .screenrc: escape ^Bb |
| No vertical split | Screen < 4.1 | horizontal Ctrl-a S, or a newer package |
| No log file | no -L and nobody pressed Ctrl-a H | enable it explicitly, check the session cwd |
Sockets live in /run/screen/S-$USER/ or ~/.screen/. Another user’s session is not something you attach to by accident: the directory permissions block it.
When screen, when not
A single non-interactive background command is enough for systemd-run --user, tmux, or even nohup. A standing workspace on a bastion, a deploy you start and then close the lid on, shared log tailing — screen covers that without extra dependencies.
tmux is nicer for splits and config. Screen wins when it is already there on the box you just SSHed into. A named session and Ctrl-a d are the reason it stays in muscle memory.
60 - Kafka: Cluster Health Check
A Kafka cluster in KRaft mode doesn’t forgive neglect until the first incident. Health checks need to be regular and quick — no graphs or dashboards, just the terminal. Here’s the command set that covers the typical checklist: processes, quorum, partition leaders, ISR, and a quick status report in one shot.
Checking KRaft Processes
KRaft mode has no separate ZooKeeper — the controller role is either co-located with the broker or isolated on dedicated nodes. First, verify the JVM processes are alive and see which mode each node started in.
For a managed cluster you expect three QuorumControllerMain processes on dedicated controllers and N kafka.Kafka processes on brokers. Mixed nodes (combined mode) run both roles in a single process — normal for smaller installations.
Then check startup logs and configuration:
The process.roles parameter accepts broker, controller, or broker,controller. If empty, the cluster is still in legacy mode with ZooKeeper — adapt the commands below.
Verify all nodes agreed on the quorum:
In the output look for leaderId, votedLeaders, and quorum size 2/3 (for three controllers) or N/N for a fully stabilized cluster.
Broker and Controller Status
Next, determine which brokers are actually responding and which dropped from the registry. Use kafka-broker-api-versions.sh — it returns supported API versions and simultaneously shows whether TCP connectivity reaches the broker.
If a node is unreachable you’ll see a timeout. This is the fastest way to distinguish “broker stuck in JVM” from “network issue.”
Full list of registered brokers and their state:
The command shows a JSON-like listing, but for a tabular report kafka-metadata-quorum.sh is more convenient:
The LEADER column shows the current quorum leader, REPLICAS shows all active nodes. Unresponsive controllers drop from the list.
The lastCaughtUpTime field in describe --status shows how far a controller lags behind the leader. 0 or fresh now — healthy. Lag in minutes — reason to check GC and network latency.
Topics and Partition Leaders
The cluster can be alive but without partition leaders — producers won’t write, consumers won’t read. So the next step is topics and their leaders.
List all topics with partition and replica counts:
The output table contains Leader, Replicas, Isr. If Leader equals -1, the partition has no active leader — this is an emergency, producers will receive NotLeaderForPartitionException.
To get only “bad” partitions in one pipeline:
To view leaders for a specific topic:
ISR and Unavailable Replicas
ISR (in-sync replicas) determines write reliability. Recommended setting is min.insync.replicas >= 2 for critical topics. Check for out-of-sync:
The command returns partitions where Isr is smaller than Replicas. Empty output — good. Any lines in the list — incident.
Full picture across all partitions with problem filtering:
kafka-topics.sh --describe on a topic with thousands of partitions outputs many lines and loads the controller. In production run with --partitions N or filter with awk, otherwise the health check itself becomes a problem.
Additionally — state of a specific replica on the broker side. If you suspect one of the disks is lagging:
In the output check the partition.error field — if non-empty, the replica has issues (offline log dir, disk full, fs in read-only).
Quick Diagnostics in One Command
For daily rounds it’s convenient to bundle all checks into a single script with clear exit codes. Below is a minimal status indicator in bash.
Save as kafka-health.sh, make executable, and wrap in cron or a systemd timer every 60 seconds:
For alerting, replace echo blocks with logger -p local0.err and configure rsyslog to your SIEM. No need for Prometheus node_exporter — events fly into the common channel.
Typical errors on first run and how to interpret them:
| Symptom | Probable Cause | Action |
|---|---|---|
Connection to node -1 could not be established | Broker not registered in cluster but process is alive | Check advertised.listeners, node.id, network ACL |
isLeader: false for all controllers | Quorum lost, controllers < __.min.insync.replicas | Check controller.quorum.voters and disk state on controllers |
--under-replicated-partitions shows entries for >5 min | Broker lagging, slow disk or GC pauses | Take jstack, check iostat -x and LogFlushRate metrics |
kafka-log-dirs.sh returns LogDirOffline | Disk full or failed | Free space, check FS for read-only, restart broker |
Empty --describe output on a running cluster | Wrong bootstrap address passed | Verify advertised.listeners and DNS |
The five command sets above cover 90% of operational “is the cluster alive?” questions. If everything is green — no need to dig deeper; if something is red — kafka-log-dirs.sh and kafka-metadata-quorum.sh describe --status will show which direction to go.
61 - kind: Local Kubernetes in Docker
kind creates a Kubernetes cluster from Docker containers: control-plane and worker nodes are kindest/node images. Primary use cases are local development and CI. GitHub Actions has an official create-kind action that makes pipelines with K8s tests straightforward.
Compared to minikube, kind has no hypervisor dependency and natively supports multi-node topologies. Compared to k3d, it requires no Rancher and talks to CRI/containerd directly.
Installation
macOS, Linux, and Windows (via WSL2) — download the binary from GitHub Releases.
Docker is required (or Podman with kind use docker driver). Make sure Docker has at least 4 GB of memory allocated for all containers.
First Cluster in One Command
kind created a cluster named kind and wrote kubeconfig to ~/.kube/config. Verify:
Done. Single-node cluster is up in a minute.
Config: Multi-node and Runtime
For a cluster with multiple worker nodes, use a YAML config. Let’s create three workers:
Cluster creation flags:
| Flag | Purpose | Example |
|---|---|---|
--name | Cluster name | kind create cluster --name prod |
--config | Path to YAML | --config ./kind.yaml |
--image | Custom node image | --image kindest/node:v1.28.0 |
--wait | Readiness timeout | --wait 5m |
--kubeconfig | Alternative kubeconfig | --kubeconfig ~/.kube/dev |
Instead of --image, you can set node.internalImage in the config — useful for air-gapped environments.
Loading Images into the Cluster
kind uses a separate container runtime inside the node. Images from your local Docker daemon are not visible. To load them:
For CI, you often load from a tar archive:
After loading, the image is available in the cluster without a registry.
extraMounts and kubeadm patches
Mount a host directory into a node — useful for a local registry or fixture files:
Tune kubelet or kube-proxy via kubeadm patches:
If the API server port 6443 is already taken, set another in the cluster config:
kind with kubeconfig
By default, kind merges the context into ~/.kube/config. For isolation:
Managing multiple clusters:
Cleanup
Without --name, the default cluster (kind) is deleted. All Docker resources are removed with the nodes. If Docker was stopped while the cluster was running, nodes remain in NotReady on the next start. Fix by recreating the cluster.
Gotchas
containerd inside the node. kubectl talks to containerd directly. Familiar docker ps and docker exec won’t show pods — they live inside the kind node. For debugging:
HostNetworking and PortMappings. For external access to pods, you need extraPortMappings in the config. Without it, hostPort does not work — pod networking in kind is isolated.
PersistentVolumes. kind creates a kind-node StorageClass backed by local-path-provisioner. Data lives on the host at /var/local-path-provisioner. For a clean environment, just delete the cluster.
cgroups v2. In newer distros (Ubuntu 22.04+, Fedora), Docker may require configuration:
kind works with both versions, but rare issues with memory limits occur with cgroups v2 on Arch Linux. Solution — pass in the config:
Versions. kind lags behind upstream Kubernetes by 1-2 minor versions. Check the compatibility matrix in the project README before setting up a prod-like environment.
Port already allocated. kind could not bind the API server or a mapped host port. Check that 6443 is free, or change networking.apiServerPort. For Ingress, extraPortMappings must not collide with host services.
No space left on device. Docker ran out of disk. Clean unused images, then recreate the cluster:
kind is a fast way to spin up Kubernetes on a developer machine or in CI without virtualization. Main loop: kind create cluster, work, kind delete cluster. For air-gapped or multi-node scenarios — YAML config and kind load docker-image. Known limitations: no GPU support, no real network stack, performance below bare metal. For everything else — it works.
62 - Too many authentication failures: SSH ran out of tries
Received disconnect from 10.0.0.5 port 22:2: Too many authentication failures followed by Permission denied (publickey) is not a broken server, and it is not necessarily a wrong password. The client spent the attempt budget while walking agent keys and never reached the method you meant to use.
The budget is MaxAuthTries on the server (6 by default). Each public key offered is one try. Five keys in ssh-agent plus one more wrong key — the connection dies before the right key or a password prompt.
Why this happens
SSH asks the server which methods it accepts, then the client walks its own order. publickey is first by default. While the agent keeps offering keys, the server is already counting failures.
| You wanted | What actually happened |
|---|---|
| Key login | The key is missing from authorized_keys, it is not in the agent, or it is fifth in line — the limit hits first |
| Password login | The server advertised publickey, the client started the key walk, the password prompt never appeared |
| PAM / external password | The client sends password, the server wants keyboard-interactive |
A file sitting in ~/.ssh is not enough: OpenSSH offers agent identities, not “the file you had in mind”. What will actually be sent:
Empty agent or the wrong key — read -vvv instead of guessing.
Client log first
Debug level 3 shows the method order and every key offered:
Look for:
Authentications that can continue— what the server allowed;Offering public key/get_agent_identities— which keys go out;Too many authentication failures— attempt budget, not “wrong password”.
Typical picture: the agent hands over id_rsa, id_ed25519, a GitLab key, a bastion key — four rejects, the fifth is still wrong, the sixth is no longer accepted.
Stop offering every key
IdentitiesOnly yes stops the client from mixing in every agent identity. After that, only IdentityFile or -i is used.
Per host, not in the global /etc/ssh/ssh_config:
One-shot:
-vvv should then show a single Offering public key. If too many turns into Permission denied (publickey), the walk is over and the key was simply rejected: it is not in authorized_keys, or it is the wrong file.
IdentitiesOnly without IdentityFile keeps the default id_ed25519 / id_rsa in the home directory and does not dump the whole agent. For a stand with its own key, always name the file.
Password when the agent is full
The client will not ask for a password until publickey is exhausted. Turn keys off and put password first:
In ~/.ssh/config:
On Ubuntu and anywhere the password goes through PAM, the server often advertises keyboard-interactive, not password. The command above then fails silently. Switch the method:
Which method the server actually offers is the Authentications that can continue line in -vvv.
What to check on the server
You need another channel: hypervisor console, VNC, serial. Otherwise you are stuck in the same error you are fixing.
Logs. On Ubuntu 24.04 the unit is ssh (sshd is an alias). For rejected-key detail, raise verbosity:
The log shows which key was rejected and whether MaxAuthTries was reached.
Permissions. With StrictModes yes (the default) sshd silently ignores files that are too open:
The public key is one line in authorized_keys; the matching private key is the one in IdentityFile.
Attempt limit. A fat agent and many keys per host can justify raising the cap (this weakens brute-force protection):
Better not to raise it: shrink the client with IdentitiesOnly and one IdentityFile per Host.
Short checklist
ssh -vvv— how many keys went out and which method remained.ssh-add -l— what sits in the agent; do not offer the rest.- In
~/.ssh/configfor the host:IdentitiesOnly yesand an explicitIdentityFile. - Need a password —
PubkeyAuthentication noand the method the server printed in debug (passwordorkeyboard-interactive). - From the server console:
journalctl -u ssh,~/.ssh/authorized_keysmodes,LogLevel VERBOSEif needed.
A “too many” error is almost always on the client: too many keys for one login. The server only counts to six and closes the session.
63 - Trusting a custom CA: system store, browsers, and CLI
A TLS error that says the certificate is untrusted almost never means the certificate is “broken”. The trust anchor is in the wrong store.
curl, openssl, Git, Python, and the browser are different clients. Some keep their own root lists. Updating the OS bundle on Linux will not fix Chrome or Firefox by itself.
What to import
You trust the CA root that signed the server certificate, not localhost.crt / app.example.internal itself.
Typical files:
| File | Role |
|---|---|
ca.crt | root (trust anchor) — this is what you import |
server.crt | service certificate signed by the CA |
server.key | service private key, server-side only |
A self-signed server certificate with no separate CA can be added as an anchor on its own. For labs and internal PKI you usually keep a CA: one root, many services.
PEM (-----BEGIN CERTIFICATE-----) or DER both work. Most OS tools expect PEM. Browsers and certutil accept DER as well.
The system store
This is what OpenSSL-based clients use: curl, wget, git, many agents. Do this first, then debug the browser.
Linux
Distros rebuild /etc/ssl/certs in different ways.
Debian, Ubuntu, and derivatives — a .crt file under /usr/local/share/ca-certificates/, then rebuild the bundle:
RHEL, Fedora, Alma, Rocky — drop the anchor, then extract:
Alpine — package ca-certificates, same update-ca-certificates, directory /usr/local/share/ca-certificates/.
Then:
If curl is clean and the browser is not, the OS trust is fine. The browser is looking elsewhere.
macOS
Use the System keychain, not the login one, if every process on the machine should trust the CA:
For a single user, the login keychain is enough: no sudo, ~/Library/Keychains/login.keychain-db. GUI: Keychain Access → System → Certificates → import → Always Trust for SSL.
Windows
The Local Machine Root store:
Same via certmgr.msc (current user) or certlm.msc (computer): Trusted Root Certification Authorities → import.
In a container or CI job the bundle lives inside the image. Installing a CA on the host does not help curl in the container. Copy ca.crt into the image and run the same update-ca-certificates / update-ca-trust at build time.
Why the browser still complains
Chrome takes public roots from the Chrome Root Store. Extra CAs come from the OS — but not on every platform.
| Client | Where your CA has to live |
|---|---|
curl / OpenSSL | OS bundle |
| Chrome / Edge on Windows and macOS | OS store (system install is often enough) |
| Chrome / Chromium on Linux | per-user NSS database, not /etc/ssl/certs |
| Firefox | per-profile store; OS roots only via policy, and not on Linux |
“Installed the CA on the system” and “opened the site in a browser” are separate steps.
Chrome, Chromium, Edge
GUI is the same idea on every OS: Settings → Privacy and security → Security → manage certificates. Authorities tab → import ca.crt. Trust for website identification is enough.
Direct URL: chrome://settings/certificates (Edge: edge://settings/certificates).
CLI on Linux. Chromium/Chrome use an NSS Shared DB. Since M146 the default is ~/.local/share/pki/nssdb; if ~/.pki/nssdb already exists, that one wins.
certutil comes from libnss3-tools (Debian family) or nss-tools (RHEL/Fedora/Alpine).
The three -t fields are SSL, email, and code signing. C means trusted CA. Need client-certificate issuance as well — "CT,,". A self-signed server cert with no CA — "P,,".
Fully quit the browser and open it again. Background Chrome processes will not pick up the new cert.
On Windows and macOS a separate NSS import is usually unnecessary: the OS store is enough.
Firefox
Its own database per profile. Snap/Flatpak/vendor builds use different paths — importing into a “normal” profile does not reach them.
GUI: Settings → Privacy & Security → Certificates → View Certificates → Authorities → Import. Enable trust for identifying websites.
CLI — same certutil, profile directory, not ~/.pki/nssdb.
Linux: ~/.mozilla/firefox/<id>.default-release/
macOS: ~/Library/Application Support/Firefox/Profiles/<id>.default-release/
Windows: %APPDATA%\Mozilla\Firefox\Profiles\<id>.default-release\
Substitute the right profile path if you have more than one.
Policy: one file for every profile
For a fleet, policies.json beats hand imports.
Policy location:
| OS | Where to put it |
|---|---|
| Linux | /etc/firefox/policies/policies.json or distribution/policies.json in the install dir |
| macOS | Firefox.app/Contents/Resources/distribution/policies.json |
| Windows | distribution\policies.json next to firefox.exe, or GPO |
ImportEnterpriseRoots makes Firefox trust roots from the OS store. It works on Windows and macOS. Mozilla does not implement it on Linux: there is no OS certificate store in the same sense.
On Linux (and as an explicit list on any OS) use Certificates.Install: a filename or an absolute path. A bare filename is searched here:
- Linux:
/usr/lib/mozilla/certificates,/usr/lib64/mozilla/certificates,~/.mozilla/certificates - macOS:
/Library/Application Support/Mozilla/Certificates,~/Library/Application Support/Mozilla/Certificates - Windows:
%LOCALAPPDATA%\Mozilla\Certificates,%APPDATA%\Mozilla\Certificates
PEM and DER both work. Restart Firefox after changing the policy.
security.enterprise_roots.enabled in about:config is the same as ImportEnterpriseRoots: Windows and macOS, not Linux. On Linux, Certificates.Install or the p11-kit-trust PKCS#11 module is the equivalent.
CLIs and runtimes that still ignore the OS
Even after the system CA is in place, parts of the stack ship their own bundle.
| Stack | What it uses | What to set |
|---|---|---|
Python (certifi, some requests/httpx) | its own Mozilla bundle | SSL_CERT_FILE, REQUESTS_CA_BUNDLE, or truststore |
| Node.js | its own list | NODE_EXTRA_CA_CERTS=/path/to/ca.crt |
| Java | JDK cacerts | keytool -importcert -alias internal-ca -file ca.crt -keystore "$JAVA_HOME/lib/security/cacerts" |
| Git | often OS OpenSSL/Schannel | OS CA; otherwise http.sslCAInfo |
macOS and Windows use different paths for the system bundle; for Node it is simpler to point at ca.crt itself.
Short checklist
- Import the CA, not the service leaf.
- Put it in the OS store and verify with
curl/openssl s_client. - Chrome on Linux needs NSS as well; on Windows/macOS step 2 is often enough.
- Firefox needs a profile import or
policies.json;ImportEnterpriseRootsis Windows/macOS only. - In a container, a JDK, and Node, check that runtime’s bundle, not only the OS.
There is no single command that covers all of this. There is a predictable order: OS → browser → runtime.
64 - ssh-connection-manager: a TUI for hosts in ~/.ssh/config
When ~/.ssh/config holds dozens of stands, bastions, and jump hosts, memorizing aliases stops being fun. ssh-connection-manager is a TUI on top of ordinary OpenSSH: a host list, a filter, a connect via the system ssh, and a way to append a new block to the config.
The CLI command is ssh-connect. Repository: gitlab.com/unsorted-projects/ssh-connection-manager.
Why not another SSH client
The client already exists: the system ssh. What is missing is navigation over the config file.
The tool:
- reads
~/.ssh/configor another file (-c); - lists concrete
Hostentries, not wildcards such asHost *; - follows
Include(up to 16 levels); - on
Entersuspends the TUI and runsssh -F <config> <alias>; - after the session ends, shows the list again.
It is Python 3.11+ and Textual. There is no custom SSH stack: keys, ProxyJump, and the agent stay with OpenSSH.
What the table shows
Column names keep the alias and the address apart:
| Field | Meaning |
|---|---|
| HostName | alias from the Host directive |
| ConnectPoint | address from HostName in the config (IP or DNS) |
| User | user, if set |
| Port | port, or 22 by default |
The details line shows IdentityFile and ProxyJump when present. If a host has no User (and none comes from Host *), the TUI asks for a username before connecting.
The filter (/) matches alias and ConnectPoint.
Keys
| Key | Action |
|---|---|
Enter | connect |
a | add a host to the config |
/ | filter |
q | quit |
Adding a host appends a block; existing sections are not rewritten. If the file does not exist yet, it is created with mode 0600. Wildcard characters are rejected in the alias: only explicit names appear in the list.
Example of what gets written:
Install
On Debian/Ubuntu do not install into the system Python (externally-managed-environment). Use a venv or pipx.
Then ssh-connect is available in the activated venv. To run it from any directory:
Or pipx install -e . — pipx isolates the environment and puts the binary in ~/.local/bin.
An interactive TTY is required: without a terminal, Textual cannot hand the screen over to ssh.
The tool does not rewrite other config blocks and ignores Match. Only named Host entries without *?[ appear in the list.
Why this fits the job
One config file stays the source of truth. The TUI does not duplicate inventory in YAML and does not store passwords: only what already lives in OpenSSH. For a Lead DevOps that is the usual loop: bastion, prod, jump — pick a row and you are in the session.