Skip to content

Kubernetes on bare metal: a three-node high-availability design

Three dedicated servers are enough for a Kubernetes cluster that survives a node failure and runs a serious production workload. Here is the design that works, the choices that matter, and the operations load you take on or hand off.

10 minute read

Why three nodes

Kubernetes stores cluster state in etcd, which needs a majority to agree. Three nodes tolerate the loss of one; two nodes tolerate nothing. Three 48-core servers also give enough capacity that losing one leaves the workload running at two-thirds, which is the point.

Distribution

For a small cluster, a lightweight distribution with batteries included is the right call. k3s bundles etcd, a container runtime, a load balancer and ingress; Talos goes further by replacing the operating system with an immutable, API-managed image that cannot be logged into. Both run the full Kubernetes API.

  • k3s: familiar Linux underneath, easy to debug, large community.
  • Talos: smaller attack surface, no SSH, upgrades are an API call. More to learn, less to patch.
  • Upstream kubeadm: maximum flexibility, maximum work. Rarely worth it at three nodes.

Control plane and workers

At three nodes, run the control plane on all three and schedule workloads on all three. Dedicated control-plane nodes are a luxury for larger clusters. Taint nothing; set resource requests honestly so the scheduler can keep the system components alive under load.

Networking

  • Private network: put the nodes on a private VLAN and bind etcd and the API to it. The public interface carries only ingress traffic.
  • CNI: Cilium is the modern default, with eBPF networking, network policies and observability built in.
  • Load balancing: with no cloud load balancer, use MetalLB or Cilium's own announcement so a service IP floats between nodes.
  • Ingress: one ingress controller, cert-manager for certificates, external-dns to publish records. DNS points at the floating IP.

Storage

This is the hard part of bare-metal Kubernetes. Local NVMe is fast and simple but ties a pod to a node. Replicated block storage across nodes gives mobility at the cost of write latency and complexity.

  • Longhorn or OpenEBS Mayastor: replicated volumes across the three nodes, pods can move. Expect two to three times the write latency of local disk.
  • Local path provisioner: raw NVMe speed, no mobility. Fine for databases that handle their own replication.
  • The pragmatic design: replicated storage for stateless-ish services with small volumes; local NVMe plus application-level replication for databases.

Databases on the cluster

Run Postgres with an operator such as CloudNativePG, using local NVMe on two or three nodes and streaming replication between them. The operator handles failover, backups to object storage and point-in-time recovery. Redis with Sentinel follows the same pattern. This is how Infraexa Cluster 3 ships.

Delivery

Use GitOps from day one. Argo CD or Flux watches your repositories and applies what is there; nobody runs kubectl apply against production. Separate namespaces for staging and production on the same cluster are fine at this scale; separate clusters are not worth the cost until compliance requires it.

Operations you take on

  • Kubernetes upgrades three times a year, one node at a time, with workloads draining and returning.
  • Operating system and firmware patching underneath, coordinated with the Kubernetes upgrades.
  • Certificate rotation for the control plane; most distributions handle it, but check.
  • etcd backups, separately from volume backups, and a tested restore of the control plane.
  • Capacity: watching the moment when losing one node would no longer leave enough room for everything.
  • Monitoring for the cluster itself, not only the applications: node pressure, pending pods, failing probes.

When not to use Kubernetes

A single application with a database does not need an orchestrator. Docker Compose on one Core 48 with a second server as a warm standby is simpler, cheaper and easier to reason about. Kubernetes earns its complexity when you have many services, several teams deploying, or clients to isolate.

Questions

Can Cluster 3 span two data centres?

A cross-site variant is quoted individually. Latency between the German and Finnish sites is fine for asynchronous replication but too high for etcd across sites, so the design changes.

Do I get kubectl access?

Yes, with roles. Day-to-day deploys go through GitOps; kubectl is there for debugging.

Tell us what you run. We will tell you what it costs to run it properly.

A quote within one business day, from an engineer rather than a sales script. No setup fee, three-month minimum, delivery in 48 hours.