Identity & Access Management

Running Keycloak in high availability

High availability in Keycloak is an architectural decision, not a configuration switch. What actually has to be decided.

Diagram of a highly available Keycloak architecture with two nodes, a shared database and an upstream load balancer
AI disclosure: Image created with AI © EicDevCon UG (haftungsbeschränkt)

High availability in Keycloak is an architectural decision, not a configuration switch. The real question is not "how many nodes" but: what happens to existing sessions when a node goes away?

The question that comes before configuration

A second Keycloak node only improves availability once it is clear what it takes over during a failure. Without that, you end up with a system that looks redundant in normal operation and logs every user out when it matters.

Three decisions therefore come before any configuration file.

Session handling

Keycloak keeps session state in memory. Whether that state is replicated between nodes determines the failure behaviour:

  • Without replication, users lose their session on node failure and have to sign in again. That is defensible — but it has to be a known property.
  • With replication, sessions survive the failure. The cost is additional distributed state that can fail on its own.

Decision

Both options are legitimate. What is not legitimate is leaving the question open and discovering during an incident which one you built.

The database is the actual single point of failure

Realms, clients, users and roles live in the database. A Keycloak cluster in front of a non-highly-available database is an illusion: the application tier is redundant, the tier underneath is not.

Engineering note

Planning high availability for Keycloak means planning high availability for the database first. Anything else is the wrong order.

Load balancers and sticky sessions

Sticky sessions are tempting because they make replication problems invisible. That is precisely the problem: the fault does not go away, it just surfaces during a failover — exactly when something else is already broken.

A health check should verify actual readiness rather than whether a port answers:

Bash
curl --fail --silent --max-time 3 \
  http://127.0.0.1:9000/health/ready \
  || exit 1

Name the operational boundaries

An architecture is only dependable once it is written down who owns which part — before the incident, not after:

Area Ownership
Keycloak configuration Application
Database operation Platform
TLS termination Load balancer
Directory service Organisation

What this means for a project

The order is: decide the failure behaviour, make the database resilient, then add nodes. Reverse it and you get a cluster whose behaviour nobody can predict.

In practice

A deliberate node failure in a test environment answers in ten minutes what otherwise stays an assumption for months.

How these questions come together in an actual engagement is covered under identity & access management.