High availability in Keycloak is an architectural decision, not a configuration switch. The real question is not "how many nodes" but: what happens to existing sessions when a node goes away?
The question that comes before configuration
A second Keycloak node only improves availability once it is clear what it takes over during a failure. Without that, you end up with a system that looks redundant in normal operation and logs every user out when it matters.
Three decisions therefore come before any configuration file.
Session handling
Keycloak keeps session state in memory. Whether that state is replicated between nodes determines the failure behaviour:
- Without replication, users lose their session on node failure and have to sign in again. That is defensible — but it has to be a known property.
- With replication, sessions survive the failure. The cost is additional distributed state that can fail on its own.
Decision
Both options are legitimate. What is not legitimate is leaving the question open and discovering during an incident which one you built.
The database is the actual single point of failure
Realms, clients, users and roles live in the database. A Keycloak cluster in front of a non-highly-available database is an illusion: the application tier is redundant, the tier underneath is not.
Engineering note
Planning high availability for Keycloak means planning high availability for the database first. Anything else is the wrong order.
Load balancers and sticky sessions
Sticky sessions are tempting because they make replication problems invisible. That is precisely the problem: the fault does not go away, it just surfaces during a failover — exactly when something else is already broken.
A health check should verify actual readiness rather than whether a port answers:
curl --fail --silent --max-time 3 \
http://127.0.0.1:9000/health/ready \
|| exit 1
Name the operational boundaries
An architecture is only dependable once it is written down who owns which part — before the incident, not after:
| Area | Ownership |
|---|---|
| Keycloak configuration | Application |
| Database operation | Platform |
| TLS termination | Load balancer |
| Directory service | Organisation |
What this means for a project
The order is: decide the failure behaviour, make the database resilient, then add nodes. Reverse it and you get a cluster whose behaviour nobody can predict.
In practice
A deliberate node failure in a test environment answers in ten minutes what otherwise stays an assumption for months.
How these questions come together in an actual engagement is covered under identity & access management.