PostgreSQL Kubernetes Operators Compared in 2026
Chat2DB TeamRunning PostgreSQL on Kubernetes means choosing an operator, and the choice matters more than it looks. An operator owns failover, backups, upgrades and connection routing — the four things that decide whether a 3am page is a five-minute event or a data-loss incident. This comparison covers the four operators that people actually deploy in 2026, and what genuinely separates them.
What an operator has to do
Before comparing, it is worth being precise about the job. A PostgreSQL operator must:
- Bootstrap a cluster: initialise a primary, clone standbys, set up replication.
- Detect failure and promote, without ever promoting two primaries at once.
- Back up continuously — base backups plus WAL archiving to object storage — and restore to a point in time.
- Route traffic: a stable endpoint for writes that follows the primary, and one for reads.
- Upgrade both itself and PostgreSQL, including major versions, without losing data.
Everything else — dashboards, pooling, extension management — is convenience on top.
CloudNativePG
CloudNativePG is a CNCF project, originally from EDB, and it has become the default recommendation for new deployments.
Its distinguishing design decision is that it does not use Patroni or an external distributed configuration store. Failover is coordinated by the operator itself through the Kubernetes API, using the API server's own consistency as the source of truth. One fewer moving part, and no etcd cluster of your own to operate.
It also runs one PostgreSQL instance per pod with no sidecars — the instance manager is PID 1 in the container. That makes the pod lifecycle unusually predictable: when the pod dies, the instance dies, and there is no supervisor process making independent decisions.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: shop-db
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:17
storage:
size: 200Gi
storageClass: fast-ssd
postgresql:
parameters:
max_connections: "200"
shared_buffers: "2GB"
work_mem: "16MB"
bootstrap:
initdb:
database: shop
owner: appBackups are declarative, to S3-compatible storage, with continuous WAL archiving and a point-in-time restore expressed as another Cluster resource that bootstraps from a recovery source. Connection pooling comes from a separate Pooler resource that runs PgBouncer.
Major version upgrades are handled through a declarative offline in-place upgrade or, for near-zero downtime, by creating a new cluster that subscribes via logical replication and cutting over — the same pattern you would use outside Kubernetes.
Choose it if: you are starting fresh, you want the fewest moving parts, and you value a vendor-neutral CNCF governance model.
Zalando Postgres Operator
The Zalando operator is the veteran of the group, running Zalando's own fleet at large scale for years, and it is built on Patroni — the battle-tested HA framework — with Spilo as the container image.
That heritage cuts both ways. Patroni is extremely well understood, has been through every failure mode in production, and gives you a REST API for inspecting and manipulating cluster state. But it means an extra layer: the operator manages Patroni, Patroni manages PostgreSQL, and debugging an incident means understanding both.
apiVersion: acid.zalan.do/v1
kind: postgresql
metadata:
name: shop-db
spec:
teamId: "data"
numberOfInstances: 3
volume:
size: 200Gi
storageClass: fast-ssd
postgresql:
version: "17"
parameters:
max_connections: "200"
shared_buffers: "2GB"
users:
app_user:
- superuser
databases:
shop: app_userIts standout features are the team-oriented model — clusters belong to a teamId, and it integrates with an identity source to grant human access — and a genuinely useful web UI for browsing clusters. It ships PgBouncer support as a connection pooler sidecar and uses WAL-E/WAL-G for archiving.
Choose it if: you already know Patroni, you want its REST API and operational maturity, or you like the team-based access model.
Crunchy Postgres for Kubernetes (PGO)
Crunchy's operator is the most enterprise-shaped of the four. It leans on pgBackRest for backups, which is a genuine differentiator: pgBackRest is the strongest open-source PostgreSQL backup tool, with parallel backup and restore, incremental and differential backups, and block-level delta restore.
apiVersion: postgres-operator.crunchydata.com/v1beta1
kind: PostgresCluster
metadata:
name: shop-db
spec:
postgresVersion: 17
instances:
- name: instance1
replicas: 3
dataVolumeClaimSpec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 200Gi
backups:
pgbackrest:
repos:
- name: repo1
s3:
bucket: shop-db-backups
endpoint: s3.amazonaws.com
region: eu-west-1
proxy:
pgBouncer:
replicas: 2It also builds on Patroni for HA, includes PgBouncer as a first-class proxy, and ships monitoring (Prometheus, Grafana, pgMonitor) as an integrated bundle. The open-source project is real and usable; the commercial offering adds support and a curated image catalogue.
Choose it if: backup and restore capability is your primary concern, or you want a vendor with a support contract behind the operator.
StackGres
StackGres takes a different philosophical position: it bundles a full stack of PostgreSQL extensions and tooling — Patroni, PgBouncer, pgBackRest, Envoy, Prometheus exporters — and exposes it through a set of custom resources plus a web console.
Its notable design idea is separating configuration into reusable resources: an SGInstanceProfile for compute, an SGPostgresConfig for parameters, an SGPoolingConfig for PgBouncer, and an SGObjectStorage for backups. Clusters then reference them.
apiVersion: stackgres.io/v1
kind: SGCluster
metadata:
name: shop-db
spec:
instances: 3
postgres:
version: "17"
pods:
persistentVolume:
size: 200Gi
configurations:
sgPostgresConfig: pgconfig-prod
sgPoolingConfig: pooling-prodThe extension catalogue is genuinely broad, and installing an extension is a declarative change rather than a container rebuild. The trade-off is more concepts to learn than any of the others.
Choose it if: you run many clusters that should share configuration, or you need a wide range of extensions without maintaining custom images.
Head to head
| CloudNativePG | Zalando | Crunchy PGO | StackGres | |
|---|---|---|---|---|
| HA mechanism | Operator + Kubernetes API | Patroni | Patroni | Patroni |
| External DCS needed | No | No (uses Kubernetes) | No | No |
| Backup tool | Built-in (Barman-based) | WAL-G / WAL-E | pgBackRest | pgBackRest |
| Pooler | PgBouncer (Pooler CRD) | PgBouncer sidecar | PgBouncer (built in) | PgBouncer (built in) |
| Web UI | No (third-party plugins) | Yes | No | Yes |
| Governance | CNCF | Zalando | Crunchy Data | Ongres |
The row that decides most evaluations is the first two: CloudNativePG's decision to skip Patroni is either a simplification you want or a departure from a proven component you trust, and reasonable teams land on both sides.
The questions that actually decide it
Have you tested failover? Every operator claims automatic failover. What differs is behaviour under partial failure — a node that is unreachable but still running, storage that is slow rather than gone. Before production, kill things: delete the primary pod, cordon its node, sever the network, fill the disk. Measure how long promotion takes and whether any writes were lost.
Have you restored a backup? Configuring backups is a spec field. Restoring is a procedure. Do it, on a real dataset, and time it.
-- After any failover or restore drill, on the new primary:
SELECT pg_is_in_recovery(); -- must be false
SELECT * FROM pg_stat_replication; -- standbys reattached?
SELECT count(*) FROM your_busiest_table; -- expected row count?
SELECT pg_last_wal_replay_lsn();Who upgrades PostgreSQL? Minor versions are an image change in all four. Major versions are where they differ substantially — read the specific procedure for your operator before you are on a version that is about to go end of life.
How do humans connect? Operators create Kubernetes Service objects for the primary and replicas, and secrets holding credentials. For day-to-day work you will port-forward and connect with a normal client:
kubectl port-forward svc/shop-db-rw 5432:5432
psql "host=127.0.0.1 port=5432 dbname=shop user=app"A GUI client such as Chat2DB (opens in a new tab) works over the same forwarded port, which matters more than it sounds when you have a dozen clusters and need to compare a staging schema against production. It is free to download at chat2db.ai/download (opens in a new tab).
A recommendation
For a new deployment with no existing investment, CloudNativePG is the default: fewest moving parts, CNCF governance, active development, and a declarative model that is pleasant to work with. If backup and restore is the thing that keeps you up at night, Crunchy PGO and its pgBackRest integration is the strongest answer. If your team already runs Patroni and knows its failure modes, Zalando lets you keep that knowledge. And if you operate many clusters that should share configuration, StackGres is built for exactly that shape.
Whichever you choose, the operator is not the hard part. The hard part is that PostgreSQL on Kubernetes puts your database's availability in the hands of a control loop, and control loops do surprising things during partial failures. Budget the time to break it deliberately, in a cluster you can afford to lose, before it breaks itself.
