Velocity Stream LogoVelocity Stream Logo
Back to Insights
Database Architecture

Zero-Downtime PostgreSQL Upgrades on AWS RDS

How we upgraded a 2TB production PostgreSQL database from version 11 to 15 with less than 5 seconds of downtime using AWS RDS logical replication.

Upgrading a major database version in production is one of the most terrifying operations an engineering team can undertake. When a SaaS client approached us to upgrade their monolithic 2TB PostgreSQL 11 database to version 15 on AWS RDS, the immediate constraint was clear: no maintenance windows exceeding 60 seconds.

The traditional in-place upgrade via AWS RDS requires shutting down the database, taking a snapshot, running pg_upgrade, and restarting. For a 2TB database, this can take 45 minutes to several hours depending on extension compatibility and catalog size. That was unacceptable.


The Strategy: Logical Replication

To bypass the in-place upgrade downtime, we utilized PostgreSQL Logical Replication via the pgoutput plugin. This allows a continuous stream of data changes to flow from the old database (Publisher) to a newly provisioned database running the new version (Subscriber).

Phase 1: Architecture & Preparation

First, we provisioned a brand new PostgreSQL 15 RDS instance. This gave us the opportunity to implement Infrastructure as Code (Terraform) best practices that were missing from the legacy instance, including better Parameter Group tuning and Graviton3 instance types.

We enabled the rds.logical_replication parameter on the source database and set wal_level = logical. Because this requires a reboot, we scheduled a brief 2-minute maintenance window at 3:00 AM purely to enable the WAL logging required for the migration.

Phase 2: The Initial Sync

We created a publication on the source database for all tables:

CREATE PUBLICATION migrate_pub FOR ALL TABLES;

On the target PostgreSQL 15 database, we restored a schema-only dump (using pg_dump -s) and then created the subscription:

CREATE SUBSCRIPTION migrate_sub CONNECTION 'host=source.rds.amazonaws.com port=5432 user=rep_user password=rep_pass dbname=prod' PUBLICATION migrate_pub;

Engineering Reality

"During the initial sync, we closely monitored the source database's Read IOPS and WAL generation. Logical replication can put significant memory pressure on the publisher if the subscriber falls behind, potentially filling up the WAL storage and crashing the primary."

Phase 3: The Cutover

Once the replication lag dropped to zero bytes, we were ready for cutover. The actual cutover procedure took less than 5 seconds:

  1. Route 53 DNS TTL was reduced to 1 minute 24 hours prior.
  2. Applications were briefly paused (circuit breaker pattern triggered via Redis flag).
  3. We verified the replication lag was absolutely zero.
  4. We updated the AWS Secrets Manager connection string to point to the new PostgreSQL 15 instance.
  5. The circuit breaker was released, and applications reconnected to the new primary.

The Outcomes

The upgrade was a resounding success. Not only did we achieve the zero-downtime requirement, but the side-effects of the migration were massive.

5s
Total Downtime

The application briefly paused writes, but no 502 errors were surfaced to the end-users.

35%
Performance Boost

Driven by PostgreSQL 15's improved sorting algorithms and the migration to Graviton3 architecture.

Verdict

While logical replication is highly complex and requires rigorous testing (especially around sequences, DDL changes, and unsupported data types), it is the only viable path for true zero-downtime database migrations at scale.

Is your architecture slowing you down?

We specialize in auditing complex cloud environments, eliminating technical debt, and building pragmatic, high-performance infrastructure. Let's simplify your stack.

Chat with an Engineer