Close-up of a blue screen error shown on a data center control terminal.

Photo by panumas nikhomkhai on Pexels.

Every production database tells a quiet story about the decisions that got it there. At SimpleWP, that story starts with a feature that was exactly right for a while, and ends with a short list of loose ends that took real effort to track down. This is a postgres database migration story end to end — why the engineers behind SimpleWP built what they built, what they found once they looked closely, and what changed as a result.

A Postgres Database Migration That Started With Neon Postgres Branches

SimpleWP built an automated workflow that gave every pull request its own isolated database copy for CI testing, using Neon’s database-branching feature. That per-pull-request isolation meant a test suite always ran against clean, dedicated data instead of a shared copy that another branch might be changing at the same time.

The workflow did its job well for a long stretch, and then it stopped making sense. Months after it shipped, per-pull-request isolation had largely fallen out of use, while the branch-and-teardown cycle kept quietly consuming Neon and Fly.io resources on every single pull request regardless. Turning it off was the obvious next step once the gap between still running and still useful became clear.

Even in its early days the workflow needed a little more care than it first shipped with: some pull-request branches failed to clean up correctly once their pull request closed, and a dedicated follow-up fix closed that gap well before the whole workflow was eventually turned off entirely.

Cutting Over to DigitalOcean Postgres

The real postgres database migration came next: every environment’s database moved off Neon and onto a shared, managed DigitalOcean Postgres cluster, retiring Neon entirely rather than leaving it half in use.

CI changed shape at the same time. Instead of creating a fresh Neon branch for every pull request, the pipeline began running each test suite against a disposable PostgreSQL container that existed only for the length of that run. One pleasant side effect: pull requests opened automatically by dependency-update tooling could finally pass CI, since they no longer needed a Neon API credential they had never reliably had in the first place.

A Schema Replay That Had Never Happened Before

A disposable container forces a kind of honesty a long-lived one never does: there is nowhere to hide an accumulated inconsistency, because the database starts from nothing every single time.

Running every test against a brand-new, disposable database meant replaying the entire schema changelog from scratch, on every run, for the first time in the project’s history. That replay surfaced an ordering problem that had been quietly hiding for years: one early schema change could not be replayed cleanly on a brand-new database at all, and two later changes written specifically to fix it ran after the broken one in changelog order — so the fix could never actually reach the problem it was meant to solve. Every existing environment had simply been nursed past it by hand, one careful deploy at a time, without anyone noticing the changelog itself was unfixable from a cold start.

Three Weeks Later: a Postgres Database Migration Tool Still Pointed at Neon

The cutover looked complete at the application layer almost immediately — the application’s own database connection was repointed to DigitalOcean cleanly. But a tooling layer sitting one step removed from the application told a different story. Three weeks after the production cutover, a routine run of the team’s schema-migration tool, Liquibase, turned up a configuration value that still pointed at the decommissioned Neon host, even though the application itself had moved on weeks earlier.

That exact symptom was not new. The same stale-host behavior had already shown up once, a few weeks before the cutover even happened, and had been worked around rather than fixed at the time. The three-week-later discovery was the second time this exact thing had been seen, not the first, which is its own small lesson about treating a workaround as a fix.

Part of the same configuration also carried an old habit that never worked the way the team’s usual run command assumed: a placeholder written as ${DB_URL}, on the belief that the tool would expand it at run time. It never did. The migration tool’s own configuration format simply does not substitute variables that way, confirmed directly against several released versions of the tool rather than assumed from memory.

A Dead Collector and a Hidden Billing Gap

Not every loose end from a postgres database migration shows up inside the database itself. A nightly job that collected hosting-cost figures from Neon’s own API kept running, and kept failing, every single night for weeks after Neon had been fully decommissioned, because nobody had gotten around to removing a job that no longer had anything left to collect.

Fixing that dead collector turned up a second, unrelated problem hiding behind it. The cost-reporting code written for the new DigitalOcean cluster had been checking for its billing credential under the wrong environment-variable name since the day it was written, in every environment. Because that lookup silently failed, it fell back to a fixed estimate instead of ever reading DigitalOcean’s real billing figures, a gap that stayed invisible until someone went looking for why the dead Neon collector kept throwing errors nobody could see.

Untangling that credential mismatch also surfaced a billing-shape detail nobody had designed for: a single DigitalOcean managed-database resource shows up as more than one line item on the real invoice, with compute and storage billed separately under the same group label. A naive one-row-per-line approach would have quietly thrown away part of the real cost rather than double-counting it, so the fix had to account for that shape directly rather than assume one resource always means one line.

PgBouncer and the Postgres Database Migration’s New Reality

As real traffic moved onto the new cluster, the team adopted PgBouncer connection pooling in front of it for every production and staging database — SimpleWP was the first project on the shared cluster to adopt pooling at all. The schema-migration tool was deliberately kept on a direct, unpooled connection throughout, since it needs session-level locks that a pooled connection cannot hold.

Shipping that pooling fix to production surfaced one more dead dependency on Neon: a separate, unrelated pre-deployment backup step had quietly kept trying to snapshot a database branch on Neon before every production and staging deploy. Once that Neon project had zero branches left to snapshot, the step failed with a permanent error every time, and it sat in the way of the real pooling fix reaching production until the dead backup step itself was removed.

The push to add pooling had a concrete trigger: within days of the first database moving onto the new cluster, every connection slot the team had originally budgeted for was already in use elsewhere, and a routine deploy’s migration step failed as a result. Pooling, plus some added headroom on the cluster itself, closed that gap for good.

What’s Different Now

Looking back across the whole postgres database migration, the common thread is not any single mistake. It is that a migration graduates in stages: the application layer first, then the tooling that sits beside it, then the monitoring and billing code that watches both from a distance. Each stage can look finished on its own while the next one is still quietly pointing at the old world.

None of these loose ends were caused by Neon, DigitalOcean, Liquibase or PgBouncer behaving badly. Every one of them traces back to SimpleWP’s own configuration, code or process, while every platform involved behaved exactly as documented. What they share is a pattern: a migration that looks finished at the application layer can still leave a tooling layer, a monitoring job or a billing integration quietly pointing at the wrong thing for weeks.

What changed as a result is simple to state, even though it took real digging to get there. The migration tool’s configuration points at the right host. The dead Neon cost job is gone. The DigitalOcean billing credential is read correctly. Pooling sits in front of every production and staging database on the cluster. Closing a postgres database migration well means going back and checking the parts that were never the headline, and SimpleWP now treats that as a matter of habit, with every follow-up fix visible on our changelog. This same postgres database migration story sits alongside SimpleWP’s own static egress IP relay, a different piece of infrastructure from around the same stretch of work.

If engineering credibility like this matters to how you pick a hosting partner, see SimpleWP’s plans and judge for yourself.