Seven hours of planned downtime sounds like a lot. Five of them went to the scripts that moved data and to checking their results. Deploying the code itself was a short part of the window, and that was the part I expected least.
3DPCC started as a 3D printing cost calculator. Rebuilding it into a production management system required a new architecture and moving live user data into a different structure. The old dashboard ran on Netlify and performed some operations directly in Supabase. The new version is a frontend on Vercel, a Node.js backend on Railway, and the same Supabase project, which still handles user accounts, the PostgreSQL database, files, and subscription functions. Instead of swapping the production project for a staging one, I extended the existing one additively. Auth accounts, user identifiers, active sessions, Storage files, and the Stripe links stayed where they were.
Below is the plan for that migration, how the window went, and what I took from it.
Why the downtime was planned
The old frontend wrote data straight to Supabase, so putting up a maintenance page froze nothing in practice. The block had to cover the interface, direct writes by the anon and authenticated roles, and privileged functions. Without a window there was a risk that someone would add a record to the old structure at the exact moment I was rewriting it. The break gave me a consistent state and the ability to check every step before the next one started.
The backup that had to work
The project ran on the Free plan for most of its life, without managed backups, so I wrote my own script that makes a logical dump. Before it starts, it verifies the production project identifier, rejects staging and directories inside the repository, and takes the connection string through a masked prompt instead of command line arguments. The result is a set of roles.sql, schema.sql and data.sql with a SHA-256 manifest. The Storage file bytes are not in that copy, they stay in the same project and need to be copied separately.
The day before the migration we moved to the Pro plan, which added managed Supabase backups. Inside the window itself I made one more final dump, after writes were already frozen, so that the restore point would hold a consistent state. A failure of that backup would abort the whole operation before any DDL ran.
A dump means little until you check that you can come back from it. The restore drill script verifies the checksums, brings up a throwaway local Postgres on random ports, restores the database in a single transaction with ON_ERROR_STOP, checks foreign keys, non-validated constraints, and orphaned records, and finally removes the containers and the volume. Two such drills, on August 30 and September 23, still ran on the Free plan and took about 54 and 40 seconds, restoring 22 public tables and the Storage metadata. So I knew from measurement, not from a hunch, that the backup would not be the main part of the window.
The write freeze and its limits
I blocked writes with reversible statement level triggers placed on every table in the public schema. They refuse INSERT, UPDATE, DELETE and TRUNCATE for the anon, authenticated and service_role roles, and for processes running under other roles. I left the postgres and supabase_admin connections open for the controlled migration, which meant that every backend deploy using the production DATABASE_URL had to be stopped before the freeze. The freeze script is idempotent, so I ran it again after the migrations and it covered the new tables, which raised the number of protected relations from 22 to 120. The status command is fail-closed and detects any mismatch between the number of tables and the number of guards, and unfreeze removes everything in a single transaction and does not accept partial success.
The block had a limit I could not get past. The auth and storage schemas belong to other Supabase owners, so a trigger could be placed there, but it could not be removed safely. I deliberately left them alone. Instead I disabled registration in Auth, paused the Supabase crons, and forbade direct uploads to Storage. Stripe webhooks were stopped for the window, and their events had to stay replayable and idempotent, so that none would be lost and none processed twice.
The sequence inside the window looked like this: freeze writes, final backup, migrations as one job, idempotent backfills with validations, backend deploy. It roughly matched the plan.
The audit that made me roll back two things
An independent migration audit upheld the no-go decision and caught two wrong assumptions of mine.
The first one concerned duplicate SKUs. I planned to append suffixes to them automatically, but the short_name of filaments is, for many users, a material label rather than a unique SKU. The automation would have changed the meaning of the data. I replaced it with migrations carrying a change map and a transactional repair script that runs after the backfills, fixes only valid records that have no link and a taken SKU, records the old and the new SKU, and aborts if anything is still without a link after the repair. A second run has to produce zero changes.
The second assumption was more convenient for me than for the users: I wanted to push filling in the missing exchange rates onto them. Instead, an NBP batch fetches the latest publication once, rejects the run if a rate is missing for any currency, saves conversion snapshots with a system audit trail, and does not touch the price in the purchase currency. The end condition is remainingConvertible=0.
Five hours on data
About 70% of the window went to running the scripts that move and map data, together with verifying the results. The order was fixed and the pattern was always the same: preview, write, repeat to prove idempotence. This covered the default SKUs, inventory backfills, SKU conflict repair, NBP rate snapshots for printers, the import of the product and variant catalog, and the pricing parameters. The links between materials and the new inventory had full coverage, and repeated runs introduced no changes.
One area I left unresolved on purpose. Some resins have no package capacity recorded, so their quantity cannot be determined unambiguously, and guessing here would mean writing a number pulled out of thin air into a user's inventory. The application flags such records in the list, shows a message after onboarding, and the editor requires the real capacity. The remaining two hours of the window were split between the backup, the freeze, the schema migration, the deploys, setting up environments, and testing.
A test on two real accounts with deliberately broken records
Before the final cutover I created records with the MIG25-TEST- prefix on two existing Pro accounts, reproducing the cases that break a migration most easily: a resin with no package capacity, two filaments with an identical SKU, a printer in a currency other than the organization currency with a lifetime of 0, two hardware items with the same SKU, and a plain filament with an old quote. After the migration I checked whether the resin editor requires the real capacity, whether the duplicate SKUs were separated with names, colors, and stock levels preserved, whether the FX snapshot left the purchase price alone, and whether the old quote is still readable. I stuck to the rule that I do not enter artificial values just to close a message, and I do not work around the validation of the production interface.
CI/CD and bringing up the environments
After the migration I had to stand up the new production environments and make sure that staging and production had the right variables. This was the stage that caused the most trouble in the whole window.
What made it demanding was the number of small settings to get right at once: CORS, environment variables, domains, and webhooks, each of them handled separately for staging and production. Most of the attention went to keeping track of which environment I was working in and whether a given value had landed where it was supposed to.
Syncing staging with main in both repositories went through ordinary pull requests with required green CI, with no force push, and I recorded the commit hashes as frozen versions. The workflow has three independent gates: the static one, meaning types, lint, format, unit tests, and build; integration tests on an ephemeral local Supabase; and a production container build with a liveness check, a non-root user, and shutdown on SIGTERM. The repository does not deploy itself. Railway builds the Dockerfile from the main branch and runs the container directly through node dist/server.js, so that SIGTERM reaches the application, and the healthcheck on /ready checks the database and Redis. The frontend on Vercel is a Vite preset with the dist directory, the API got its own domain api-prod.3dpcc.com, and app.3dpcc.com switches over through a CNAME added only at the end of the window.
No asynchronous process runs inside the API process. Instead there are three separate Cron services on Railway built from the same image, run every 5 minutes: expiring reservations, the event outbox with notifications, and transactional email delivery. Leases and FOR UPDATE SKIP LOCKED let consecutive runs overlap safely.
Tests before reopening traffic
The order of the tests followed from the fact that reads can be checked without risk and writes cannot. First, with the block still active, I went through logging in with an old and a new session, the dashboard, old materials, a printer with FX, an old quote, a product, an image and a PDF, the Pro and Free plans, and the lack of access to other people's data. Only after a positive result did I lift the block and test the full write flow: new quote, reservation, approval, consumption, order. On top of that, registration, subscriptions, invoices and webhooks, data consistency, and the absence of errors.
There is one thing worth remembering about this stage. From lifting the block to switching DNS, anyone with a valid session or an open old tab could write data, because a private Netlify and a protected Vercel do not block direct requests to Supabase. Subscription payments had already gone through a full test in the Stripe Sandbox, so in production I did not pay with a real card, but I opened the production Checkout and confirmed the Managed Payments and automatic tax configuration.
I lowered the TTL of the app record to 60 seconds before the window, and switching DNS to Vercel was the last step of reopening traffic. The old Netlify stayed in Private mode for a few days as a fallback, and during that time I watched whether the new version behaved correctly.
The rollback I did not have to use
Apart from the expected warnings about incomplete data in a few inventory records, which I had a separate screen for, nothing broke. The rollback was prepared anyway. All migrations were expand-only, so the preferred option was going back to the previous deployments, and I left restoring the database strictly for the case of corrupted data, because it means losing everything written after the final backup. An automatic rollback was triggered by, among other things, mismatched counts or orphaned records, an unresolved backfill blocker, a tenant isolation failure, the inability to log in, persistent 5xx errors, and unavailable reads of critical Storage files. If a no-go decision had come before traffic was reopened, the plan was to keep the block in place and stop the Railway services, and with DNS not switched the customers would not have noticed anything.
Next time I will estimate this differently
The largest part of the window went to the data scripts and the verification of their results. In the next operation of this kind I will calculate the downtime from the actual data volume and the number of repairs needed, not from the time it takes to publish the application.
In summary
The whole operation closed in seven hours, inside the six to eight hour range I had planned for. A plan that could be tested before the migration, keeping the production Supabase project instead of swapping it out, a reversible write freeze, and expand-only migrations meant the worst case stayed recoverable the whole time. The rest was working through the steps calmly and checking the result after each one.
3DPCC now runs on the new architecture. If you cost 3D prints or run a print shop and want to see what the system looks like after the rebuild, take a look at 3dpcc.com.


