Infrastructure AWS ca-central-1 Proposed — not built
2026-10-08

AWS production, at a glance

What the RequiemOS production stack on AWS looks like, how a request reaches it, how data is protected, and how a blue/green deploy swaps one version for the next without downtime.

Written for Zareef. This is the design in aws-production-sizing-scenarios.md v0.4 and railway-to-aws-migration.md, drawn on one page. Nothing here is built yet. Production goes live on Railway first; the move to AWS has a go / no-go when a quiet window opens, and must be done before the first non-Kearney paying customer or the Mortware import.
3
availability zones in Montreal (ca-central-1). Losing one costs a third of capacity, not the service.
~$869
a month for production at 50 tenants (Scenario A, year 1, committed). ~$519 for a single-AZ pilot.
$0
extra for blue/green. Green exists only for the hour of a deploy, beside blue.
0
servers to patch. ECS Fargate runs the containers; AWS owns the hosts.

From a director's browser to the database

Read top to bottom. Every request passes through Cloudflare; the load balancer accepts traffic from nowhere else. The app runs in private subnets across three zones, and nothing in the data tier is reachable from the internet.

Edge Network & compute Database Storage & backups Delivery & off-AWS

How blue/green works here

Blue is the version serving users. Green is the new one. ECS starts green beside blue on the same load balancer, checks it, moves traffic over, then drains blue. There is no second stack — green lives for about an hour, so it costs cents, not $480–700 a month.

Blue target group — v41 (live)

v41
v41
v41

Keeps serving until green proves healthy. Drained, not killed: in-flight requests finish.

ALB listenershifts traffic between the two target groups

Automatic rollback to blue if a CloudWatch alarm fires during the bake.

Green target group — v42 (new)

v42
v42
v42

Same image digest that staging ran. Health-checked, and optionally smoke-tested through a test listener, before any user sees it.

both colours talk to the same database
One RDS PostgreSQL 17 — shared by blue and greena migration runs once, before green starts — so both versions must work with the new schema
1

Build

Merge to main, CI green, image pushed to ECR by digest.

blue serving
2

Release step

db:prepare → env:provision → env:doctor as a one-off task. Fails? Stop here; blue untouched.

gate
3

Start green

New tasks start beside blue. Task count briefly doubles.

blue serving
4

Bake

Health checks and alarms watched for a set period. Any alarm: roll back.

gate
5

Shift

The listener moves traffic from blue to green.

green serving
6

Drain blue

Blue finishes in-flight requests and stops. Green is the new blue.

green serving

The shared database is the real risk

Blue/green protects against a bad application release. It does not protect against a bad migration, because both colours use one database. The safety net is expand/contract: add the new column, deploy, backfill, switch over, and drop the old column only in a later release (OQ-0406).

Green is not staging

Green runs new code against production data and then becomes production. QA and Kearney UAT happen on the separate staging environment, which never holds production data. Production ships the exact image digest staging ran.

Risky schema changes

For the rare change expand/contract can't cover, use RDS Blue/Green Deployments on demand: AWS builds a synced copy of the database, you switch over, and it is billed only for that window.

From a merge to production

Close to today's Railway flow: merge, CI, deploy. The one new step is a deliberate promotion to production.

automatic

Merge to main

The same CI as today: scans, lint, client build, e2e, RSpec.

automatic

Build & push

OIDC sign-in, Docker build, push to ECR. Record the digest.

automatic

Deploy staging

Release step against staging, then blue/green.

manual

Promote

"Promote to production" with the digest staging is running.

automatic

Deploy production

Release step against production, blue/green, release tag.

Tools: GitHub Actions, Terraform, the AWS CLI and four standard AWS actions. No CodePipeline, CodeBuild or Jenkins. Rollback: automatic on a failed bake, otherwise re-promote the previous digest. Detail: CI-CD Operations Guide.md §4.

Environments and what they cost

EnvironmentHomeData
ProductionAWS · Multi-AZReal customer data
StagingAWS · same VPCSynthetic only
Import validationAWS · ephemeralLegacy export, destroyed at sign-off
PR environmentsRailwaySeeded demo data
SandboxesRailwayDeveloper-owned
DemoRailwaySynthetic Kearney data
Monthly, Scenario A, year 1USD
AWS production (committed)$869
AWS staging (always on)$159
GitHub — Team + Actions$76
Railway — PR envs, sandboxes, demo$60
Platform and delivery~$1,164
+ DocuSeal (remote-only → full adoption)$103–437

List prices, USD, pre-tax, deliberately biased high. Database and backups are ~80% of the AWS bill and run 24/7; compute is the cheap part. Source: aws-production-sizing-scenarios.md §6.

Decided, and still open

Decided

  • ECS Fargate, not EC2 + Kamal — no hosts to patch
  • Multi-AZ RDS from the first paying tenant; non-burstable m7g
  • Cloudflare Pro in front, no AWS WAF; ALB locked to Cloudflare
  • Blue/green inside one stack, not a standing parallel environment
  • PostgreSQL 17 everywhere (CI moving to 17 — #344)
  • Staging never holds production data
  • Demo, PR environments and sandboxes stay on Railway
  • Move before the first non-Kearney customer or the Mortware import

Still open

  • Secrets: AWS Secrets Manager or SSM Parameter Store
  • Staging: own database instance (proposed) — OQ-0422
  • Seven-year logs: CloudWatch or S3 — OQ-0128
  • Offsite backup mechanism at scale — OQ-0404
  • Does ECS Express Mode simplify Terraform? — OQ-0409
  • Job queue (Solid Queue) in Phase 2? — OQ-0421
  • Which Railway environment becomes production first (roadmap V5)