InfrastructureAWS ca-central-1Proposed — not built
2026-10-08
AWS production, at a glance
What the RequiemOS production stack on AWS looks like, how a request reaches it, how data is protected, and how a blue/green deploy swaps one version for the next without downtime.
Written for Zareef. This is the design in aws-production-sizing-scenarios.md v0.4 and railway-to-aws-migration.md, drawn on one page. Nothing here is built yet. Production goes live on Railway first; the move to AWS has a go / no-go when a quiet window opens, and must be done before the first non-Kearney paying customer or the Mortware import.
The short version
3
availability zones in Montreal (ca-central-1). Losing one costs a third of capacity, not the service.
~$869
a month for production at 50 tenants (Scenario A, year 1, committed). ~$519 for a single-AZ pilot.
$0
extra for blue/green. Green exists only for the hour of a deploy, beside blue.
0
servers to patch. ECS Fargate runs the containers; AWS owns the hosts.
The stack
From a director's browser to the database
Read top to bottom. Every request passes through Cloudflare; the load balancer accepts traffic from nowhere else. The app runs in private subnets across three zones, and nothing in the data tier is reachable from the internet.
Funeral directors & staffbrowser · phone · station scanner
Familiesfamily-facing links (ADR-055)
Platform staffplatform console · ECS Exec
HTTPS
Cloudflare — Pro plan (ADR-008)DNS · TLS · CDN · DDoS · managed WAF rules · maintenance-page rule for the cutover · no AWS WAF behind it
only Cloudflare gets through — security group limited to Cloudflare's ranges + Authenticated Origin Pulls (mTLS)
AWS · ca-central-1 (Montreal)
VPC · 3 availability zones
NAT Gateway ×1outbound only
Application Load Balancerpublic subnets · two target groups (blue / green) · staging on its own hostname rule
VPC endpointsS3 · ECR · Logs · SSM · KMS
Zone A
private · app
app-core web taskFargate · 1 vCPU / 2 GB · Puma
Stirling-PDF taskdocument fill / stamp / flatten · internal only
Zone B
private · app
app-core web taskFargate · 1 vCPU / 2 GB
Worker tasks ×2budgeted, not planned — no job queue (OQ-0421)
Zone C
private · app
app-core web taskFargate · 1 vCPU / 2 GB
Stagingsame VPC & ALB · 1 app + 1 Stirling task · never production data
private · data
RDS PostgreSQL 17 — Multi-AZdb.m7g.large (non-burstable) · primary in one zone, synchronous standby in another · failover 60–120 s · encrypted (KMS) · 300 GB gp3 at Scenario A
Monthly snapshots — 90 days90-day ceiling: erasure must reach every copy
Copy to ca-west-1 (Calgary) — 30 daysregion-loss recovery, still in Canada
Offsite logical backupoutside AWS · survives losing the account (OQ-0404)
Missing-backup alarma silent backup failure is worse than none
Delivery
GitHub ActionsCI on every PR · deploy job · promote workflow
OIDC → IAM deploy roleno AWS keys stored in GitHub
Terraformplan on PR · apply on merge · S3 state
Stays on Railway
PR environmentsseeded per pull request
Sandboxesper developer
Demosynthetic Kearney data
No production data, everown credentials · can't read the prod bucket
Zones are drawn as A / B / C; task placement across zones is ECS's job, not fixed. Task counts are Scenario A (50 tenants). Never fewer than two web tasks — first calls happen at 03:00.
Deploys
How blue/green works here
Blue is the version serving users. Green is the new one. ECS starts green beside blue on the same load balancer, checks it, moves traffic over, then drains blue. There is no second stack — green lives for about an hour, so it costs cents, not $480–700 a month.
Blue target group — v41 (live)
v41
v41
v41
Keeps serving until green proves healthy. Drained, not killed: in-flight requests finish.
ALB listenershifts traffic between the two target groups
blue
green
Automatic rollback to blue if a CloudWatch alarm fires during the bake.
Green target group — v42 (new)
v42
v42
v42
Same image digest that staging ran. Health-checked, and optionally smoke-tested through a test listener, before any user sees it.
both colours talk to the same database
One RDS PostgreSQL 17 — shared by blue and greena migration runs once, before green starts — so both versions must work with the new schema
1
Build
Merge to main, CI green, image pushed to ECR by digest.
blue serving
2
Release step
db:prepare → env:provision → env:doctor as a one-off task. Fails? Stop here; blue untouched.
gate
3
Start green
New tasks start beside blue. Task count briefly doubles.
blue serving
4
Bake
Health checks and alarms watched for a set period. Any alarm: roll back.
gate
5
Shift
The listener moves traffic from blue to green.
green serving
6
Drain blue
Blue finishes in-flight requests and stops. Green is the new blue.
green serving
The shared database is the real risk
Blue/green protects against a bad application release. It does not protect against a bad migration, because both colours use one database. The safety net is expand/contract: add the new column, deploy, backfill, switch over, and drop the old column only in a later release (OQ-0406).
Green is not staging
Green runs new code against production data and then becomes production. QA and Kearney UAT happen on the separate staging environment, which never holds production data. Production ships the exact image digest staging ran.
Risky schema changes
For the rare change expand/contract can't cover, use RDS Blue/Green Deployments on demand: AWS builds a synced copy of the database, you switch over, and it is billed only for that window.
Delivery
From a merge to production
Close to today's Railway flow: merge, CI, deploy. The one new step is a deliberate promotion to production.
automatic
Merge to main
The same CI as today: scans, lint, client build, e2e, RSpec.
automatic
Build & push
OIDC sign-in, Docker build, push to ECR. Record the digest.
automatic
Deploy staging
Release step against staging, then blue/green.
manual
Promote
"Promote to production" with the digest staging is running.
automatic
Deploy production
Release step against production, blue/green, release tag.
Tools: GitHub Actions, Terraform, the AWS CLI and four standard AWS actions. No CodePipeline, CodeBuild or Jenkins. Rollback: automatic on a failed bake, otherwise re-promote the previous digest. Detail: CI-CD Operations Guide.md §4.
Where everything runs
Environments and what they cost
Environment
Home
Data
Production
AWS · Multi-AZ
Real customer data
Staging
AWS · same VPC
Synthetic only
Import validation
AWS · ephemeral
Legacy export, destroyed at sign-off
PR environments
Railway
Seeded demo data
Sandboxes
Railway
Developer-owned
Demo
Railway
Synthetic Kearney data
Monthly, Scenario A, year 1
USD
AWS production (committed)
$869
AWS staging (always on)
$159
GitHub — Team + Actions
$76
Railway — PR envs, sandboxes, demo
$60
Platform and delivery
~$1,164
+ DocuSeal (remote-only → full adoption)
$103–437
List prices, USD, pre-tax, deliberately biased high. Database and backups are ~80% of the AWS bill and run 24/7; compute is the cheap part. Source: aws-production-sizing-scenarios.md §6.
Status
Decided, and still open
Decided
ECS Fargate, not EC2 + Kamal — no hosts to patch
Multi-AZ RDS from the first paying tenant; non-burstable m7g
Cloudflare Pro in front, no AWS WAF; ALB locked to Cloudflare
Blue/green inside one stack, not a standing parallel environment
PostgreSQL 17 everywhere (CI moving to 17 — #344)
Staging never holds production data
Demo, PR environments and sandboxes stay on Railway
Move before the first non-Kearney customer or the Mortware import
Still open
Secrets: AWS Secrets Manager or SSM Parameter Store
Staging: own database instance (proposed) — OQ-0422
Seven-year logs: CloudWatch or S3 — OQ-0128
Offsite backup mechanism at scale — OQ-0404
Does ECS Express Mode simplify Terraform? — OQ-0409
Job queue (Solid Queue) in Phase 2? — OQ-0421
Which Railway environment becomes production first (roadmap V5)