Written from direct SSH inspection of both real clusters (node1-3 prod HA, dev-k8s-1-3 dev) done while building the GitLab CI + ArgoCD pipeline (MR !1) - not re-derived from the aspirational docs/manifests that predated that inspection. ARCHITECTURE.md: system diagram (clients, both clusters, Stalwart, EJBCA CA, the CI+ArgoCD flow) plus the storage-coupling fact that everything else hinges on - 4 RWO PVCs + strategy:Recreate is why the app is single-replica today. SANDBOX-DEV-MANUAL.md: day-to-day branch/MR/CI/ArgoCD flow, one-time bootstrap, troubleshooting, and what's explicitly out of scope for normal dev work (the CA, the still-inert prod overlay). PRODUCTION-SCALE-OUT-PLAN.md: phased path to a 100k+-user production deployment on node1-3 - breaking the storage coupling first (rook-ceph CephFS RWX as the fast path, migrating mutable state into the already-installed-but-unused CNPG Postgres as the correct one), then autoscaling, Stalwart's own scaling track, networking/edge, the observability gap (none found on either cluster), security hardening, load testing, DR, and the go-live sequence. Includes a "scale at any time" manual lever, not just HPA.
9.0 KiB
VNCmail+ — Production Scale-Out Plan (target: 100k+ users, scalable on demand)
Goal: take VNCmail+ from "doesn't exist on node1-3" to a production
deployment that can grow past 100k users and be scaled at any time —
both automatically (load-driven) and on a single manual command (ahead of an
expected spike), not just reactively.
Read ARCHITECTURE.md first, especially "The storage coupling" section — it's the reason this is phased the way it is.
Where things stand today (verified by direct inspection, not assumed)
node1-3is a healthy 3-node HA microk8s cluster (rook-ceph, traefik, metallb, cert-manager) with zero application workloads and zero ClusterIssuers. It's a clean slate, not a half-finished deployment.rook-cephis already there, with bothceph-rbd(RWO) andceph-cephfs(RWX, distributed) StorageClasses available — the key piece that makes multi-replica VNCmail+ possible without inventing new infrastructure.cnpg-system(CloudNativePG, a Postgres operator) is already installed on both clusters and currently unused by anything. This is the natural home for the app's mutable state once it moves off local files (Phase 1).- No Prometheus/Grafana/logging stack was found on either cluster — this is a real gap, not a "probably fine," and it's a prerequisite for safe autoscaling (HPA needs a metrics pipeline) and for running anything at 100k-user scale with any visibility into it.
- Stalwart's own scaling story is not covered here — it's a separate system owned by the backend/infra side of this decision. It's called out explicitly at each phase below because VNCmail+ scaling is moot if Stalwart can't handle the same load; plan the two together, not sequentially.
Phase 1 — Break the storage coupling (blocking; do this first)
Today: 4 RWO PVCs, strategy: Recreate, one pod max, ever. Two ways to fix,
pick based on how much time you have before you need >1 replica:
Tactical (fast, days): switch the 4 PVCs to the ceph-cephfs StorageClass
(RWX) and the Deployment strategy to RollingUpdate. This alone unblocks
multiple replicas with no code changes. Real risk: admin-state/telemetry
are multi-writer files on a shared filesystem — fine at low write volume
(login timestamps, audit log, version-check state), but it's a shortcut, not
the target architecture. settings (per-user, keyed by hash(username:serverUrl))
has no cross-writer conflict risk since each user only ever writes their own
file — this one is safe on RWX indefinitely.
Structural (correct, weeks): migrate admin-state and telemetry into
CNPG Postgres (already installed, unused) — proper multi-writer semantics,
no filesystem-locking edge cases, and it's the natural place for this kind
of low-volume operational state anyway. Keep admin (config) as a
read-only mount after setup — ADMIN_CONFIG_READONLY=true is already a
supported mode (lib/admin/paths.ts), so this can be baked into the image
or a ConfigMap at deploy time instead of a writable volume at all. settings
can either stay on CephFS RWX (it's genuinely safe there) or also move to
Postgres if you want zero PVCs in the final state.
Either way: this is the one item that has to happen before Phase 2 means anything. Everything downstream assumes replicas > 1 is possible.
Phase 2 — Autoscaling & headroom
- Install a metrics pipeline (
metrics-serverat minimum for HPA; Prometheus+Grafana for real visibility — see Phase 5, do it once, not twice). HorizontalPodAutoscaleron CPU/memory to start; revisit with a custom metric (JMAP request rate, active WebSocket/SSE connections) once you have real traffic shape.PodDisruptionBudgetso rolling updates and node maintenance don't drop below your minimum replica count.- Re-size
resources.requests/limitsfrom real load-test numbers (Phase 7) — the sandbox's100m/256Mirequests are sandbox-appropriate, not production-appropriate; don't carry them forward by default. - The "scale at any time" requirement: HPA covers load-driven scaling,
but also document (and rehearse once) a single manual command to add
capacity ahead of a known event, before HPA would react:
kubectl -n vncmail-prod scale deploy/vncmail-plus --replicas=Nor bumping the HPA'sminReplicas. This should be a one-line runbook entry, not something someone has to figure out under pressure.
Phase 3 — Stalwart scaling (parallel track, not this repo's code)
VNCmail+ has no database and does no caching of its own — every request is a live JMAP call to Stalwart. At 100k users, Stalwart's own architecture decision matters as much as anything in this repo:
- Storage backend: Stalwart supports RocksDB (single-node) or FoundationDB (distributed, HA) — FoundationDB is the one that scales past a single node.
- Blob storage: point Stalwart's message-blob storage at an S3-compatible backend — rook-ceph's object gateway (RGW), if enabled, is already sitting on the same cluster.
- Confirm Stalwart's own capacity plan (connections, IOPS, memory) against the same 100k-user target this doc is aiming for, ideally before Phase 7's load test, not after it fails.
Phase 4 — Networking & edge
- Create a real
ClusterIssueronnode1-3— none exists today. Decide ACME account + DNS-01 or HTTP-01 solver before anything else in this phase. - Decide the real production hostname (still an open decision — see
deploy/k8s/overlays/prod/patch-ingress.yaml's placeholder). - Rate limiting at the Traefik ingress (a
MiddlewareCRD) before opening up publicly at this scale — nothing enforces this today. - Consider a CDN in front of
_next/staticand other cacheable assets to keep origin load down as user count grows.
Phase 5 — Observability
Stand up Prometheus + Grafana (or point at existing org tooling if one already covers this cluster — worth checking before installing a second stack) before Phase 2's HPA and before Phase 7's load test — you need to see what's happening in both. At minimum: request rate/latency/error rate per pod, JMAP call latency to Stalwart, PVC/CephFS I/O if Phase 1 went the tactical route, and alerting on pod restarts / ImagePullBackOff / cert expiry.
Phase 6 — Security hardening
NetworkPolicyforvncmail-prod, mirroring thevnc-canamespace's existing default-deny-plus-narrow-allow pattern — nothing enforces network isolation forvncmail-prodtoday.- Confirm the microk8s CNI on
node1-3actually enforcesNetworkPolicy(Calico does, flannel-without-a-policy-plugin silently doesn't — thevnc-caREADME already flags this exact trap, re-verify for this namespace too rather than assuming). - Image scanning in the CI build stage.
- S/MIME CA promotion to prod is its own separate, human-only runbook
(
deploy/k8s/ca/README.md§9) — sequence it, don't bundle it into this plan's steps.
Phase 7 — Load testing & capacity planning
Model the actual target before guessing replica counts: concurrent users, JMAP poll/push connection count, expected sync volume per user, attachment upload size/frequency. Run a load test against a prod-shaped deployment (real storage backend from Phase 1, real Stalwart capacity from Phase 3, HPA from Phase 2) before the real cutover — a load test against the sandbox's single-hostPath-replica setup would tell you nothing useful about 100k users.
Recommend a staged ramp for the actual rollout (soft-launch a cohort → watch Phase 5's dashboards → widen) rather than a single cutover to the full 100k target on day one.
Phase 8 — Backup & DR
- rook-ceph snapshot policy for whatever PVCs remain after Phase 1.
- Stalwart's own backup strategy (backend-owned, but confirm it exists and is tested — a mail server's data loss is a much worse incident than this app's).
- A written, rehearsed restore runbook — not just "backups exist."
Go-live sequence (once Phases 1–6 are actually done, not just planned)
- Register
node1-3as an ArgoCD-managed cluster (argocd cluster add, or an equivalent ServiceAccount+kubeconfig secret) — not done yet, and deliberately not done before this point. - Fill in the real values in
deploy/k8s/overlays/prod/(hostname, prod Stalwart'sJMAP_SERVER_URL) and applydeploy/argocd/vncmail-prod-app.yaml. - Create the real
vncmail-envsecret + registry pull secret invncmail-prod, by hand, same as dev's one-time bootstrap. - Merge
dev→main(fast-forward only —git log dev..mainmust be empty first). - Click Sync on
vncmail-prodin the ArgoCD UI. This stays a permanent manual gate — there is no plan to automate this step, ever. - Smoke test against the real hostname, watch Phase 5's dashboards, then proceed with Phase 7's staged ramp.
Nothing in Phases 1–8 requires the go-live sequence to happen first — build
and verify the scaling story in isolation (e.g. on dev-k8s at smaller
scale, or in a throwaway prod-shaped namespace) before the actual cutover.