# VNCmail+ — Production Scale-Out Plan (target: 100k+ users, scalable on demand) Goal: take VNCmail+ from "doesn't exist on `node1-3`" to a production deployment that can grow past 100k users and be scaled **at any time** — both automatically (load-driven) and on a single manual command (ahead of an expected spike), not just reactively. Read [ARCHITECTURE.md](ARCHITECTURE.md) first, especially "The storage coupling" section — it's the reason this is phased the way it is. ## Where things stand today (verified by direct inspection, not assumed) - `node1-3` is a healthy 3-node HA microk8s cluster (rook-ceph, traefik, metallb, cert-manager) with **zero application workloads and zero ClusterIssuers**. It's a clean slate, not a half-finished deployment. - `rook-ceph` is already there, with both `ceph-rbd` (RWO) and **`ceph-cephfs` (RWX, distributed)** StorageClasses available — the key piece that makes multi-replica VNCmail+ possible without inventing new infrastructure. - `cnpg-system` (CloudNativePG, a Postgres operator) is **already installed on both clusters** and currently unused by anything. This is the natural home for the app's mutable state once it moves off local files (Phase 1). - No Prometheus/Grafana/logging stack was found on either cluster — this is a real gap, not a "probably fine," and it's a prerequisite for safe autoscaling (HPA needs a metrics pipeline) and for running anything at 100k-user scale with any visibility into it. - Stalwart's own scaling story is **not covered here** — it's a separate system owned by the backend/infra side of this decision. It's called out explicitly at each phase below because VNCmail+ scaling is moot if Stalwart can't handle the same load; plan the two together, not sequentially. ## Phase 1 — Break the storage coupling (blocking; do this first) Today: 4 RWO PVCs, `strategy: Recreate`, one pod max, ever. Two ways to fix, pick based on how much time you have before you need >1 replica: **Tactical (fast, days)**: switch the 4 PVCs to the `ceph-cephfs` StorageClass (RWX) and the Deployment `strategy` to `RollingUpdate`. This alone unblocks multiple replicas with no code changes. Real risk: `admin-state`/`telemetry` are multi-writer files on a shared filesystem — fine at low write volume (login timestamps, audit log, version-check state), but it's a shortcut, not the target architecture. `settings` (per-user, keyed by `hash(username:serverUrl)`) has no cross-writer conflict risk since each user only ever writes their own file — this one is safe on RWX indefinitely. **Structural (correct, weeks)**: migrate `admin-state` and `telemetry` into CNPG Postgres (already installed, unused) — proper multi-writer semantics, no filesystem-locking edge cases, and it's the natural place for this kind of low-volume operational state anyway. Keep `admin` (config) as a **read-only mount** after setup — `ADMIN_CONFIG_READONLY=true` is already a supported mode (`lib/admin/paths.ts`), so this can be baked into the image or a ConfigMap at deploy time instead of a writable volume at all. `settings` can either stay on CephFS RWX (it's genuinely safe there) or also move to Postgres if you want zero PVCs in the final state. Either way: this is the one item that has to happen before Phase 2 means anything. Everything downstream assumes replicas > 1 is possible. ## Phase 2 — Autoscaling & headroom - Install a metrics pipeline (`metrics-server` at minimum for HPA; Prometheus+Grafana for real visibility — see Phase 5, do it once, not twice). - `HorizontalPodAutoscaler` on CPU/memory to start; revisit with a custom metric (JMAP request rate, active WebSocket/SSE connections) once you have real traffic shape. - `PodDisruptionBudget` so rolling updates and node maintenance don't drop below your minimum replica count. - Re-size `resources.requests/limits` from real load-test numbers (Phase 7) — the sandbox's `100m/256Mi` requests are sandbox-appropriate, not production-appropriate; don't carry them forward by default. - **The "scale at any time" requirement**: HPA covers load-driven scaling, but also document (and rehearse once) a single manual command to add capacity ahead of a known event, before HPA would react: `kubectl -n vncmail-prod scale deploy/vncmail-plus --replicas=N` or bumping the HPA's `minReplicas`. This should be a one-line runbook entry, not something someone has to figure out under pressure. ## Phase 3 — Stalwart scaling (parallel track, not this repo's code) VNCmail+ has no database and does no caching of its own — every request is a live JMAP call to Stalwart. At 100k users, Stalwart's own architecture decision matters as much as anything in this repo: - Storage backend: Stalwart supports RocksDB (single-node) or FoundationDB (distributed, HA) — FoundationDB is the one that scales past a single node. - Blob storage: point Stalwart's message-blob storage at an S3-compatible backend — rook-ceph's object gateway (RGW), if enabled, is already sitting on the same cluster. - Confirm Stalwart's own capacity plan (connections, IOPS, memory) against the same 100k-user target this doc is aiming for, ideally before Phase 7's load test, not after it fails. ## Phase 4 — Networking & edge - Create a real `ClusterIssuer` on `node1-3` — **none exists today**. Decide ACME account + DNS-01 or HTTP-01 solver before anything else in this phase. - Decide the real production hostname (still an open decision — see `deploy/k8s/overlays/prod/patch-ingress.yaml`'s placeholder). - Rate limiting at the Traefik ingress (a `Middleware` CRD) before opening up publicly at this scale — nothing enforces this today. - Consider a CDN in front of `_next/static` and other cacheable assets to keep origin load down as user count grows. ## Phase 5 — Observability Stand up Prometheus + Grafana (or point at existing org tooling if one already covers this cluster — worth checking before installing a second stack) **before** Phase 2's HPA and **before** Phase 7's load test — you need to see what's happening in both. At minimum: request rate/latency/error rate per pod, JMAP call latency to Stalwart, PVC/CephFS I/O if Phase 1 went the tactical route, and alerting on pod restarts / ImagePullBackOff / cert expiry. ## Phase 6 — Security hardening - `NetworkPolicy` for `vncmail-prod`, mirroring the `vnc-ca` namespace's existing default-deny-plus-narrow-allow pattern — nothing enforces network isolation for `vncmail-prod` today. - Confirm the microk8s CNI on `node1-3` actually enforces `NetworkPolicy` (Calico does, flannel-without-a-policy-plugin silently doesn't — the `vnc-ca` README already flags this exact trap, re-verify for this namespace too rather than assuming). - Image scanning in the CI build stage. - S/MIME CA promotion to prod is its own separate, human-only runbook (`deploy/k8s/ca/README.md` §9) — sequence it, don't bundle it into this plan's steps. ## Phase 7 — Load testing & capacity planning Model the actual target before guessing replica counts: concurrent users, JMAP poll/push connection count, expected sync volume per user, attachment upload size/frequency. Run a load test against a **prod-shaped** deployment (real storage backend from Phase 1, real Stalwart capacity from Phase 3, HPA from Phase 2) before the real cutover — a load test against the sandbox's single-hostPath-replica setup would tell you nothing useful about 100k users. Recommend a staged ramp for the actual rollout (soft-launch a cohort → watch Phase 5's dashboards → widen) rather than a single cutover to the full 100k target on day one. ## Phase 8 — Backup & DR - rook-ceph snapshot policy for whatever PVCs remain after Phase 1. - Stalwart's own backup strategy (backend-owned, but confirm it exists and is tested — a mail server's data loss is a much worse incident than this app's). - A written, rehearsed restore runbook — not just "backups exist." ## Go-live sequence (once Phases 1–6 are actually done, not just planned) 1. Register `node1-3` as an ArgoCD-managed cluster (`argocd cluster add`, or an equivalent ServiceAccount+kubeconfig secret) — not done yet, and deliberately not done before this point. 2. Fill in the real values in `deploy/k8s/overlays/prod/` (hostname, prod Stalwart's `JMAP_SERVER_URL`) and apply `deploy/argocd/vncmail-prod-app.yaml`. 3. Create the real `vncmail-env` secret + registry pull secret in `vncmail-prod`, by hand, same as dev's one-time bootstrap. 4. Merge `dev` → `main` (fast-forward only — `git log dev..main` must be empty first). 5. Click **Sync** on `vncmail-prod` in the ArgoCD UI. This stays a permanent manual gate — there is no plan to automate this step, ever. 6. Smoke test against the real hostname, watch Phase 5's dashboards, then proceed with Phase 7's staged ramp. Nothing in Phases 1–8 requires the go-live sequence to happen first — build and verify the scaling story in isolation (e.g. on `dev-k8s` at smaller scale, or in a throwaway prod-shaped namespace) before the actual cutover.