# Cloud choice: Google Cloud, with a documented AWS escape hatch Decided 2026-09-12. Scope: `infra/terraform/`, `infra/k8s/`, `pipeline/nextflow.config`. ## Decision rarelens deploys to **Google Cloud**. AWS was the serious alternative, and it is genuinely better on two points (below), but not by enough to justify rebuilding an estate that already works. ## Why Google Cloud | Reason | Detail | |---|---| | One Kubernetes control plane is effectively free | GKE's free tier gives $74.40/month in credits per billing account, which covers one Autopilot or zonal cluster. EKS charges $0.10/hour per cluster (~$73/month) with no equivalent credit. For a self-funded lab this is the largest fixed monthly difference. | | The executor question is already settled here | Google retired Cloud Life Sciences on 8 July 2025; Batch is its successor, and Nextflow upstream moved to Google Batch in April 2025. `pipeline/nextflow.config` uses `google-batch`, which is the supported path rather than a legacy one. | | The estate exists and is verified | Terraform (custom VPC, private Cloud SQL, Workload Identity, Secret Manager, Batch IAM), Kustomize overlays, Argo Workflows/Events and CI all render, validate and pass tests today. Rebuilding this on AWS costs 1–2 weeks and mostly repeats learning already banked. | | No data lock-in | Every dataset the platform uses is readable from either cloud (see [data.md](data.md)): gnomAD publishes to GCP, AWS and Azure; GIAB and 1000 Genomes are open on AWS and NCBI; ClinVar is a plain NCBI download. | ## What AWS is genuinely better at - **Managed Nextflow.** AWS HealthOmics runs Nextflow (up to 26.04), WDL and CWL as a managed service, and is available in London (`eu-west-2`). GCP has no equivalent: you operate the driver yourself, which is exactly what `infra/argo-workflows/annotate.yaml` does. - **UK life-sciences gravity.** The UK Biobank Research Analysis Platform is DNAnexus running on AWS, hosted in the UK. If the aim is to mirror what Cambridge-area employers run day to day, AWS is the more common answer. ## When to revisit this Move to AWS if any of these becomes true: - The lab wants a managed pipeline runner instead of an Argo + Batch driver it maintains. - Matching an AWS-first employer's stack matters more than the two weeks it costs. - The shape changes: several clusters, or enough managed-service spend that one free control plane stops being material. ## Running this on a hobby budget The cloud is not the cost driver; the always-on shape is. Estimates below are list price, and rounded — treat them as orders of magnitude, not quotes. ### What the Kubernetes estate costs at rest GKE Autopilot bills what pods *request*, not what they use, with a per-pod floor (250m vCPU / 512 MiB). The free tier credit covers the cluster fee only, not pod-hours. | Always-on | Requests | ~Monthly (us-central1 rates: $0.0445/vCPU-h, $0.0049/GiB-h) | |---|---|---| | api + web (2 replicas each, incl. Cloud SQL proxy sidecar) | ~1.2 vCPU, ~2.3 GiB | ~$47 | | ArgoCD, Argo Workflows, Argo Events + NATS EventBus (~11 pods at the floor) | ~2.8 vCPU, ~5.5 GiB | ~$110 | | Cloud SQL `db-f1-micro` | — | ~$8–12 | | **Total** | | **~$165–170, London a bit more** | That is the wrong shape for a portfolio that is idle 99% of the time. ### The shape that costs ~£1/month This is what `terraform apply` builds by default (`deploy_kubernetes` and `deploy_cloud_sql` are both `false`). Kubernetes becomes something you switch on to show, not something you rent: | Piece | Service | Idle cost | |---|---|---| | api, web | Cloud Run, `min-instances=0`, `max_instances` capped | £0 — Always Free covers 2M requests, 180k vCPU-s, 360k GiB-s per month | | Nextflow driver | Cloud Run **job**, started by the API through the Jobs API (`roles/run.jobsExecutorWithOverrides`, one job only) | £0 idle, pennies per run | | Pipeline tasks | Google Batch on **Spot** VMs | £0 idle; a chr22 VEP run is a few pence | | Database | Neon free tier (scale-to-zero, 0.5 GB) via `TF_VAR_database_url`, or `-var deploy_cloud_sql=true` | £0 (or ~$8–12) | | Model | pyfunc artifact loaded straight from GCS (`MODEL_URI`), no MLflow server running | £0 | | Storage | GCS + Artifact Registry | ~£1 (VEP cache dominates; Nearline halves it) | Trade-offs worth knowing: Cloud Run cold starts add 1–3 s to the first request after idle; the Cloud Run path drops Pub/Sub, Argo Events and Argo Workflows from the critical path (the API calls the Jobs API directly); and 0.5 GB of Neon does not fit a whole chromosome once `annotations` stores the full VEP record — demo a gene panel, or store only the annotation keys the UI uses. Only the web service needs to be public: it serves the UI and proxies `/api` to the API service (`web/src/routes/api/[...path]`), which is the same shape the ingress gives the Kubernetes track, so the frontend code is identical either way. ### Keep the Kubernetes story, stop paying rent for it `infra/k8s/` and `infra/argo-workflows/` stay in the repo and stay deployable. Bring the estate up with `terraform apply` for an interview or a recording (roughly $0.25/hour while running, so a two-hour demo is small change), then `terraform destroy -var deletion_protection=false`. `make kind` runs the same manifests locally for free. ### Guardrails - A billing budget with alerts at £5/£10, before anything else. - `max-instances` on every Cloud Run service: scale-to-zero protects the floor, a cap protects the ceiling. - Spot VMs for Batch, and the existing 30-day lifecycle rule on `work/` in the bucket. - New accounts get $300 of Google Cloud credit for 90 days, which covers the experimenting phase. ### One more reason not to switch to AWS AWS replaced its 12-month free tier on 15 July 2025 with credits ($100, up to $200) on a Free plan that closes after six months or when the credits run out. Google's Always Free quotas, including Cloud Run's, are permanent. For a demo meant to stay reachable indefinitely at near-zero cost, that difference matters more than any feature comparison above. ## The escape hatch Nextflow is the portability layer: executors are configuration, not code. An AWS run needs a new profile in `pipeline/nextflow.config` (`process.executor = 'awsbatch'`, an S3 work directory and a job queue), or a HealthOmics workflow definition. The processes themselves do not change. Keeping `pipeline/bin/` cloud-agnostic (the scripts read `DATABASE_URL` from the environment, never from a command line) is what keeps that true. Sources: [Migrate to Batch from Cloud Life Sciences](https://docs.cloud.google.com/batch/docs/migrate-to-batch-from-cloud-life-sciences), [GKE pricing](https://cloud.google.com/kubernetes-engine/pricing), [HealthOmics supported languages](https://docs.aws.amazon.com/omics/latest/dev/workflows-supported-languages.html), [HealthOmics Nextflow 26.04](https://aws.amazon.com/about-aws/whats-new/2026/06/aws-healthomics-nextflow-version-26-04/), [UK Biobank Research Analysis Platform](https://www.ukbiobank.ac.uk/use-our-data/research-analysis-platform/).