An end-to-end audit found the repo could not build, test or run as shipped. This fixes every finding, then adds a Cloud Run track so the demo costs about £1/month idle instead of ~£150. CI (red on its first run) - api: setuptools could not build the package (flat layout with app/ and alembic/) - web: missing @types/node; `vitest run` exited 1 with no test files - pipeline: the stub run needed a gitignored VCF, and no process had a stub block - ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI) ML serving (scores were meaningless) - the registered model now carries its own feature engineering and returns predict_proba, so serving sends raw columns and cannot drift from training - resolve by registry alias (stages are deprecated in MLflow 3) and record the real version; re-scoring upserts instead of failing on the unique constraint - ClinVar labels parsed from VEP's lowercase terms Pipeline - exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent - job status reaches running/failed/succeeded, so the UI stops polling dead jobs - DATABASE_URL travels in the environment or a Nextflow secret, never on a command line - VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch Deployment - the API serves /api (matching the ingress); the web app reads its API URL at runtime - migrations run in an init container under a Postgres advisory lock - terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload Identity bindings, Secret Manager, deletion protection - serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md. Correctness and security - 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider VEP text columns, enum dropped on downgrade, the sample's assembly actually used - vcf_uri restricted to gs:// objects or files under the data root, blocking option injection - CORS restricted to configured origins; `make down` no longer deletes volumes Data - docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar, gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform validate and both kustomize overlays clean.
4.7 KiB
Architecture
flowchart LR
U[Scientist] -->|browser| W[SvelteKit web]
W -->|REST /api| A[FastAPI]
A --> P[(PostgreSQL / Cloud SQL)]
A -->|publish vcf-uploaded| Q[Pub/Sub]
Q --> E[Argo Events sensor]
E --> AW[Argo Workflow: Nextflow driver]
AW -->|tasks| B[Google Batch: bcftools norm, VEP, load_db]
B -->|reads VCF, VEP cache| G[(GCS bucket)]
B -->|writes variants, marks job succeeded| P
AW -.->|exit handler marks job failed| P
A -->|models:/rarelens-pathogenicity@production| M[MLflow registry]
T[ml/train.py] --> M
GH[GitHub Actions] -->|images via WIF| AR[Artifact Registry]
GH -->|bumps overlay tags| R[(git: infra/k8s/overlays/gcp)]
R --> CD[ArgoCD] --> K[GKE Autopilot]
Two deployment tracks
The same images and the same pipeline, deployed two ways (infra/terraform/variables.tf):
| Serverless (default) | Kubernetes (-var deploy_kubernetes=true) |
|
|---|---|---|
| api, web | Cloud Run, scale to zero | Deployments behind an ingress |
| dispatch | the API executes a Cloud Run job | Pub/Sub -> Argo Events -> Argo Workflow |
| pipeline tasks | Google Batch | Google Batch |
/api routing |
the web service proxies it | the ingress routes it |
| idle cost | ~£1/month | ~£130+/month |
app.services.events.launch() picks the dispatch backend from configuration: a Cloud Run job when
CLOUDRUN_JOB is set, Pub/Sub when PUBSUB_TOPIC is, and a local Nextflow process otherwise.
See cloud.md for why the serverless one is the default.
Why these choices
One monorepo. The four components share a schema (variants table, feature columns) and the
point of the exercise is to see them evolve together. Separate repos would hide the coupling.
Nextflow for the science, Argo Workflows for the trigger. Nextflow is the lingua franca for
bioinformatics pipelines. Argo is what the platform team already runs. So Argo owns when a
pipeline runs; Nextflow owns what it does. The API never talks to Kubernetes directly; it
publishes an event and gets on with its life. The Nextflow driver runs in the Argo pod and sends
each task to Google Batch: a gs:// work directory needs an executor that stages through GCS
(Nextflow's Kubernetes executor needs a shared ReadWriteMany volume instead).
Job lifecycle. The API creates a job as running once the pipeline is dispatched, or failed
with the reason in jobs.log when dispatch is impossible. The loader marks it succeeded in the
same transaction that stores the variants. Anything else (Nextflow error, eviction) is caught by
the Argo exit handler, or locally by the API watching the Nextflow process, and marked failed,
so the UI never polls a dead job.
Variant identity. NORMALISE sets each VCF ID to CHROM_POS_REF_ALT; VEP echoes it as
Uploaded_variation and the loader takes exact VCF alleles from it, because VEP's own
Location/Allele columns trim indel alleles.
FastAPI + Pydantic v2 + SQLAlchemy 2.0 async. Typed at both boundaries: request/response models and ORM models are separate on purpose so the database can change without breaking the frontend contract. Alembic owns the schema; the loader script writes raw SQL against that schema, not the ORM, because the pipeline container should not import the API.
SvelteKit. Small runtime, no virtual DOM, and Svelte 5 runes make server-driven state simple.
The UI has exactly two pages; the goal is a table a scientist actually wants to filter, not a dashboard.
PUBLIC_API_URL is read at runtime, so one image works behind the ingress (/api) and elsewhere.
GKE Autopilot + Cloud SQL, not self-managed. The lab is about the platform patterns (Workload Identity, GitOps, Kustomize overlays, private networking), not about running etcd. Cloud SQL has only a private IP; the API reaches it through a Cloud SQL Proxy sidecar, pipeline tasks directly in the VPC. Database URLs live in Secret Manager.
GitOps. CI builds and tests; it never runs kubectl apply. It edits image tags in the gcp overlay
and ArgoCD reconciles. Rollback is git revert. Migrations run in an init container under a Postgres
advisory lock, so replicas starting together migrate once.
MLflow registry as the model contract. The API loads whichever version the production alias
points at and records that version on every prediction. The registered model is a pyfunc that owns its
feature engineering (rarelens_ml.features ships inside it) and returns P(pathogenic), so serving only
sends raw columns and cannot drift from training.
What is deliberately missing
Authentication, PHI handling, audit logs, clinical validation. This is a learning platform on public data. Adding Identity-Aware Proxy in front of the ingress is the first step if that ever changes.