Files
rarelens/docs/architecture.md
T
Kemal Yaylali 11fb6b3d73 fix: overhaul the platform skeleton, add a serverless deployment track
An end-to-end audit found the repo could not build, test or run as shipped. This
fixes every finding, then adds a Cloud Run track so the demo costs about £1/month
idle instead of ~£150.

CI (red on its first run)
- api: setuptools could not build the package (flat layout with app/ and alembic/)
- web: missing @types/node; `vitest run` exited 1 with no test files
- pipeline: the stub run needed a gitignored VCF, and no process had a stub block
- ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI)

ML serving (scores were meaningless)
- the registered model now carries its own feature engineering and returns predict_proba,
  so serving sends raw columns and cannot drift from training
- resolve by registry alias (stages are deprecated in MLflow 3) and record the real
  version; re-scoring upserts instead of failing on the unique constraint
- ClinVar labels parsed from VEP's lowercase terms

Pipeline
- exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent
- job status reaches running/failed/succeeded, so the UI stops polling dead jobs
- DATABASE_URL travels in the environment or a Nextflow secret, never on a command line
- VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch

Deployment
- the API serves /api (matching the ingress); the web app reads its API URL at runtime
- migrations run in an init container under a Postgres advisory lock
- terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload
  Identity bindings, Secret Manager, deletion protection
- serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job
  for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and
  Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md.

Correctness and security
- 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider
  VEP text columns, enum dropped on downgrade, the sample's assembly actually used
- vcf_uri restricted to gs:// objects or files under the data root, blocking option injection
- CORS restricted to configured origins; `make down` no longer deletes volumes

Data
- docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar,
  gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice

Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform
validate and both kustomize overlays clean.
2026-09-12 07:21:11 +01:00

4.7 KiB

Architecture

flowchart LR
  U[Scientist] -->|browser| W[SvelteKit web]
  W -->|REST /api| A[FastAPI]
  A --> P[(PostgreSQL / Cloud SQL)]
  A -->|publish vcf-uploaded| Q[Pub/Sub]
  Q --> E[Argo Events sensor]
  E --> AW[Argo Workflow: Nextflow driver]
  AW -->|tasks| B[Google Batch: bcftools norm, VEP, load_db]
  B -->|reads VCF, VEP cache| G[(GCS bucket)]
  B -->|writes variants, marks job succeeded| P
  AW -.->|exit handler marks job failed| P
  A -->|models:/rarelens-pathogenicity@production| M[MLflow registry]
  T[ml/train.py] --> M
  GH[GitHub Actions] -->|images via WIF| AR[Artifact Registry]
  GH -->|bumps overlay tags| R[(git: infra/k8s/overlays/gcp)]
  R --> CD[ArgoCD] --> K[GKE Autopilot]

Two deployment tracks

The same images and the same pipeline, deployed two ways (infra/terraform/variables.tf):

Serverless (default) Kubernetes (-var deploy_kubernetes=true)
api, web Cloud Run, scale to zero Deployments behind an ingress
dispatch the API executes a Cloud Run job Pub/Sub -> Argo Events -> Argo Workflow
pipeline tasks Google Batch Google Batch
/api routing the web service proxies it the ingress routes it
idle cost ~£1/month ~£130+/month

app.services.events.launch() picks the dispatch backend from configuration: a Cloud Run job when CLOUDRUN_JOB is set, Pub/Sub when PUBSUB_TOPIC is, and a local Nextflow process otherwise. See cloud.md for why the serverless one is the default.

Why these choices

One monorepo. The four components share a schema (variants table, feature columns) and the point of the exercise is to see them evolve together. Separate repos would hide the coupling.

Nextflow for the science, Argo Workflows for the trigger. Nextflow is the lingua franca for bioinformatics pipelines. Argo is what the platform team already runs. So Argo owns when a pipeline runs; Nextflow owns what it does. The API never talks to Kubernetes directly; it publishes an event and gets on with its life. The Nextflow driver runs in the Argo pod and sends each task to Google Batch: a gs:// work directory needs an executor that stages through GCS (Nextflow's Kubernetes executor needs a shared ReadWriteMany volume instead).

Job lifecycle. The API creates a job as running once the pipeline is dispatched, or failed with the reason in jobs.log when dispatch is impossible. The loader marks it succeeded in the same transaction that stores the variants. Anything else (Nextflow error, eviction) is caught by the Argo exit handler, or locally by the API watching the Nextflow process, and marked failed, so the UI never polls a dead job.

Variant identity. NORMALISE sets each VCF ID to CHROM_POS_REF_ALT; VEP echoes it as Uploaded_variation and the loader takes exact VCF alleles from it, because VEP's own Location/Allele columns trim indel alleles.

FastAPI + Pydantic v2 + SQLAlchemy 2.0 async. Typed at both boundaries: request/response models and ORM models are separate on purpose so the database can change without breaking the frontend contract. Alembic owns the schema; the loader script writes raw SQL against that schema, not the ORM, because the pipeline container should not import the API.

SvelteKit. Small runtime, no virtual DOM, and Svelte 5 runes make server-driven state simple. The UI has exactly two pages; the goal is a table a scientist actually wants to filter, not a dashboard. PUBLIC_API_URL is read at runtime, so one image works behind the ingress (/api) and elsewhere.

GKE Autopilot + Cloud SQL, not self-managed. The lab is about the platform patterns (Workload Identity, GitOps, Kustomize overlays, private networking), not about running etcd. Cloud SQL has only a private IP; the API reaches it through a Cloud SQL Proxy sidecar, pipeline tasks directly in the VPC. Database URLs live in Secret Manager.

GitOps. CI builds and tests; it never runs kubectl apply. It edits image tags in the gcp overlay and ArgoCD reconciles. Rollback is git revert. Migrations run in an init container under a Postgres advisory lock, so replicas starting together migrate once.

MLflow registry as the model contract. The API loads whichever version the production alias points at and records that version on every prediction. The registered model is a pyfunc that owns its feature engineering (rarelens_ml.features ships inside it) and returns P(pathogenic), so serving only sends raw columns and cannot drift from training.

What is deliberately missing

Authentication, PHI handling, audit logs, clinical validation. This is a learning platform on public data. Adding Identity-Aware Proxy in front of the ingress is the first step if that ever changes.