fix: overhaul the platform skeleton, add a serverless deployment track

An end-to-end audit found the repo could not build, test or run as shipped. This
fixes every finding, then adds a Cloud Run track so the demo costs about £1/month
idle instead of ~£150.

CI (red on its first run)
- api: setuptools could not build the package (flat layout with app/ and alembic/)
- web: missing @types/node; `vitest run` exited 1 with no test files
- pipeline: the stub run needed a gitignored VCF, and no process had a stub block
- ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI)

ML serving (scores were meaningless)
- the registered model now carries its own feature engineering and returns predict_proba,
  so serving sends raw columns and cannot drift from training
- resolve by registry alias (stages are deprecated in MLflow 3) and record the real
  version; re-scoring upserts instead of failing on the unique constraint
- ClinVar labels parsed from VEP's lowercase terms

Pipeline
- exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent
- job status reaches running/failed/succeeded, so the UI stops polling dead jobs
- DATABASE_URL travels in the environment or a Nextflow secret, never on a command line
- VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch

Deployment
- the API serves /api (matching the ingress); the web app reads its API URL at runtime
- migrations run in an init container under a Postgres advisory lock
- terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload
  Identity bindings, Secret Manager, deletion protection
- serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job
  for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and
  Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md.

Correctness and security
- 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider
  VEP text columns, enum dropped on downgrade, the sample's assembly actually used
- vcf_uri restricted to gs:// objects or files under the data root, blocking option injection
- CORS restricted to configured origins; `make down` no longer deletes volumes

Data
- docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar,
  gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice

Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform
validate and both kustomize overlays clean.
This commit is contained in:
Kemal Yaylali
2026-09-12 07:21:11 +01:00
parent 5463f489a3
commit 11fb6b3d73
100 changed files with 3431 additions and 340 deletions
+70 -8
View File
@@ -13,31 +13,93 @@ clinical tool and makes no diagnostic claims.
| Layer | Technology | Directory |
|------------|--------------------------------------------------------|----------------------|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker | `pipeline/` |
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
| Database | PostgreSQL 16 | `docker-compose.yml` |
| Frontend | SvelteKit, TypeScript | `web/` |
| ML | LightGBM pathogenicity scorer, MLflow tracking | `ml/` |
| Orchestration | Argo Workflows (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Artifact Registry | `infra/terraform/` |
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` |
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
## Quick start (local)
```bash
make up # postgres + api + web via docker-compose
make up # postgres + api + web + mlflow via docker-compose
make migrate # alembic upgrade head
make pipeline # nextflow run pipeline/main.nf -profile docker --vcf data/example.vcf.gz
make kind # spin up a local kind cluster and apply infra/k8s/overlays/local
make data # real public data: GIAB HG002 + ClinVar, chr22 (needs bcftools)
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
```
Then open http://localhost:5173.
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
and a VCF under `data/` (see [data/README.md](data/README.md)):
```bash
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
```
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
```bash
cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \
uv run python -m rarelens_ml.train --tsv ../pipeline/results/<sample>.vep.tsv --register
```
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
`infra/k8s/overlays/local`.
## Deploying to GCP
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
[docs/cloud.md](docs/cloud.md) has the numbers.
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
driver runs as a Cloud Run job only while a pipeline is running.
```bash
cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
```
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
there is one public address and no CORS. Upload the VEP cache to
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
MLflow server. Set a billing budget first — the demo has no authentication.
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
then destroy it.
```bash
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id> # once; commit the result
make gcp-secrets PROJECT=<project id>
```
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
`make serverless-destroy PROJECT=<project id>` tears everything down.
## Data
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources,
licences, citations and how the model should be evaluated honestly are in
[docs/data.md](docs/data.md).
## Architecture
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
each choice.
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
than AWS.
## Status