Kemal Yaylali 11fb6b3d73 fix: overhaul the platform skeleton, add a serverless deployment track
An end-to-end audit found the repo could not build, test or run as shipped. This
fixes every finding, then adds a Cloud Run track so the demo costs about £1/month
idle instead of ~£150.

CI (red on its first run)
- api: setuptools could not build the package (flat layout with app/ and alembic/)
- web: missing @types/node; `vitest run` exited 1 with no test files
- pipeline: the stub run needed a gitignored VCF, and no process had a stub block
- ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI)

ML serving (scores were meaningless)
- the registered model now carries its own feature engineering and returns predict_proba,
  so serving sends raw columns and cannot drift from training
- resolve by registry alias (stages are deprecated in MLflow 3) and record the real
  version; re-scoring upserts instead of failing on the unique constraint
- ClinVar labels parsed from VEP's lowercase terms

Pipeline
- exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent
- job status reaches running/failed/succeeded, so the UI stops polling dead jobs
- DATABASE_URL travels in the environment or a Nextflow secret, never on a command line
- VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch

Deployment
- the API serves /api (matching the ingress); the web app reads its API URL at runtime
- migrations run in an init container under a Postgres advisory lock
- terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload
  Identity bindings, Secret Manager, deletion protection
- serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job
  for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and
  Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md.

Correctness and security
- 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider
  VEP text columns, enum dropped on downgrade, the sample's assembly actually used
- vcf_uri restricted to gs:// objects or files under the data root, blocking option injection
- CORS restricted to configured origins; `make down` no longer deletes volumes

Data
- docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar,
  gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice

Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform
validate and both kustomize overlays clean.
2026-09-12 07:21:11 +01:00

rarelens

A small, end-to-end variant interpretation platform for rare genetic disease research. Scientists upload a VCF, a Nextflow workflow annotates it with Ensembl VEP, a machine learning model scores each variant, and results are browsable in a web app.

This repository is a self-training lab. It exists so that one engineer can learn, in public, how a modern life-sciences platform is built end to end: full-stack application, scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a clinical tool and makes no diagnostic claims.

What is in the box

Layer Technology Directory
Pipeline Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch pipeline/
API FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic api/
Database PostgreSQL 16 docker-compose.yml
Frontend SvelteKit, TypeScript web/
ML LightGBM pathogenicity scorer, MLflow registry ml/
Orchestration Argo Workflows + Argo Events (pipeline), Pub/Sub (events) infra/argo-workflows/
Platform Kubernetes (Kustomize), ArgoCD (GitOps) infra/k8s/, infra/argocd/
Cloud GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry infra/terraform/
CI/CD GitHub Actions, Workload Identity Federation .github/workflows/

Quick start (local)

make up          # postgres + api + web + mlflow via docker-compose
make migrate     # alembic upgrade head
make data        # real public data: GIAB HG002 + ClinVar, chr22 (needs bcftools)
make test        # api, ml, loader and web tests (no Docker needed for the DB tests)

Then open http://localhost:5173.

The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the command to run instead. With Nextflow and Docker on the host, a VEP cache in pipeline/cache/vep and a VCF under data/ (see data/README.md):

make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz   # dry run: annotate without touching the database

To train and register a model (the API scores with models:/rarelens-pathogenicity@production):

cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \
  uv run python -m rarelens_ml.train --tsv ../pipeline/results/<sample>.vep.tsv --register

Local Kubernetes: make kind builds the images, loads them into a kind cluster and applies infra/k8s/overlays/local.

Deploying to GCP

Two tracks, same code. The serverless one is the default because it costs about £1/month idle; docs/cloud.md has the numbers.

Serverless (Cloud Run + Google Batch). The API and the UI scale to zero, and the Nextflow driver runs as a Cloud Run job only while a pipeline is running.

cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require'  # e.g. Neon's free tier
terraform apply -var project=<project id>          # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha>   # redeploy a new build

terraform output web_url is the URL to share; it serves the UI and proxies /api to the API, so there is one public address and no CORS. Upload the VEP cache to gs://<project>-rarelens-data/refs/vep before running a real annotation, and set -var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1 to score without running an MLflow server. Set a billing budget first — the demo has no authentication.

Kubernetes (GKE + Argo + ArgoCD). Off by default; turn it on to demonstrate the GitOps path, then destroy it.

terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id>   # once; commit the result
make gcp-secrets PROJECT=<project id>

Then install Argo Workflows, Argo Events and ArgoCD, and kubectl apply -f infra/argocd/app.yaml. Every green CI run on main bumps image tags in the gcp overlay and ArgoCD deploys them. make serverless-destroy PROJECT=<project id> tears everything down.

Data

The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002 benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources, licences, citations and how the model should be evaluated honestly are in docs/data.md.

Architecture

See docs/architecture.md for the diagram and the reasoning behind each choice, and docs/cloud.md for why this deploys to Google Cloud rather than AWS.

Status

Work in progress. Milestones, in order:

  1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
  2. SvelteKit UI: sample list, variant table with filters, job status
  3. Kubernetes manifests, kind, Argo Workflows trigger
  4. Terraform for GCP, ArgoCD GitOps deploy
  5. Pathogenicity model, MLflow registry, prediction endpoint

Licence

AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.

S
Description
End-to-end variant interpretation platform for rare genetic disease research. Public test data only; no clinical claims. AGPL-3.0.
Readme AGPL-3.0
12 MiB
Languages
Python 62.4%
TypeScript 10.4%
Svelte 8%
HCL 7.1%
Shell 3.4%
Other 8.6%