From clicking through the redesigned UI: - scoring a case without a model registry painted a red failure across a case that had in fact analysed fine. It is now a quiet note saying the model term contributes 0, because scoring is an optional fourth of the rank, not the analysis. - the MLflow default moves to port 5001. On macOS, AirPlay Receiver owns 5000, which is why the registry answered "403" rather than refusing the connection; docker-compose publishes 5001 to match. - a funnel step that kept nothing drew a visible bar. Zero now draws zero. - "1 candidates". - the funnel's fixed grid columns forced a horizontal scrollbar on the report. The report also lists the top undecided candidates now: the first thing anyone opens has no decisions in it, and "Shortlisted (0)" alone said nothing about what the tool found. Tests: api 77, web 32; ruff, mypy, svelte-check clean.
rarelens
A small, end-to-end variant interpretation platform for rare genetic disease research. A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow annotates the variants with Ensembl VEP, a model scores each one, and the app narrows thousands of variants to a handful of candidates ranked against that phenotype — each carrying the evidence for its rank, and each able to be shortlisted or dismissed with a reason that ends up in a case report.
This repository is a self-training lab. It exists so that one engineer can learn, in public, how a modern life-sciences platform is built end to end: full-stack application, scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a clinical tool and makes no diagnostic claims.
What is in the box
| Layer | Technology | Directory |
|---|---|---|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | pipeline/ |
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | api/ |
| Database | PostgreSQL 16 | docker-compose.yml |
| Frontend | SvelteKit, TypeScript | web/ |
| ML | LightGBM pathogenicity scorer, MLflow registry | ml/ |
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | infra/argo-workflows/ |
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | infra/k8s/, infra/argocd/ |
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | infra/terraform/ |
| CI/CD | GitHub Actions, Workload Identity Federation | .github/workflows/ |
Quick start (local)
make up # postgres + api + web + mlflow via docker-compose
make migrate # alembic upgrade head
make hpo # HPO gene-to-phenotype annotations: what the ranking matches against
make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
Then open http://localhost:5173, create a case pointing at data/proband-simulated.vcf.gz,
give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular
schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it.
The planted variant should come back ranked first.
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
command to run instead. With Nextflow and Docker on the host, a VEP cache in pipeline/cache/vep
and a VCF under data/ (see data/README.md):
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
No cache? VEP_DATABASE=true queries Ensembl's public database instead. It is slow per variant
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
pipeline on a handful of variants:
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and Docker are) rather than in docker-compose:
docker compose up -d db
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
uv run --extra dev uvicorn app.main:app --port 8000
PIPELINE_DATABASE_URL is what the loader container gets: inside it, the API's own localhost
would be the container itself. LOCAL_DATA_ROOT is the directory a sample's vcf_uri must sit under.
To train and register a model (the API scores with models:/rarelens-pathogenicity@production):
cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \
uv run python -m rarelens_ml.train --tsv ../pipeline/results/<sample>.vep.tsv --register
Local Kubernetes: make kind builds the images, loads them into a kind cluster and applies
infra/k8s/overlays/local.
Deploying to GCP
Two tracks, same code. The serverless one is the default because it costs about £1/month idle; docs/cloud.md has the numbers.
Serverless (Cloud Run + Google Batch). The API and the UI scale to zero, and the Nextflow driver runs as a Cloud Run job only while a pipeline is running.
cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
terraform output web_url is the URL to share; it serves the UI and proxies /api to the API, so
there is one public address and no CORS. Upload the VEP cache to
gs://<project>-rarelens-data/refs/vep before running a real annotation, and set
-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1 to score without running an
MLflow server. Set a billing budget first — the demo has no authentication.
Kubernetes (GKE + Argo + ArgoCD). Off by default; turn it on to demonstrate the GitOps path, then destroy it.
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id> # once; commit the result
make gcp-secrets PROJECT=<project id>
Then install Argo Workflows, Argo Events and ArgoCD, and kubectl apply -f infra/argocd/app.yaml.
Every green CI run on main bumps image tags in the gcp overlay and ArgoCD deploys them.
make serverless-destroy PROJECT=<project id> tears everything down.
Data
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002 benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources, licences, citations and how the model should be evaluated honestly are in docs/data.md.
Architecture
See docs/architecture.md for the diagram and the reasoning behind each choice, and docs/cloud.md for why this deploys to Google Cloud rather than AWS.
Status
Work in progress. Milestones, in order:
- Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
- SvelteKit UI: sample list, variant table with filters, job status
- Kubernetes manifests, kind, Argo Workflows trigger
- Terraform for GCP, ArgoCD GitOps deploy
- Pathogenicity model, MLflow registry, prediction endpoint
Licence
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.