Files
rarelens/README.md
T
Kemal Yaylali 07a01715fd feat: redesign around phenotype-driven triage, not variant filtering
A table with filters made the user do the work. Rare disease triage is a different task:
which few variants could explain *this* patient's phenotype, and why. The app now answers
that, and lets a reviewer act on the answer.

Domain
- a case is a proband: a VCF plus the HPO terms observed in the patient (samples -> cases)
- HPO's gene-to-phenotype annotations are loaded as reference data (scripts/load-hpo.py)
- each candidate can be shortlisted or dismissed with a reason and a note

Ranking (app/services/triage.py, 21 tests)
- weighted sum of phenotype match, rarity, consequence severity and the model's score,
  with every component shown next to the candidate
- rarity and consequence filter; phenotype only ranks, because a real diagnosis can sit in
  a gene nobody has annotated yet and filtering on it would hide exactly that case
- ClinVar is deliberately not an input: it appears beside the result as independent
  confirmation, so nothing ranks highly merely because ClinVar already said pathogenic

UI
- the funnel is the headline: variants called -> rare -> coding candidates -> phenotype-matched
- ranked candidates with evidence chips, not a grid of everything; filters are demoted
- a variant panel showing the score breakdown, the matched HPO terms, the raw VEP record and
  links out to Ensembl/gnomAD/ClinVar, with the decision controls
- a printable case report: phenotype, funnel, shortlisted variants with reasons, provenance

API: /cases with phenotypes, /cases/{id}/candidates (funnel + ranked + weights),
/variants/{id}, /variants/{id}/decision, /cases/{id}/report, /phenotypes for the picker.
Scoring moved under the case and now answers 503 with the reason when no model registry is
reachable, instead of a 500.

Verified end to end on a simulated proband (scripts/make-demo-case.sh: real GIAB HG002
background + one real ClinVar 2-star pathogenic NF2 variant). 13 variants called -> 1 coding
candidate, and the planted variant ranks first at 0.80 on phenotype 1.00, rarity 1.00 and
consequence 1.00, with ClinVar agreeing afterwards.

Tests: api 75, ml 18, loader 16, web 27; ruff, mypy, svelte-check, terraform validate, both
kustomize overlays and the Nextflow stub run all clean.
2026-09-12 08:30:44 +01:00

146 lines
7.0 KiB
Markdown

# rarelens
A small, end-to-end variant interpretation platform for rare genetic disease research.
A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow
annotates the variants with Ensembl VEP, a model scores each one, and the app narrows
thousands of variants to a handful of candidates ranked against that phenotype — each
carrying the evidence for its rank, and each able to be shortlisted or dismissed with a
reason that ends up in a case report.
This repository is a **self-training lab**. It exists so that one engineer can learn, in
public, how a modern life-sciences platform is built end to end: full-stack application,
scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a
clinical tool and makes no diagnostic claims.
## What is in the box
| Layer | Technology | Directory |
|------------|--------------------------------------------------------|----------------------|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
| Database | PostgreSQL 16 | `docker-compose.yml` |
| Frontend | SvelteKit, TypeScript | `web/` |
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` |
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
## Quick start (local)
```bash
make up # postgres + api + web + mlflow via docker-compose
make migrate # alembic upgrade head
make hpo # HPO gene-to-phenotype annotations: what the ranking matches against
make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
```
Then open http://localhost:5173, create a case pointing at `data/proband-simulated.vcf.gz`,
give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular
schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it.
The planted variant should come back ranked first.
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
and a VCF under `data/` (see [data/README.md](data/README.md)):
```bash
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
```
No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
pipeline on a handful of variants:
```bash
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
```
To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and
Docker are) rather than in docker-compose:
```bash
docker compose up -d db
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
uv run --extra dev uvicorn app.main:app --port 8000
```
`PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost`
would be the container itself. `LOCAL_DATA_ROOT` is the directory a sample's `vcf_uri` must sit under.
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
```bash
cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \
uv run python -m rarelens_ml.train --tsv ../pipeline/results/<sample>.vep.tsv --register
```
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
`infra/k8s/overlays/local`.
## Deploying to GCP
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
[docs/cloud.md](docs/cloud.md) has the numbers.
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
driver runs as a Cloud Run job only while a pipeline is running.
```bash
cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
```
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
there is one public address and no CORS. Upload the VEP cache to
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
MLflow server. Set a billing budget first — the demo has no authentication.
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
then destroy it.
```bash
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id> # once; commit the result
make gcp-secrets PROJECT=<project id>
```
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
`make serverless-destroy PROJECT=<project id>` tears everything down.
## Data
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources,
licences, citations and how the model should be evaluated honestly are in
[docs/data.md](docs/data.md).
## Architecture
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
than AWS.
## Status
Work in progress. Milestones, in order:
1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
2. SvelteKit UI: sample list, variant table with filters, job status
3. Kubernetes manifests, kind, Argo Workflows trigger
4. Terraform for GCP, ArgoCD GitOps deploy
5. Pathogenicity model, MLflow registry, prediction endpoint
## Licence
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.