Files
rarelens/README.md
T
Kemal Yaylali 5588c9391d feat(data): build a case from a real published patient
`make published-case` reads a GA4GH phenopacket from Monarch's Phenopacket
Store and takes two things from it verbatim: the HPO terms the authors
reported and the variant they called causal. The default is the TGFBR2
proband from Loeys et al., Nat Genet 2005 (doi:10.1038/ng1511), the paper
that first defined Loeys-Dietz syndrome -- 30 reported terms and
NM_003242.6:c.1069G>T p.(Gly357Trp).

The rest of that patient's genome is not public, so background variants come
from GIAB HG002 around the locus. They are drawn from coding exons where
possible, via Ensembl's REST API: of ~4,000 HG002 variants in the window only
9 are coding, so a random sample is entirely intronic, the consequence filter
discards all of it, and the causal variant is left as the only candidate --
a funnel that proves nothing.

The real run ranks TGFBR2 first at 0.897 against an OSBPL10 missense at
0.547. Both are rare missense variants the model scores identically (0.887);
only the phenotype separates them, which is the argument for phenotype-driven
triage in one table.

Documented with three caveats rather than left implicit: the phenotype match
is partly circular because HPO's gene annotations are themselves curated from
published cases; rarity contributes nothing without the VEP cache
(--af_gnomade is rejected with --database, and plain --af returns nothing
even for rs429358 at ~15% global frequency); and one healthy genome is not a
diagnostic exome.
2026-09-12 10:27:08 +01:00

159 lines
8.0 KiB
Markdown

# rarelens
A small, end-to-end variant interpretation platform for rare genetic disease research.
A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow
annotates the variants with Ensembl VEP, a model scores each one, and the app narrows
thousands of variants to a handful of candidates ranked against that phenotype — each
carrying the evidence for its rank, and each able to be shortlisted or dismissed with a
reason that ends up in a case report.
This repository is a **self-training lab**. It exists so that one engineer can learn, in
public, how a modern life-sciences platform is built end to end: full-stack application,
scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a
clinical tool and makes no diagnostic claims.
## What is in the box
| Layer | Technology | Directory |
|------------|--------------------------------------------------------|----------------------|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
| Database | PostgreSQL 16 | `docker-compose.yml` |
| Frontend | SvelteKit, TypeScript | `web/` |
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` |
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
## Quick start (local)
```bash
make up # postgres + api + web + mlflow via docker-compose
make migrate # alembic upgrade head
make hpo # HPO gene-to-phenotype annotations: what the ranking matches against
make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant
make published-case # a real published patient: their reported phenotype and causal variant
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
```
Then open http://localhost:5173, create a case pointing at `data/proband-simulated.vcf.gz`,
give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular
schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it.
The planted variant should come back ranked first.
`make published-case` is the same idea with nothing invented. It builds a case from a GA4GH
phenopacket curated from a peer-reviewed case report — by default the *TGFBR2* proband from
Loeys et al., *Nat Genet* 2005, [10.1038/ng1511](https://doi.org/10.1038/ng1511), the paper that
first described Loeys-Dietz syndrome. The patient's 30 reported HPO terms and their causal
variant come straight from the publication; the background variants come from GIAB HG002, because
the rest of that patient's genome is not public. It writes the phenotype list alongside the VCF,
so the case can be created exactly as reported. See [docs/data.md](docs/data.md) for the
provenance and for what this case does and does not demonstrate.
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
and a VCF under `data/` (see [data/README.md](data/README.md)):
```bash
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
```
No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
pipeline on a handful of variants:
```bash
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
```
To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and
Docker are) rather than in docker-compose:
```bash
docker compose up -d db
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
uv run --extra dev uvicorn app.main:app --port 8000
```
`PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost`
would be the container itself. `LOCAL_DATA_ROOT` is the directory a sample's `vcf_uri` must sit under.
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
```bash
make training-set # a ClinVar-derived training table, ~370k labelled variants
make train # fits, reports held-out metrics by gene split, moves the production alias
```
What those metrics do and do not mean is in [docs/data.md](docs/data.md); the headline AUROC
flatters a model whose strongest feature is the consequence class.
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
`infra/k8s/overlays/local`.
## Deploying to GCP
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
[docs/cloud.md](docs/cloud.md) has the numbers.
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
driver runs as a Cloud Run job only while a pipeline is running.
```bash
cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
```
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
there is one public address and no CORS. Upload the VEP cache to
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
MLflow server. Set a billing budget first — the demo has no authentication.
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
then destroy it.
```bash
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id> # once; commit the result
make gcp-secrets PROJECT=<project id>
```
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
`make serverless-destroy PROJECT=<project id>` tears everything down.
## Data
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources,
licences, citations and how the model should be evaluated honestly are in
[docs/data.md](docs/data.md).
## Architecture
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
than AWS.
## Status
Work in progress. Milestones, in order:
1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
2. SvelteKit UI: sample list, variant table with filters, job status
3. Kubernetes manifests, kind, Argo Workflows trigger
4. Terraform for GCP, ArgoCD GitOps deploy
5. Pathogenicity model, MLflow registry, prediction endpoint
## Licence
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.