"Variants are unscored" was accurate: nothing was ever trained, so a quarter of every rank was dead weight and the UI leaked a connection error at the reader. - scripts/make-training-set.sh derives a training table from ClinVar directly. ClinVar already carries the molecular consequence, the gene and an allele frequency, which is the feature set serving sends, so this avoids running VEP over hundreds of thousands of variants. 2-star records only. - train.py now holds out whole genes (GroupShuffleSplit). docs/data.md had said to do this since the data pass; the code was still doing a random split, which is the leak Grimm 2015 describes. - evaluate() reports missense on its own. On the last run: AUROC 0.986 over 74,239 held-out variants, but 0.872 over the 13,553 missense ones, and the docs say plainly why even that is flattered — within missense the only live feature is allele frequency, and ClinVar's benign calls often use allele frequency as evidence (ACMG BA1/BS1), so the feature partly caused the label. - the 503 now names what is missing (model@alias via tracking URI) and leaves the exception in the server log instead of the UI. - make training-set / make train; the 58 MB table is gitignored. Verified end to end: model registered as v2, the simulated NF2 case scores 0.999 on the planted variant, and it now ranks 1.00 with all four components live. Tests: api 77, ml 22, loader 16, web 32; ruff, mypy, svelte-check clean.
149 lines
7.2 KiB
Markdown
149 lines
7.2 KiB
Markdown
# rarelens
|
|
|
|
A small, end-to-end variant interpretation platform for rare genetic disease research.
|
|
A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow
|
|
annotates the variants with Ensembl VEP, a model scores each one, and the app narrows
|
|
thousands of variants to a handful of candidates ranked against that phenotype — each
|
|
carrying the evidence for its rank, and each able to be shortlisted or dismissed with a
|
|
reason that ends up in a case report.
|
|
|
|
This repository is a **self-training lab**. It exists so that one engineer can learn, in
|
|
public, how a modern life-sciences platform is built end to end: full-stack application,
|
|
scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a
|
|
clinical tool and makes no diagnostic claims.
|
|
|
|
## What is in the box
|
|
|
|
| Layer | Technology | Directory |
|
|
|------------|--------------------------------------------------------|----------------------|
|
|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
|
|
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
|
|
| Database | PostgreSQL 16 | `docker-compose.yml` |
|
|
| Frontend | SvelteKit, TypeScript | `web/` |
|
|
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
|
|
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
|
|
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
|
|
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` |
|
|
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
|
|
|
|
## Quick start (local)
|
|
|
|
```bash
|
|
make up # postgres + api + web + mlflow via docker-compose
|
|
make migrate # alembic upgrade head
|
|
make hpo # HPO gene-to-phenotype annotations: what the ranking matches against
|
|
make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant
|
|
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
|
|
```
|
|
|
|
Then open http://localhost:5173, create a case pointing at `data/proband-simulated.vcf.gz`,
|
|
give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular
|
|
schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it.
|
|
The planted variant should come back ranked first.
|
|
|
|
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
|
|
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
|
|
and a VCF under `data/` (see [data/README.md](data/README.md)):
|
|
|
|
```bash
|
|
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
|
|
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
|
|
```
|
|
|
|
No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant
|
|
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
|
|
pipeline on a handful of variants:
|
|
|
|
```bash
|
|
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
|
|
```
|
|
|
|
To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and
|
|
Docker are) rather than in docker-compose:
|
|
|
|
```bash
|
|
docker compose up -d db
|
|
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
|
|
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
|
|
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
|
|
uv run --extra dev uvicorn app.main:app --port 8000
|
|
```
|
|
|
|
`PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost`
|
|
would be the container itself. `LOCAL_DATA_ROOT` is the directory a sample's `vcf_uri` must sit under.
|
|
|
|
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
|
|
|
|
```bash
|
|
make training-set # a ClinVar-derived training table, ~370k labelled variants
|
|
make train # fits, reports held-out metrics by gene split, moves the production alias
|
|
```
|
|
|
|
What those metrics do and do not mean is in [docs/data.md](docs/data.md); the headline AUROC
|
|
flatters a model whose strongest feature is the consequence class.
|
|
|
|
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
|
|
`infra/k8s/overlays/local`.
|
|
|
|
## Deploying to GCP
|
|
|
|
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
|
|
[docs/cloud.md](docs/cloud.md) has the numbers.
|
|
|
|
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
|
|
driver runs as a Cloud Run job only while a pipeline is running.
|
|
|
|
```bash
|
|
cd infra/terraform
|
|
terraform init -backend-config="bucket=<tfstate bucket>"
|
|
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
|
|
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
|
|
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
|
|
```
|
|
|
|
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
|
|
there is one public address and no CORS. Upload the VEP cache to
|
|
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
|
|
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
|
|
MLflow server. Set a billing budget first — the demo has no authentication.
|
|
|
|
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
|
|
then destroy it.
|
|
|
|
```bash
|
|
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
|
|
make gcp-configure PROJECT=<project id> # once; commit the result
|
|
make gcp-secrets PROJECT=<project id>
|
|
```
|
|
|
|
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
|
|
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
|
|
`make serverless-destroy PROJECT=<project id>` tears everything down.
|
|
|
|
## Data
|
|
|
|
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
|
|
benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources,
|
|
licences, citations and how the model should be evaluated honestly are in
|
|
[docs/data.md](docs/data.md).
|
|
|
|
## Architecture
|
|
|
|
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
|
|
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
|
|
than AWS.
|
|
|
|
## Status
|
|
|
|
Work in progress. Milestones, in order:
|
|
|
|
1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
|
|
2. SvelteKit UI: sample list, variant table with filters, job status
|
|
3. Kubernetes manifests, kind, Argo Workflows trigger
|
|
4. Terraform for GCP, ArgoCD GitOps deploy
|
|
5. Pathogenicity model, MLflow registry, prediction endpoint
|
|
|
|
## Licence
|
|
|
|
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.
|