Files
rarelens/README.md
Kemal Yaylali e76ae847a1 fix(science): stop scoring evidence that was never looked up
A review of the ranking's arithmetic found four things wrong, all of which
made the score look better informed than it was. Measurements below are from
this repo, not estimates.

**Components now abstain instead of inventing a number.** A run without a VEP
cache returns no allele frequencies, and rarity_score(None) read that as
"absent from gnomAD, therefore maximally rare" and awarded every variant a
free 0.25. jobs.has_frequencies / has_effect_scores record what the run
actually produced, absent components are dropped from the weighted mean, and
the remaining weights are renormalised so the score keeps its meaning. The UI
shows "not looked up" rather than a bar, and the funnel stops calling a step
"rare" when nothing was filtered.

**Allele frequency is no longer a model feature.** It dominated: the same
missense variant scored 0.887 at AF 0 and 0.0003 at AF 0.01. That double-
counted, because the ranking already scores frequency explicitly, putting
~45% of every rank on one measurement; and it was circular, because ACMG
assigns ClinVar's benign labels using frequency (BA1/BS1). Retraining without
it moves missense AUROC from 0.872 to 0.500 — exactly random. The old figure
was allele frequency, not variant-effect knowledge. The model therefore
abstains unless CADD or AlphaMissense is present, since otherwise it only
restates the consequence class.

**Phenotype matching is weighted by information content** and HPO annotations
are propagated up the ontology. Counting terms alike let "global
developmental delay" (IC 0.93) count as much as "dilated left subclavian
artery" (IC 7.88).

**A real bug in the propagation, found by checking it.** The ancestor walk
read a pre-order DFS backwards, which on a DAG lets a term resolve before one
of its parents and inherit that parent alone instead of its lineage. It
dropped 399 terms out of the phenotype branch, Camptodactyly and Chiari
malformation among them. Now a true post-order, tested against a reference
transitive closure.

The ontology arithmetic moved to rarelens_ml.hpo so it is covered by tests,
and rarelens_ml.benchmark measures the whole thing: across 10,178 published
cases the causal gene ranks first 45.9-81.0% of the time against 5,269 genes,
versus 0.02% for chance. docs/data.md reports that with its contamination
(HPO's annotations come from these same case reports), and includes the
measurement showing information-content weighting earns its place while
propagation does not - kept anyway, for a reason the docs argue rather than
assume.
2026-09-12 11:32:46 +01:00

9.5 KiB

rarelens

A small, end-to-end variant interpretation platform for rare genetic disease research. A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow annotates the variants with Ensembl VEP, a model scores each one, and the app narrows thousands of variants to a handful of candidates ranked against that phenotype — each carrying the evidence for its rank, and each able to be shortlisted or dismissed with a reason that ends up in a case report.

This repository is a self-training lab. It exists so that one engineer can learn, in public, how a modern life-sciences platform is built end to end: full-stack application, scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a clinical tool and makes no diagnostic claims.

What is in the box

Layer Technology Directory
Pipeline Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch pipeline/
API FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic api/
Database PostgreSQL 16 docker-compose.yml
Frontend SvelteKit, TypeScript web/
ML LightGBM pathogenicity scorer, MLflow registry ml/
Orchestration Cloud Run job (default), Argo Workflows + Argo Events, Pub/Sub infra/terraform/, infra/argo-workflows/
Platform Kubernetes (Kustomize), ArgoCD (GitOps) infra/k8s/, infra/argocd/
Cloud GCP: Cloud Run, Google Batch, GCS, Secret Manager, Artifact Registry; GKE Autopilot and Cloud SQL behind flags infra/terraform/
CI/CD GitHub Actions, Workload Identity Federation .github/workflows/

Quick start (local)

make up          # postgres + api + web + mlflow via docker-compose
make migrate     # alembic upgrade head
make hpo         # HPO gene-to-phenotype annotations: what the ranking matches against
make demo-case   # a simulated proband: GIAB background + one ClinVar pathogenic variant
make published-case  # a real published patient: their reported phenotype and causal variant
make test        # api, ml, loader and web tests (no Docker needed for the DB tests)

Then open http://localhost:5173, create a case pointing at data/proband-simulated.vcf.gz, give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it. The planted variant should come back ranked first.

make published-case is the same idea with nothing invented. It builds a case from a GA4GH phenopacket curated from a peer-reviewed case report — by default the TGFBR2 proband from Loeys et al., Nat Genet 2005, 10.1038/ng1511, the paper that first described Loeys-Dietz syndrome. The patient's 30 reported HPO terms and their causal variant come straight from the publication; the background variants come from GIAB HG002, because the rest of that patient's genome is not public. It writes the phenotype list alongside the VCF, so the case can be created exactly as reported. See docs/data.md for the provenance and for what this case does and does not demonstrate.

The docker-compose API has no Nextflow, so "Analyse case" marks the job failed with the command to run instead. With Nextflow and Docker on the host, a VEP cache in pipeline/cache/vep and a VCF under data/ (see data/README.md):

make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
make pipeline VCF=data/example.vcf.gz   # dry run: annotate without touching the database

No cache? VEP_DATABASE=true queries Ensembl's public database instead. It is slow per variant and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the pipeline on a handful of variants:

VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf

To make the UI's "Analyse case" button work, run the API on the host (where Nextflow and Docker are) rather than in docker-compose:

docker compose up -d db
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
  PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
  LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
  uv run --extra dev uvicorn app.main:app --port 8000

PIPELINE_DATABASE_URL is what the loader container gets: inside it, the API's own localhost would be the container itself. LOCAL_DATA_ROOT is the directory a case's vcf_uri must sit under.

To train and register a model (the API scores with models:/rarelens-pathogenicity@production):

make training-set   # a ClinVar-derived training table, ~370k labelled variants
make train          # fits, reports held-out metrics by gene split, moves the production alias

What those metrics do and do not mean is in docs/data.md. The short version: with allele frequency removed as a feature, the model scores AUROC 0.500 — exactly random — on missense variants, because nothing is left but the consequence class the ranking already uses. It therefore abstains from the ranking unless CADD or AlphaMissense scores are available. The frequency feature is what made the old 0.872 look respectable, and ACMG assigns ClinVar's benign labels using frequency, so the feature had partly caused the label.

make benchmark      # rank every published case in Phenopacket Store by phenotype alone

Across 10,178 published cases the causal gene is ranked first 45.9-81.0% of the time (the range is ties; random would be 0.02%). That benchmark is contaminated — HPO's gene annotations come from the same case reports — so read it as an upper bound. docs/data.md has the full table, including the measurement that says information-content weighting earns its place and ontology propagation does not.

Local Kubernetes: make kind builds the images, loads them into a kind cluster and applies infra/k8s/overlays/local.

Deploying to GCP

Two tracks, same code. The serverless one is the default because it costs about £1/month idle; docs/cloud.md has the numbers.

Serverless (Cloud Run + Google Batch). The API and the UI scale to zero, and the Nextflow driver runs as a Cloud Run job only while a pipeline is running.

cd infra/terraform
terraform init -backend-config="bucket=<tfstate bucket>"
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require'  # e.g. Neon's free tier
terraform apply -var project=<project id>          # add -var deploy_cloud_sql=true to use Cloud SQL instead
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha>   # redeploy a new build

terraform output web_url is the URL to share; it serves the UI and proxies /api to the API, so there is one public address and no CORS. Upload the VEP cache to gs://<project>-rarelens-data/refs/vep before running a real annotation, and set -var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1 to score without running an MLflow server. Set a billing budget first — the demo has no authentication.

Kubernetes (GKE + Argo + ArgoCD). Off by default; turn it on to demonstrate the GitOps path, then destroy it.

terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
make gcp-configure PROJECT=<project id>   # once; commit the result
make gcp-secrets PROJECT=<project id>

Then install Argo Workflows, Argo Events and ArgoCD, and kubectl apply -f infra/argocd/app.yaml. Every green CI run on main bumps image tags in the gcp overlay and ArgoCD deploys them. make serverless-destroy PROJECT=<project id> tears everything down.

Data

The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002 benchmark genome as the background sample, ClinVar for labels, gnomAD for allele frequencies, the Human Phenotype Ontology's gene-to-phenotype annotations as what the ranking matches against, and GA4GH phenopackets curated from case reports for the published case. Sources, licences, citations and how the model should be evaluated honestly are in docs/data.md.

Architecture

See docs/architecture.md for the diagram and the reasoning behind each choice, and docs/cloud.md for why this deploys to Google Cloud rather than AWS.

Status

A self-training lab, built in the open. Working end to end:

  1. Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
  2. SvelteKit UI: cases, a phenotype-ranked candidate list showing the evidence behind each rank, shortlist/dismiss decisions, and a case report
  3. Kubernetes manifests, kind, Argo Workflows trigger
  4. Terraform for GCP: serverless Cloud Run + Batch by default, GKE and ArgoCD behind a flag
  5. Pathogenicity model trained on ClinVar, MLflow registry, scoring endpoint
  6. A demo case built from a published patient, with the citations behind it

Known gaps: allele frequencies need the 25 GB VEP cache, because VEP's database mode returns none, so the rarity term does no work without it; and the deployed demo has no authentication.

Licence

AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.