Files
rarelens/Makefile
T
Kemal Yaylali 07a01715fd feat: redesign around phenotype-driven triage, not variant filtering
A table with filters made the user do the work. Rare disease triage is a different task:
which few variants could explain *this* patient's phenotype, and why. The app now answers
that, and lets a reviewer act on the answer.

Domain
- a case is a proband: a VCF plus the HPO terms observed in the patient (samples -> cases)
- HPO's gene-to-phenotype annotations are loaded as reference data (scripts/load-hpo.py)
- each candidate can be shortlisted or dismissed with a reason and a note

Ranking (app/services/triage.py, 21 tests)
- weighted sum of phenotype match, rarity, consequence severity and the model's score,
  with every component shown next to the candidate
- rarity and consequence filter; phenotype only ranks, because a real diagnosis can sit in
  a gene nobody has annotated yet and filtering on it would hide exactly that case
- ClinVar is deliberately not an input: it appears beside the result as independent
  confirmation, so nothing ranks highly merely because ClinVar already said pathogenic

UI
- the funnel is the headline: variants called -> rare -> coding candidates -> phenotype-matched
- ranked candidates with evidence chips, not a grid of everything; filters are demoted
- a variant panel showing the score breakdown, the matched HPO terms, the raw VEP record and
  links out to Ensembl/gnomAD/ClinVar, with the decision controls
- a printable case report: phenotype, funnel, shortlisted variants with reasons, provenance

API: /cases with phenotypes, /cases/{id}/candidates (funnel + ranked + weights),
/variants/{id}, /variants/{id}/decision, /cases/{id}/report, /phenotypes for the picker.
Scoring moved under the case and now answers 503 with the reason when no model registry is
reachable, instead of a 500.

Verified end to end on a simulated proband (scripts/make-demo-case.sh: real GIAB HG002
background + one real ClinVar 2-star pathogenic NF2 variant). 13 variants called -> 1 coding
candidate, and the planted variant ranks first at 0.80 on phenotype 1.00, rarity 1.00 and
consequence 1.00, with ClinVar agreeing afterwards.

Tests: api 75, ml 18, loader 16, web 27; ruff, mypy, svelte-check, terraform validate, both
kustomize overlays and the Nextflow stub run all clean.
2026-09-12 08:30:44 +01:00

84 lines
3.7 KiB
Makefile

.PHONY: up down clean migrate test lint data hpo demo-case loader pipeline annotate images kind serverless-deploy serverless-destroy gcp-configure gcp-secrets
VCF ?= data/example.vcf.gz
TAG ?= latest
# The loader container reaches docker-compose's Postgres through the host.
HOST_DB_URL ?= postgresql://rarelens:[email protected]:5432/rarelens
up:
docker compose up -d --build
down:
docker compose down
clean: ## also deletes the Postgres and MLflow volumes
docker compose down -v
migrate:
docker compose exec api alembic upgrade head
test:
cd api && uv run --extra dev pytest -q
cd ml && uv run --extra dev pytest -q
cd pipeline && uv run --no-project --with-requirements requirements.txt --with pytest --with pgserver pytest -q tests
cd web && npm test
lint:
cd api && uv run --extra dev ruff check . && uv run --extra dev mypy app
cd web && npm run check
data: ## download the public demo slice: GIAB HG002 + ClinVar, chr22 (see docs/data.md)
scripts/fetch-demo-data.sh
hpo: ## load HPO gene-to-phenotype annotations, which the ranking matches against
scripts/load-hpo.py
demo-case: ## build the simulated proband: GIAB background + one ClinVar pathogenic variant
scripts/make-demo-case.sh
loader:
docker build -t rarelens/loader:dev -f pipeline/loader.Dockerfile pipeline
pipeline: loader ## dry run: annotate $(VCF) without touching the database
cd pipeline && nextflow run main.nf -profile docker --vcf ../$(VCF)
annotate: loader ## make annotate JOB=<job id> [VCF=data/x.vcf.gz]
@test -n "$(JOB)" || (echo "usage: make annotate JOB=<job id> [VCF=...]"; exit 1)
cd pipeline && DATABASE_URL=$(HOST_DB_URL) nextflow run main.nf -profile docker \
--vcf ../$(VCF) --job_id $(JOB)
images:
docker build -t rarelens-api:dev api
docker build -t rarelens-web:dev web
kind: images
kind create cluster --name rarelens 2>/dev/null || true
kind load docker-image rarelens-api:dev rarelens-web:dev --name rarelens
kubectl apply -k infra/k8s/overlays/local
kubectl -n rarelens rollout status deploy/postgres deploy/api deploy/web
@echo "kubectl -n rarelens port-forward svc/web 8080:80 (UI)"
@echo "kubectl -n rarelens port-forward svc/api 8000:80 (API, used by the UI)"
serverless-deploy: ## deploy the Cloud Run track: make serverless-deploy PROJECT=<id> [TAG=<sha>]
@test -n "$(PROJECT)" || (echo "usage: make serverless-deploy PROJECT=<gcp project id> [TAG=<image tag>]"; exit 1)
@test -n "$$TF_VAR_database_url" || echo "note: TF_VAR_database_url is unset; add -var deploy_cloud_sql=true or export a Postgres URL"
cd infra/terraform && terraform apply -var project=$(PROJECT) -var image_tag=$(TAG)
serverless-destroy: ## tear it all down
cd infra/terraform && terraform destroy -var project=$(PROJECT) -var deletion_protection=false
gcp-configure: ## one-time: write your GCP project id into the gcp overlay and Argo manifests
@test -n "$(PROJECT)" || (echo "usage: make gcp-configure PROJECT=<gcp project id>"; exit 1)
grep -rl __GCP_PROJECT__ infra/k8s/overlays/gcp infra/argo-workflows \
| xargs sed -i.bak "s/__GCP_PROJECT__/$(PROJECT)/g"
find infra -name '*.bak' -delete
gcp-secrets: ## after terraform apply: copy DB URLs from Secret Manager into k8s secrets
@test -n "$(PROJECT)" || (echo "usage: make gcp-secrets PROJECT=<gcp project id>"; exit 1)
kubectl -n rarelens create secret generic api-secrets --dry-run=client -o yaml \
--from-literal=DATABASE_URL="$$(gcloud secrets versions access latest --project $(PROJECT) --secret rarelens-api-database-url)" \
| kubectl apply -f -
kubectl -n rarelens create secret generic pipeline-secrets --dry-run=client -o yaml \
--from-literal=DATABASE_URL="$$(gcloud secrets versions access latest --project $(PROJECT) --secret DATABASE_URL)" \
| kubectl apply -f -