Three services -- Postgres, API, UI -- with the API on Railway's private network only, so the UI's /api proxy is the single public entry point and there is no CORS. The pipeline cannot run there. Nextflow shells out to `docker run` for VEP and bcftools, and Railway gives you a container, not a Docker daemon. Rather than leave a button that always fails, cases are annotated locally and copied up by scripts/seed-remote.sh, and PUBLIC_PIPELINE_ENABLED=false hides the analyse/score actions and the create-case form. DATABASE_IDLE_CONNECTIONS=false is what makes idling work. Railway decides a service is idle from its *outbound* traffic and sleeps it after ~5-10 minutes; a pooled database connection is outbound traffic, so SQLAlchemy's default pool would have kept the API awake and billable for ever. Setting it false switches to NullPool, which costs a connection per request -- nothing at demo traffic, the wrong trade under real load, hence the flag rather than a rewrite. BASIC_AUTH_USER / BASIC_AUTH_PASSWORD put one shared credential in front of the site. Nothing deployed is patient data, so this stops the URL being wandered into rather than protecting anyone's privacy; unset, the site is open, which is what local development wants. Compared in constant time, and both halves of the credential are checked even when the first fails.
216 lines
12 KiB
Markdown
216 lines
12 KiB
Markdown
# rarelens
|
|
|
|
A small, end-to-end variant interpretation platform for rare genetic disease research.
|
|
A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow
|
|
annotates the variants with Ensembl VEP, a model scores each one, and the app narrows
|
|
thousands of variants to a handful of candidates ranked against that phenotype — each
|
|
carrying the evidence for its rank, and each able to be shortlisted or dismissed with a
|
|
reason that ends up in a case report.
|
|
|
|
This repository is a **self-training lab**. It exists so that one engineer can learn, in
|
|
public, how a modern life-sciences platform is built end to end: full-stack application,
|
|
scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a
|
|
clinical tool and makes no diagnostic claims.
|
|
|
|
## What is in the box
|
|
|
|
| Layer | Technology | Directory |
|
|
|------------|--------------------------------------------------------|----------------------|
|
|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
|
|
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
|
|
| Database | PostgreSQL 16 | `docker-compose.yml` |
|
|
| Frontend | SvelteKit, TypeScript | `web/` |
|
|
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
|
|
| Orchestration | Cloud Run job (default), Argo Workflows + Argo Events, Pub/Sub | `infra/terraform/`, `infra/argo-workflows/` |
|
|
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
|
|
| Cloud | GCP: Cloud Run, Google Batch, GCS, Secret Manager, Artifact Registry; GKE Autopilot and Cloud SQL behind flags | `infra/terraform/` |
|
|
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
|
|
|
|
## Quick start (local)
|
|
|
|
```bash
|
|
make up # postgres + api + web + mlflow via docker-compose
|
|
make migrate # alembic upgrade head
|
|
make hpo # HPO gene-to-phenotype annotations: what the ranking matches against
|
|
make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant
|
|
make published-case # a real published patient: their reported phenotype and causal variant
|
|
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
|
|
```
|
|
|
|
Then open http://localhost:5173, create a case pointing at `data/proband-simulated.vcf.gz`,
|
|
give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular
|
|
schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it.
|
|
The planted variant should come back ranked first.
|
|
|
|
`make published-case` is the same idea with nothing invented. It builds a case from a GA4GH
|
|
phenopacket curated from a peer-reviewed case report — by default the *TGFBR2* proband from
|
|
Loeys et al., *Nat Genet* 2005, [10.1038/ng1511](https://doi.org/10.1038/ng1511), the paper that
|
|
first described Loeys-Dietz syndrome. The patient's 30 reported HPO terms and their causal
|
|
variant come straight from the publication; the background variants come from GIAB HG002, because
|
|
the rest of that patient's genome is not public. It writes the phenotype list alongside the VCF,
|
|
so the case can be created exactly as reported. See [docs/data.md](docs/data.md) for the
|
|
provenance and for what this case does and does not demonstrate.
|
|
|
|
The docker-compose API has no Nextflow, so "Analyse case" marks the job failed with the
|
|
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
|
|
and a VCF under `data/` (see [data/README.md](data/README.md)):
|
|
|
|
```bash
|
|
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
|
|
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
|
|
```
|
|
|
|
No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant
|
|
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
|
|
pipeline on a handful of variants:
|
|
|
|
```bash
|
|
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
|
|
```
|
|
|
|
To make the UI's "Analyse case" button work, run the API on the host (where Nextflow and
|
|
Docker are) rather than in docker-compose:
|
|
|
|
```bash
|
|
docker compose up -d db
|
|
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
|
|
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
|
|
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
|
|
uv run --extra dev uvicorn app.main:app --port 8000
|
|
```
|
|
|
|
`PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost`
|
|
would be the container itself. `LOCAL_DATA_ROOT` is the directory a case's `vcf_uri` must sit under.
|
|
|
|
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
|
|
|
|
```bash
|
|
make training-set # a ClinVar-derived training table, ~370k labelled variants
|
|
make train # fits, reports held-out metrics by gene split, moves the production alias
|
|
```
|
|
|
|
What those metrics do and do not mean is in [docs/data.md](docs/data.md). The short version:
|
|
with allele frequency removed as a feature, the model scores AUROC **0.500 — exactly random — on
|
|
missense variants**, because nothing is left but the consequence class the ranking already uses.
|
|
It therefore abstains from the ranking unless CADD or AlphaMissense scores are available. The
|
|
frequency feature is what made the old 0.872 look respectable, and ACMG assigns ClinVar's benign
|
|
labels using frequency, so the feature had partly caused the label.
|
|
|
|
```bash
|
|
make benchmark # rank every published case in Phenopacket Store by phenotype alone
|
|
```
|
|
|
|
Across 10,178 published cases the causal gene is ranked first 45.9-81.0% of the time (the range is
|
|
ties; random would be 0.02%). That benchmark is contaminated — HPO's gene annotations come from the
|
|
same case reports — so read it as an upper bound. [docs/data.md](docs/data.md) has the full table,
|
|
including the measurement that says information-content weighting earns its place and ontology
|
|
propagation does not.
|
|
|
|
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
|
|
`infra/k8s/overlays/local`.
|
|
|
|
## Deploying to Railway
|
|
|
|
The quickest way to put it in front of people. Three services — Postgres, the API, the UI — with
|
|
the API reachable only over Railway's private network, so the UI's `/api` proxy is the single
|
|
public entry point and there is no CORS.
|
|
|
|
```bash
|
|
railway link --project <id> --environment production --service api
|
|
cd api && railway up --service api # the repo is on Gitea, so deploy from the working copy
|
|
cd ../web && railway up --service web
|
|
```
|
|
|
|
Set on the API: `DATABASE_URL` (pointing at `postgres.railway.internal`) and
|
|
`DATABASE_IDLE_CONNECTIONS=false`. On the UI: `API_INTERNAL_URL=http://api.railway.internal:8000`,
|
|
`PUBLIC_PIPELINE_ENABLED=false`, and `BASIC_AUTH_USER` / `BASIC_AUTH_PASSWORD`.
|
|
|
|
Three things are worth knowing before copying this:
|
|
|
|
- **The pipeline cannot run there.** Nextflow shells out to `docker run` for VEP and bcftools, and
|
|
Railway gives you a container, not a Docker daemon. So the cases are annotated here and copied
|
|
up with `scripts/seed-remote.sh`, and `PUBLIC_PIPELINE_ENABLED=false` hides the buttons that
|
|
would otherwise be left to fail. Visitors explore real analysed cases; they do not run VEP.
|
|
- **`DATABASE_IDLE_CONNECTIONS=false` is what makes sleeping work.** Railway decides a service is
|
|
idle from its *outbound* traffic, and a pooled database connection is outbound traffic, so the
|
|
default pool would keep the API awake and billable for ever. It uses `NullPool` instead, which
|
|
costs a connection per request and is the wrong trade under real load.
|
|
- **Serverless must be enabled per service and only takes effect on the next deploy.** Leave it off
|
|
for Postgres, which holds the volume.
|
|
|
|
`BASIC_AUTH_USER`/`BASIC_AUTH_PASSWORD` put one shared credential in front of the whole site
|
|
(`web/src/hooks.server.ts`); unset, the site is open, which is what local development wants.
|
|
Nothing here is patient data, so this stops the URL being wandered into rather than protecting
|
|
anyone's privacy.
|
|
|
|
Cost: Railway's Hobby plan is $5/month flat including $5 of usage, so it costs that whether or not
|
|
anyone visits — more at rest than the GCP track below, and much less to operate.
|
|
|
|
## Deploying to GCP
|
|
|
|
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
|
|
[docs/cloud.md](docs/cloud.md) has the numbers.
|
|
|
|
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
|
|
driver runs as a Cloud Run job only while a pipeline is running.
|
|
|
|
```bash
|
|
cd infra/terraform
|
|
terraform init -backend-config="bucket=<tfstate bucket>"
|
|
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
|
|
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
|
|
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
|
|
```
|
|
|
|
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
|
|
there is one public address and no CORS. Upload the VEP cache to
|
|
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
|
|
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
|
|
MLflow server. Set a billing budget first — the demo has no authentication.
|
|
|
|
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
|
|
then destroy it.
|
|
|
|
```bash
|
|
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
|
|
make gcp-configure PROJECT=<project id> # once; commit the result
|
|
make gcp-secrets PROJECT=<project id>
|
|
```
|
|
|
|
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
|
|
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
|
|
`make serverless-destroy PROJECT=<project id>` tears everything down.
|
|
|
|
## Data
|
|
|
|
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
|
|
benchmark genome as the background sample, ClinVar for labels, gnomAD for allele frequencies, the
|
|
Human Phenotype Ontology's gene-to-phenotype annotations as what the ranking matches against, and
|
|
GA4GH phenopackets curated from case reports for the published case. Sources, licences, citations
|
|
and how the model should be evaluated honestly are in [docs/data.md](docs/data.md).
|
|
|
|
## Architecture
|
|
|
|
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
|
|
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
|
|
than AWS.
|
|
|
|
## Status
|
|
|
|
A self-training lab, built in the open. Working end to end:
|
|
|
|
1. Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
|
|
2. SvelteKit UI: cases, a phenotype-ranked candidate list showing the evidence behind each rank,
|
|
shortlist/dismiss decisions, and a case report
|
|
3. Kubernetes manifests, kind, Argo Workflows trigger
|
|
4. Terraform for GCP: serverless Cloud Run + Batch by default, GKE and ArgoCD behind a flag
|
|
5. Pathogenicity model trained on ClinVar, MLflow registry, scoring endpoint
|
|
6. A demo case built from a published patient, with the citations behind it
|
|
|
|
Known gaps: allele frequencies need the 25 GB VEP cache, because VEP's database mode returns
|
|
none, so the rarity term does no work without it; and the deployed demo has no authentication.
|
|
|
|
## Licence
|
|
|
|
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.
|