Makes a real annotation runnable locally without the 25 GB VEP cache, which is what the demo needs and what a reviewer can reproduce in minutes. - params.vep_database (VEP_DATABASE=true) queries Ensembl's public database instead of a local cache. Slower per variant and fewer fields, so --everything is swapped for the flags the loader actually stores. Its cache placeholder is NO_CACHE, not NO_FILE: Nextflow rejects two staged inputs sharing a filename. - PIPELINE_DATABASE_URL is handed to the pipeline when set. The loader runs inside a container, where the API's own localhost URL would point at the container itself. - README: how to run the UI's annotate button locally against host Nextflow + Docker. Verified end to end on pipeline/tests/data/tiny.vcf: bcftools norm split the multiallelic record, VEP 113 annotated 4 variants live, the loader wrote them and marked the job succeeded, and the UI shows them. The deletion came back as 22:42126611 CT>C with exact VCF alleles, which is the case the audit's ID-tagging fix exists for. Tests: api 51, loader 16, stub run 3/3; ruff, mypy clean.
139 lines
6.4 KiB
Markdown
139 lines
6.4 KiB
Markdown
# rarelens
|
|
|
|
A small, end-to-end variant interpretation platform for rare genetic disease research.
|
|
Scientists upload a VCF, a Nextflow workflow annotates it with Ensembl VEP, a machine
|
|
learning model scores each variant, and results are browsable in a web app.
|
|
|
|
This repository is a **self-training lab**. It exists so that one engineer can learn, in
|
|
public, how a modern life-sciences platform is built end to end: full-stack application,
|
|
scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a
|
|
clinical tool and makes no diagnostic claims.
|
|
|
|
## What is in the box
|
|
|
|
| Layer | Technology | Directory |
|
|
|------------|--------------------------------------------------------|----------------------|
|
|
| Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` |
|
|
| API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` |
|
|
| Database | PostgreSQL 16 | `docker-compose.yml` |
|
|
| Frontend | SvelteKit, TypeScript | `web/` |
|
|
| ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` |
|
|
| Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` |
|
|
| Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` |
|
|
| Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` |
|
|
| CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` |
|
|
|
|
## Quick start (local)
|
|
|
|
```bash
|
|
make up # postgres + api + web + mlflow via docker-compose
|
|
make migrate # alembic upgrade head
|
|
make data # real public data: GIAB HG002 + ClinVar, chr22 (needs bcftools)
|
|
make test # api, ml, loader and web tests (no Docker needed for the DB tests)
|
|
```
|
|
|
|
Then open http://localhost:5173.
|
|
|
|
The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the
|
|
command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep`
|
|
and a VCF under `data/` (see [data/README.md](data/README.md)):
|
|
|
|
```bash
|
|
make annotate JOB=<job id from the UI> VCF=data/example.vcf.gz
|
|
make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database
|
|
```
|
|
|
|
No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant
|
|
and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the
|
|
pipeline on a handful of variants:
|
|
|
|
```bash
|
|
VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf
|
|
```
|
|
|
|
To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and
|
|
Docker are) rather than in docker-compose:
|
|
|
|
```bash
|
|
docker compose up -d db
|
|
cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \
|
|
PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:[email protected]:5432/rarelens \
|
|
LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \
|
|
uv run --extra dev uvicorn app.main:app --port 8000
|
|
```
|
|
|
|
`PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost`
|
|
would be the container itself. `LOCAL_DATA_ROOT` is the directory a sample's `vcf_uri` must sit under.
|
|
|
|
To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`):
|
|
|
|
```bash
|
|
cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \
|
|
uv run python -m rarelens_ml.train --tsv ../pipeline/results/<sample>.vep.tsv --register
|
|
```
|
|
|
|
Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies
|
|
`infra/k8s/overlays/local`.
|
|
|
|
## Deploying to GCP
|
|
|
|
Two tracks, same code. The serverless one is the default because it costs about £1/month idle;
|
|
[docs/cloud.md](docs/cloud.md) has the numbers.
|
|
|
|
**Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow
|
|
driver runs as a Cloud Run job only while a pipeline is running.
|
|
|
|
```bash
|
|
cd infra/terraform
|
|
terraform init -backend-config="bucket=<tfstate bucket>"
|
|
export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier
|
|
terraform apply -var project=<project id> # add -var deploy_cloud_sql=true to use Cloud SQL instead
|
|
cd ../.. && make serverless-deploy PROJECT=<project id> TAG=<commit sha> # redeploy a new build
|
|
```
|
|
|
|
`terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so
|
|
there is one public address and no CORS. Upload the VEP cache to
|
|
`gs://<project>-rarelens-data/refs/vep` before running a real annotation, and set
|
|
`-var model_uri=gs://<project>-rarelens-data/models/pathogenicity/1` to score without running an
|
|
MLflow server. Set a billing budget first — the demo has no authentication.
|
|
|
|
**Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path,
|
|
then destroy it.
|
|
|
|
```bash
|
|
terraform apply -var project=<project id> -var deploy_kubernetes=true -var deploy_cloud_sql=true
|
|
make gcp-configure PROJECT=<project id> # once; commit the result
|
|
make gcp-secrets PROJECT=<project id>
|
|
```
|
|
|
|
Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`.
|
|
Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them.
|
|
`make serverless-destroy PROJECT=<project id>` tears everything down.
|
|
|
|
## Data
|
|
|
|
The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002
|
|
benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources,
|
|
licences, citations and how the model should be evaluated honestly are in
|
|
[docs/data.md](docs/data.md).
|
|
|
|
## Architecture
|
|
|
|
See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind
|
|
each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather
|
|
than AWS.
|
|
|
|
## Status
|
|
|
|
Work in progress. Milestones, in order:
|
|
|
|
1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green
|
|
2. SvelteKit UI: sample list, variant table with filters, job status
|
|
3. Kubernetes manifests, kind, Argo Workflows trigger
|
|
4. Terraform for GCP, ArgoCD GitOps deploy
|
|
5. Pathogenicity model, MLflow registry, prediction endpoint
|
|
|
|
## Licence
|
|
|
|
AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.
|