# rarelens A small, end-to-end variant interpretation platform for rare genetic disease research. A case is a proband: a VCF plus the patient's phenotype (HPO terms). A Nextflow workflow annotates the variants with Ensembl VEP, a model scores each one, and the app narrows thousands of variants to a handful of candidates ranked against that phenotype — each carrying the evidence for its rank, and each able to be shortlisted or dismissed with a reason that ends up in a case report. This repository is a **self-training lab**. It exists so that one engineer can learn, in public, how a modern life-sciences platform is built end to end: full-stack application, scientific pipeline, ML serving, and cloud infrastructure, all in one monorepo. It is not a clinical tool and makes no diagnostic claims. ## What is in the box | Layer | Technology | Directory | |------------|--------------------------------------------------------|----------------------| | Pipeline | Nextflow DSL2, bcftools, Ensembl VEP, Docker, Google Batch | `pipeline/` | | API | FastAPI, Pydantic v2, SQLAlchemy 2.0 (async), Alembic | `api/` | | Database | PostgreSQL 16 | `docker-compose.yml` | | Frontend | SvelteKit, TypeScript | `web/` | | ML | LightGBM pathogenicity scorer, MLflow registry | `ml/` | | Orchestration | Argo Workflows + Argo Events (pipeline), Pub/Sub (events) | `infra/argo-workflows/` | | Platform | Kubernetes (Kustomize), ArgoCD (GitOps) | `infra/k8s/`, `infra/argocd/` | | Cloud | GCP: GKE Autopilot, Cloud SQL, GCS, Batch, Secret Manager, Artifact Registry | `infra/terraform/` | | CI/CD | GitHub Actions, Workload Identity Federation | `.github/workflows/` | ## Quick start (local) ```bash make up # postgres + api + web + mlflow via docker-compose make migrate # alembic upgrade head make hpo # HPO gene-to-phenotype annotations: what the ranking matches against make demo-case # a simulated proband: GIAB background + one ClinVar pathogenic variant make test # api, ml, loader and web tests (no Docker needed for the DB tests) ``` Then open http://localhost:5173, create a case pointing at `data/proband-simulated.vcf.gz`, give it the phenotype of the planted disease (for the default NF2 case: bilateral vestibular schwannoma, sensorineural hearing impairment, tinnitus, meningioma, cataract), and analyse it. The planted variant should come back ranked first. The docker-compose API has no Nextflow, so "Run VEP annotation" marks the job failed with the command to run instead. With Nextflow and Docker on the host, a VEP cache in `pipeline/cache/vep` and a VCF under `data/` (see [data/README.md](data/README.md)): ```bash make annotate JOB= VCF=data/example.vcf.gz make pipeline VCF=data/example.vcf.gz # dry run: annotate without touching the database ``` No cache? `VEP_DATABASE=true` queries Ensembl's public database instead. It is slow per variant and returns fewer fields, but it needs no 25 GB download, which is enough to demonstrate the pipeline on a handful of variants: ```bash VEP_DATABASE=true make pipeline VCF=pipeline/tests/data/tiny.vcf ``` To make the UI's "Run VEP annotation" button work, run the API on the host (where Nextflow and Docker are) rather than in docker-compose: ```bash docker compose up -d db cd api && DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@localhost:5432/rarelens \ PIPELINE_DATABASE_URL=postgresql+asyncpg://rarelens:rarelens@host.docker.internal:5432/rarelens \ LOCAL_DATA_ROOT=$PWD/.. VEP_DATABASE=true \ uv run --extra dev uvicorn app.main:app --port 8000 ``` `PIPELINE_DATABASE_URL` is what the loader container gets: inside it, the API's own `localhost` would be the container itself. `LOCAL_DATA_ROOT` is the directory a sample's `vcf_uri` must sit under. To train and register a model (the API scores with `models:/rarelens-pathogenicity@production`): ```bash cd ml && MLFLOW_TRACKING_URI=http://localhost:5000 \ uv run python -m rarelens_ml.train --tsv ../pipeline/results/.vep.tsv --register ``` Local Kubernetes: `make kind` builds the images, loads them into a kind cluster and applies `infra/k8s/overlays/local`. ## Deploying to GCP Two tracks, same code. The serverless one is the default because it costs about £1/month idle; [docs/cloud.md](docs/cloud.md) has the numbers. **Serverless (Cloud Run + Google Batch).** The API and the UI scale to zero, and the Nextflow driver runs as a Cloud Run job only while a pipeline is running. ```bash cd infra/terraform terraform init -backend-config="bucket=" export TF_VAR_database_url='postgresql+asyncpg://user:pass@host/db?sslmode=require' # e.g. Neon's free tier terraform apply -var project= # add -var deploy_cloud_sql=true to use Cloud SQL instead cd ../.. && make serverless-deploy PROJECT= TAG= # redeploy a new build ``` `terraform output web_url` is the URL to share; it serves the UI and proxies `/api` to the API, so there is one public address and no CORS. Upload the VEP cache to `gs://-rarelens-data/refs/vep` before running a real annotation, and set `-var model_uri=gs://-rarelens-data/models/pathogenicity/1` to score without running an MLflow server. Set a billing budget first — the demo has no authentication. **Kubernetes (GKE + Argo + ArgoCD).** Off by default; turn it on to demonstrate the GitOps path, then destroy it. ```bash terraform apply -var project= -var deploy_kubernetes=true -var deploy_cloud_sql=true make gcp-configure PROJECT= # once; commit the result make gcp-secrets PROJECT= ``` Then install Argo Workflows, Argo Events and ArgoCD, and `kubectl apply -f infra/argocd/app.yaml`. Every green CI run on `main` bumps image tags in the gcp overlay and ArgoCD deploys them. `make serverless-destroy PROJECT=` tears everything down. ## Data The demo runs on published, openly licensed human data: the NIST Genome in a Bottle HG002 benchmark genome as the sample, ClinVar for labels, gnomAD for allele frequencies. Sources, licences, citations and how the model should be evaluated honestly are in [docs/data.md](docs/data.md). ## Architecture See [docs/architecture.md](docs/architecture.md) for the diagram and the reasoning behind each choice, and [docs/cloud.md](docs/cloud.md) for why this deploys to Google Cloud rather than AWS. ## Status Work in progress. Milestones, in order: 1. Skeleton, Postgres, FastAPI, Nextflow VEP annotation on a public VCF, CI green 2. SvelteKit UI: sample list, variant table with filters, job status 3. Kubernetes manifests, kind, Argo Workflows trigger 4. Terraform for GCP, ArgoCD GitOps deploy 5. Pathogenicity model, MLflow registry, prediction endpoint ## Licence AGPL-3.0. Test data are public (ClinVar, gnomAD subsets); no patient data are used or accepted.