An end-to-end audit found the repo could not build, test or run as shipped. This fixes every finding, then adds a Cloud Run track so the demo costs about £1/month idle instead of ~£150. CI (red on its first run) - api: setuptools could not build the package (flat layout with app/ and alembic/) - web: missing @types/node; `vitest run` exited 1 with no test files - pipeline: the stub run needed a gitignored VCF, and no process had a stub block - ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI) ML serving (scores were meaningless) - the registered model now carries its own feature engineering and returns predict_proba, so serving sends raw columns and cannot drift from training - resolve by registry alias (stages are deprecated in MLflow 3) and record the real version; re-scoring upserts instead of failing on the unique constraint - ClinVar labels parsed from VEP's lowercase terms Pipeline - exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent - job status reaches running/failed/succeeded, so the UI stops polling dead jobs - DATABASE_URL travels in the environment or a Nextflow secret, never on a command line - VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch Deployment - the API serves /api (matching the ingress); the web app reads its API URL at runtime - migrations run in an init container under a Postgres advisory lock - terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload Identity bindings, Secret Manager, deletion protection - serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md. Correctness and security - 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider VEP text columns, enum dropped on downgrade, the sample's assembly actually used - vcf_uri restricted to gs:// objects or files under the data root, blocking option injection - CORS restricted to configured origins; `make down` no longer deletes volumes Data - docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar, gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform validate and both kustomize overlays clean.
87 lines
4.7 KiB
Markdown
87 lines
4.7 KiB
Markdown
# Architecture
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
U[Scientist] -->|browser| W[SvelteKit web]
|
|
W -->|REST /api| A[FastAPI]
|
|
A --> P[(PostgreSQL / Cloud SQL)]
|
|
A -->|publish vcf-uploaded| Q[Pub/Sub]
|
|
Q --> E[Argo Events sensor]
|
|
E --> AW[Argo Workflow: Nextflow driver]
|
|
AW -->|tasks| B[Google Batch: bcftools norm, VEP, load_db]
|
|
B -->|reads VCF, VEP cache| G[(GCS bucket)]
|
|
B -->|writes variants, marks job succeeded| P
|
|
AW -.->|exit handler marks job failed| P
|
|
A -->|models:/rarelens-pathogenicity@production| M[MLflow registry]
|
|
T[ml/train.py] --> M
|
|
GH[GitHub Actions] -->|images via WIF| AR[Artifact Registry]
|
|
GH -->|bumps overlay tags| R[(git: infra/k8s/overlays/gcp)]
|
|
R --> CD[ArgoCD] --> K[GKE Autopilot]
|
|
```
|
|
|
|
## Two deployment tracks
|
|
|
|
The same images and the same pipeline, deployed two ways (`infra/terraform/variables.tf`):
|
|
|
|
| | Serverless (default) | Kubernetes (`-var deploy_kubernetes=true`) |
|
|
|---|---|---|
|
|
| api, web | Cloud Run, scale to zero | Deployments behind an ingress |
|
|
| dispatch | the API executes a Cloud Run job | Pub/Sub -> Argo Events -> Argo Workflow |
|
|
| pipeline tasks | Google Batch | Google Batch |
|
|
| `/api` routing | the web service proxies it | the ingress routes it |
|
|
| idle cost | ~£1/month | ~£130+/month |
|
|
|
|
`app.services.events.launch()` picks the dispatch backend from configuration: a Cloud Run job when
|
|
`CLOUDRUN_JOB` is set, Pub/Sub when `PUBSUB_TOPIC` is, and a local Nextflow process otherwise.
|
|
See [cloud.md](cloud.md) for why the serverless one is the default.
|
|
|
|
## Why these choices
|
|
|
|
**One monorepo.** The four components share a schema (`variants` table, feature columns) and the
|
|
point of the exercise is to see them evolve together. Separate repos would hide the coupling.
|
|
|
|
**Nextflow for the science, Argo Workflows for the trigger.** Nextflow is the lingua franca for
|
|
bioinformatics pipelines. Argo is what the platform team already runs. So Argo owns *when* a
|
|
pipeline runs; Nextflow owns *what* it does. The API never talks to Kubernetes directly; it
|
|
publishes an event and gets on with its life. The Nextflow driver runs in the Argo pod and sends
|
|
each task to Google Batch: a `gs://` work directory needs an executor that stages through GCS
|
|
(Nextflow's Kubernetes executor needs a shared ReadWriteMany volume instead).
|
|
|
|
**Job lifecycle.** The API creates a job as `running` once the pipeline is dispatched, or `failed`
|
|
with the reason in `jobs.log` when dispatch is impossible. The loader marks it `succeeded` in the
|
|
same transaction that stores the variants. Anything else (Nextflow error, eviction) is caught by
|
|
the Argo exit handler, or locally by the API watching the Nextflow process, and marked `failed`,
|
|
so the UI never polls a dead job.
|
|
|
|
**Variant identity.** NORMALISE sets each VCF ID to `CHROM_POS_REF_ALT`; VEP echoes it as
|
|
`Uploaded_variation` and the loader takes exact VCF alleles from it, because VEP's own
|
|
`Location`/`Allele` columns trim indel alleles.
|
|
|
|
**FastAPI + Pydantic v2 + SQLAlchemy 2.0 async.** Typed at both boundaries: request/response models
|
|
and ORM models are separate on purpose so the database can change without breaking the frontend
|
|
contract. Alembic owns the schema; the loader script writes raw SQL against that schema, not the ORM,
|
|
because the pipeline container should not import the API.
|
|
|
|
**SvelteKit.** Small runtime, no virtual DOM, and Svelte 5 runes make server-driven state simple.
|
|
The UI has exactly two pages; the goal is a table a scientist actually wants to filter, not a dashboard.
|
|
`PUBLIC_API_URL` is read at runtime, so one image works behind the ingress (`/api`) and elsewhere.
|
|
|
|
**GKE Autopilot + Cloud SQL, not self-managed.** The lab is about the platform patterns (Workload
|
|
Identity, GitOps, Kustomize overlays, private networking), not about running etcd. Cloud SQL has only
|
|
a private IP; the API reaches it through a Cloud SQL Proxy sidecar, pipeline tasks directly in the VPC.
|
|
Database URLs live in Secret Manager.
|
|
|
|
**GitOps.** CI builds and tests; it never runs `kubectl apply`. It edits image tags in the `gcp` overlay
|
|
and ArgoCD reconciles. Rollback is `git revert`. Migrations run in an init container under a Postgres
|
|
advisory lock, so replicas starting together migrate once.
|
|
|
|
**MLflow registry as the model contract.** The API loads whichever version the `production` alias
|
|
points at and records that version on every prediction. The registered model is a pyfunc that owns its
|
|
feature engineering (`rarelens_ml.features` ships inside it) and returns P(pathogenic), so serving only
|
|
sends raw columns and cannot drift from training.
|
|
|
|
## What is deliberately missing
|
|
|
|
Authentication, PHI handling, audit logs, clinical validation. This is a learning platform on public
|
|
data. Adding Identity-Aware Proxy in front of the ingress is the first step if that ever changes.
|