fix: overhaul the platform skeleton, add a serverless deployment track
An end-to-end audit found the repo could not build, test or run as shipped. This fixes every finding, then adds a Cloud Run track so the demo costs about £1/month idle instead of ~£150. CI (red on its first run) - api: setuptools could not build the package (flat layout with app/ and alembic/) - web: missing @types/node; `vitest run` exited 1 with no test files - pipeline: the stub run needed a gitignored VCF, and no process had a stub block - ruff pinned, mypy configured, DB tests on real Postgres (pgserver locally, service in CI) ML serving (scores were meaningless) - the registered model now carries its own feature engineering and returns predict_proba, so serving sends raw columns and cannot drift from training - resolve by registry alias (stages are deprecated in MLflow 3) and record the real version; re-scoring upserts instead of failing on the unique constraint - ClinVar labels parsed from VEP's lowercase terms Pipeline - exact ref/alt recovered from a CHROM_POS_REF_ALT VCF ID; loading is idempotent - job status reaches running/failed/succeeded, so the UI stops polling dead jobs - DATABASE_URL travels in the environment or a Nextflow secret, never on a command line - VEP cache and plugins staged as inputs; the gcp profile runs tasks on Google Batch Deployment - the API serves /api (matching the ingress); the web app reads its API URL at runtime - migrations run in an init container under a Postgres advisory lock - terraform: custom VPC shared with Batch, private Cloud SQL, API enablement, Workload Identity bindings, Secret Manager, deletion protection - serverless track, now the default: Cloud Run services scaling to zero, a Cloud Run job for the Nextflow driver, and Neon or Cloud SQL behind one DATABASE_URL secret. GKE and Argo remain, behind -var deploy_kubernetes=true. See docs/cloud.md. Correctness and security - 409 on duplicate sample names, 422 on bad paging, natural chromosome ordering, wider VEP text columns, enum dropped on downgrade, the sample's assembly actually used - vcf_uri restricted to gs:// objects or files under the data root, blocking option injection - CORS restricted to configured origins; `make down` no longer deletes volumes Data - docs/data.md records the peer-reviewed, openly licensed sources (GIAB HG002, ClinVar, gnomAD) with citations and an honest evaluation plan; `make data` fetches a chr22 slice Verified: api 50 tests, ml 18, loader 16, web 12; ruff, mypy, svelte-check, terraform validate and both kustomize overlays clean.
This commit is contained in:
@@ -1,37 +1,54 @@
|
||||
# Triggered by an Argo Events sensor listening on the Pub/Sub topic "vcf-uploaded".
|
||||
# Submitted by the Argo Events sensor in events.yaml for each "vcf-uploaded" Pub/Sub message.
|
||||
# The Nextflow driver runs here; its tasks run on Google Batch (see the gcp profile in
|
||||
# pipeline/nextflow.config). Image names are rewritten by the gcp overlay and bumped by CI.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata: { name: annotate-vcf, namespace: rarelens }
|
||||
metadata: { name: annotate-vcf }
|
||||
spec:
|
||||
serviceAccountName: rarelens-pipeline
|
||||
entrypoint: nextflow
|
||||
onExit: exit-handler
|
||||
arguments:
|
||||
parameters:
|
||||
- { name: job_id }
|
||||
- { name: vcf_uri }
|
||||
- { name: assembly, value: GRCh38 }
|
||||
templates:
|
||||
- name: nextflow
|
||||
inputs:
|
||||
parameters: [{ name: job_id }, { name: vcf_uri }]
|
||||
serviceAccountName: rarelens-pipeline
|
||||
container:
|
||||
image: europe-west2-docker.pkg.dev/PROJECT/rarelens/pipeline:latest
|
||||
command: [nextflow]
|
||||
image: rarelens/pipeline
|
||||
args:
|
||||
- run
|
||||
- /pipeline/main.nf
|
||||
- -profile
|
||||
- gcp
|
||||
- --vcf
|
||||
- "{{inputs.parameters.vcf_uri}}"
|
||||
- "{{workflow.parameters.vcf_uri}}"
|
||||
- --job_id
|
||||
- "{{inputs.parameters.job_id}}"
|
||||
- --db_url
|
||||
- "$(DATABASE_URL)"
|
||||
envFrom: [{ secretRef: { name: api-secrets } }]
|
||||
resources: { requests: { cpu: "2", memory: 4Gi } }
|
||||
- name: score
|
||||
# Optional GPU step for the deep-learning baseline; Autopilot schedules on an L4 node.
|
||||
nodeSelector: { cloud.google.com/gke-accelerator: nvidia-l4 }
|
||||
- "{{workflow.parameters.job_id}}"
|
||||
- --assembly
|
||||
- "{{workflow.parameters.assembly}}"
|
||||
# GCP_PROJECT / GCS_BUCKET / GCP_REGION feed params in nextflow.config.
|
||||
envFrom: [{ configMapRef: { name: pipeline-config } }]
|
||||
resources: { requests: { cpu: "1", memory: 2Gi } }
|
||||
|
||||
# The loader marks success; anything else (Nextflow error, OOM, eviction) is marked here
|
||||
# so the UI never polls a dead job forever.
|
||||
- name: exit-handler
|
||||
steps:
|
||||
- - name: mark-failed
|
||||
template: mark-failed
|
||||
when: "{{workflow.status}} != Succeeded"
|
||||
- name: mark-failed
|
||||
container:
|
||||
image: europe-west2-docker.pkg.dev/PROJECT/rarelens/ml:latest
|
||||
resources: { limits: { nvidia.com/gpu: 1 } }
|
||||
image: rarelens/loader
|
||||
command: [set_job_status.py]
|
||||
args:
|
||||
- --job-id
|
||||
- "{{workflow.parameters.job_id}}"
|
||||
- --status
|
||||
- failed
|
||||
- --log
|
||||
- "Argo workflow {{workflow.name}} ended {{workflow.status}}"
|
||||
envFrom: [{ secretRef: { name: pipeline-secrets } }]
|
||||
resources: { requests: { cpu: 100m, memory: 256Mi } }
|
||||
|
||||
Reference in New Issue
Block a user