# Test data `make data` fetches the demo slice automatically. Provenance, licences, citations and the evaluation plan live in [../docs/data.md](../docs/data.md). No patient data. Use public sources only: - ClinVar VCF (GRCh38): https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/ - gnomAD exomes subset for allele frequencies - A small HG002 (GIAB) chr22 slice for a realistic germline sample ```bash # Example: 2,000 ClinVar variants on chr22 as a smoke-test VCF wget -O clinvar.vcf.gz https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/clinvar.vcf.gz tabix -p vcf clinvar.vcf.gz bcftools view -r 22 clinvar.vcf.gz | bcftools view -H | head -2000 > body.vcf (bcftools view -h clinvar.vcf.gz; cat body.vcf) | bgzip > example.vcf.gz tabix -p vcf example.vcf.gz ``` Samples added in the UI must point at a `gs://bucket/object` or at a file under `/data` (`LOCAL_DATA_ROOT`), ending in `.vcf`, `.vcf.gz`, `.vcf.bgz` or `.bcf`. ## VEP cache and plugins The pipeline runs VEP offline. Install the cache once (about 25 GB for GRCh38): ```bash docker run --rm -v $PWD/pipeline/cache/vep:/cache ensemblorg/ensembl-vep:release_113.0 \ INSTALL.pl -a cf -s homo_sapiens -y GRCh38 -c /cache ``` CADD and AlphaMissense are optional; the model treats their scores as missing without them. To enable them, put the plugin modules and data files in one directory and pass `--vep_plugin_data`: ```bash docker run --rm -v $PWD/pipeline/cache/plugins:/plugins ensemblorg/ensembl-vep:release_113.0 \ INSTALL.pl -a p -g CADD,AlphaMissense -r /plugins # then add whole_genome_SNVs.tsv.gz, gnomad.genomes.r4.0.indel.tsv.gz (CADD) and # AlphaMissense_hg38.tsv.gz, each with its .tbi index, to pipeline/cache/plugins ``` `pipeline/tests/data/tiny.vcf` is a synthetic three-record fixture used by CI's stub run.