ensembl_vep/vep
variant effect prediction
annotation
genomics
consequence
missense
Description
Variant Effect Predictor (VEP) determines the effect of variants on genes, transcripts, and protein sequences.
VEP annotates variants with:
Consequence types (missense, nonsense, splice site, etc.)
Gene and transcript information
Protein effect predictions (SIFT, PolyPhen)
Population frequencies and clinical significance
Conservation scores and regulatory features
See the VEP documentation for comprehensive usage.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input_file -i | file required | Input variant file for annotation. **Formats:** VCF, VEP, variant_identifier, HGVS, ID, region, SPDI **Compression:** Supports gzip (.gz) and bgzip compressed files |
--format | string | Input file format. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output_file -o | file required output | Output annotated variants to file. **Compression:** Automatically compress if filename ends with .gz |
--vcf | boolean_true | Output results in VCF format. **Default:** Tab-delimited text format **VCF mode:** Adds INFO field with VEP annotations |
Reference Data
Name | Type & Properties | Description |
|---|---|---|
--species | string | Species for annotation. |
--assembly | string | Genome assembly version. |
--cache | boolean_true | Use cache files for annotation (recommended). **Performance:** Much faster than database queries **Location:** Specify with --dir or use default ~/.vep/ |
--dir | file | Directory containing cache files. |
--cache_version | integer | Use specific cache version. |
--offline | boolean_true | Enable offline mode (no database connections). **Requires:** Cache files and FASTA files **Use case:** Isolated environments or reproducible runs |
Annotation Options
Name | Type & Properties | Description |
|---|---|---|
--everything | boolean_true | Shortcut to switch on commonly used options. **Includes:** --sift, --polyphen, --ccds, --hgvs, --symbol, --numbers, --domains, --regulatory, --canonical, --protein, --biotype, --uniprot, --tsl, --appris, --gene_phenotype, --af, --af_1kg, --af_esp, --af_gnomad, --max_af, --pubmed, --variant_class, --mane |
--canonical | boolean_true | Add canonical transcript flag to output. **Canonical:** Transcript selected as representative for each gene |
--ccds | boolean_true | Add CCDS transcript identifiers. |
--protein | boolean_true | Add Ensembl protein identifiers. |
--symbol | boolean_true | Add gene symbol (e.g., HGNC). |
--hgvs | boolean_true | Add HGVS nomenclature for variants. |
--sift | string | Add SIFT prediction and/or score. |
--polyphen | string | Add PolyPhen prediction and/or score. |
Frequency Options
Name | Type & Properties | Description |
|---|---|---|
--af | boolean_true | Add global allele frequency from 1000 Genomes. |
--af_1kg | boolean_true | Add continental allele frequencies from 1000 Genomes. |
--af_gnomad | boolean_true | Add allele frequencies from gnomAD exomes collections. |
--max_af | boolean_true | Add maximum observed allele frequency across populations. |
Filtering Options
Name | Type & Properties | Description |
|---|---|---|
--pick | boolean_true | Pick one consequence annotation per variant. **Selection:** Most severe consequence per gene **Output:** Single annotation line per variant |
--pick_allele | boolean_true | Pick one consequence annotation per allele. |
--flag_pick | boolean_true | Flag picked consequence with PICK=1. |
--per_gene | boolean_true | Use one consequence annotation per gene. |
--pick_order | string | Customize annotation selection criteria order. |
Output Filtering
Name | Type & Properties | Description |
|---|---|---|
--most_severe | boolean_true | Output only most severe consequence per variant. |
--summary | boolean_true | Output only summary of consequences for each variant. |
--filter_common | boolean_true | Shortcut to exclude common variants (frequency > 0.01). |
Advanced Options
Name | Type & Properties | Description |
|---|---|---|
--buffer_size | integer | Number of variants to read at once. |
--no_check_variants_order | boolean_true | Permit variants not ordered by position. |
--allow_non_variant | boolean_true | Allow non-variant VCF lines through analysis. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
format: [ "vcf" ]
output_file: "$id.$key.output_file.txt"
species: [ "homo_sapiens" ]
dir: [ "/root/.vep" ]
buffer_size: [ 5000 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/ensembl_vep/vep/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
ensembl_vep/vepbiobox v0.4.2
Uses
0 relationships
No component dependencies found.