bcftools/bcftools_norm
Normalize
VCF
BCF
Description
Left-align and normalize indels, check if REF alleles match the reference, split multiallelic sites into multiple rows;
recover multiallelics from multiple rows.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input VCF/BCF file. The file to be normalized, left-aligned, and/or processed. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Write output to a file. If not specified, output goes to standard output. |
Options
Name | Type & Properties | Description |
|---|---|---|
--atomize -a | boolean_true | Decompose complex variants (e.g., MNVs become consecutive SNVs). Breaks down complex variants into simpler components. |
--atom_overlaps | string | Use the star allele (*) for overlapping alleles or set to missing (.). **Options:** - `*`: Use star allele for overlaps (default) - `.`: Set overlapping alleles to missing |
--check_ref -c | string | Check REF alleles and exit (e), warn (w), exclude (x), or set (s) bad sites. **Options:** - `e`: exit on REF mismatch (default) - `w`: warn about REF mismatches - `x`: exclude sites with REF mismatches - `s`: set/fix REF mismatches |
--remove_duplicates_flag -D | boolean_true | Remove duplicate lines of the same type. Shorthand for --rm_dup exact. |
--rm_dup -d | string | Remove duplicate snps|indels|both|all|exact. **Options:** - `snps`: Remove duplicate SNPs - `indels`: Remove duplicate indels - `both`: Remove duplicate SNPs and indels - `all`: Remove all duplicates - `exact`: Remove exact duplicates only |
--exclude -e | string | Do not normalize records for which the expression is true. Uses bcftools expression syntax (see man page for details). |
--fasta_ref -f | file | Reference sequence file. Required for checking REF alleles and left-alignment. |
--force | boolean_true | Try to proceed even if malformed tags are encountered. **Warning:** Experimental feature, use at your own risk. |
--gff_annot -g | file | Follow HGVS 3'rule and right-align variants in transcripts on the forward strand. Uses GFF annotation file for transcript information. |
--include -i | string | Normalize only records for which the expression is true. Uses bcftools expression syntax (see man page for details). |
--keep_sum | string | Keep vector sum constant when splitting multiallelics. Comma-separated list of INFO tags (see github issue #360). |
--multiallelics -m | string | Split multiallelics (-) or join biallelics (+), type: snps|indels|both|any. **Options:** - `-both`: Split multiallelic sites (default) - `+both`: Join biallelic sites - Use `snps`, `indels`, `any` for specific variant types |
--multi_overlaps | string | Fill in the reference (0) or missing (.) allele when splitting multiallelics. **Options:** - `0`: Fill with reference allele (default) - `.`: Fill with missing allele |
--no_version | boolean_true | Do not append version and command line to the header. Produces cleaner output headers. |
--do_not_normalize -N | boolean_true | Do not normalize indels (with -m or -c s). Skips indel left-alignment and normalization. |
--old_rec_tag | string | Annotate modified records with INFO/STR indicating the original variant. Adds specified INFO tag to track original variants. |
--output_type -O | string | Output type and compression level. **Options:** - `u`: uncompressed BCF - `b`: compressed BCF - `v`: uncompressed VCF (default) - `z`: compressed VCF (with optional compression level 0-9) |
--regions -r | string | Restrict to comma-separated list of regions. **Formats supported:** chr|chr:pos|chr:beg-end|chr:beg-[,…] |
--regions_file -R | file | Restrict to regions listed in a file. Regions can be specified in VCF, BED, or tab-delimited format. |
--regions_overlap | string | Include if POS in the region (0), record overlaps (1), variant overlaps (2). **Options:** - `0`: POS inside region (default for -t/-T) - `1`: overlapping records included (default for -r/-R) - `2`: true overlapping variation only |
--strict_filter -s | boolean_true | When merging (-m+), merged site is PASS only if all sites being merged PASS. Stricter FILTER handling during multiallelic joining. |
--sort -S | string | Sort order: chr_pos,lex. **Options:** - `chr_pos`: Sort by chromosome and position (default) - `lex`: Lexicographic sort |
--targets -t | string | Similar to --regions but streams rather than index-jumps. More efficient for processing many small regions. |
--targets_file -T | file | Similar to --regions_file but streams rather than index-jumps. More efficient for processing many regions from file. |
--targets_overlap | string | Include if POS in the region (0), record overlaps (1), variant overlaps (2). Similar to --regions_overlap but for streaming mode. |
--verbosity -v | integer | Verbosity level. Controls amount of diagnostic output. |
--site_win -w | integer | Buffer for sorting lines which changed position during realignment. Larger values use more memory but handle more complex rearrangements. |
--write_index -W | string | Automatically index the output files. **Format:** Specify index format or use default. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output.gz"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.0 \
-main-script target/nextflow/bcftools/bcftools_norm/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
bcftools/bcftools_normbiobox v0.4.0
Uses
0 relationships
No component dependencies found.