bedtools/bedtools_getfasta
sequencing
fasta
BED
GFF
VCF
sequence extraction
Description
Extract DNA sequences from a FASTA file based on feature coordinates.
Given intervals specified in BED/GFF/VCF format and a FASTA file, this tool
extracts the corresponding sequences from the FASTA file. Various output formats
are supported including FASTA (default), tab-delimited, and BED format with sequences.
Input arguments
Name | Type & Properties | Description |
|---|---|---|
--input_fasta -fi | file required | Input FASTA file containing sequences for extraction. The headers in the input FASTA file must exactly match the chromosome column in the BED file. |
--input_bed -bed | file required | BED/GFF/VCF file containing intervals to extract from the FASTA file. BED files containing a single region require a newline character at the end of the line, otherwise a blank output file is produced. |
--rna | boolean_true | The FASTA is RNA not DNA. Reverse complementation handled accordingly. |
Processing options
Name | Type & Properties | Description |
|---|---|---|
--strandedness -s | boolean_true | Force strandedness. If the feature occupies the antisense strand, the output sequence will be reverse complemented. By default strandedness is not taken into account. |
--split | boolean_true | When input is in BED12 format, create a separate FASTA entry for each block in a BED12 record. Blocks are described in the 11th and 12th columns of the BED format. |
--full_header -fullHeader | boolean_true | Use full FASTA header. By default, only the word before the first space or tab is used. |
Output arguments
Name | Type & Properties | Description |
|---|---|---|
--output -o -fo | file required output | Output file where the extracted sequences will be written. By default, output is in FASTA format unless --tab or --bed_out is specified. |
--name | boolean_true | Set the FASTA header for each extracted sequence to be the "name" and coordinate columns from the BED feature (format: name::chr:start-end). |
--name_only -nameOnly | boolean_true | Set the FASTA header for each extracted sequence to be only the "name" column from the BED feature. |
--tab | boolean_true | Report extracted sequences in a tab-delimited format instead of FASTA format. Output format: name<tab>sequence. |
--bed_out -bedOut | boolean_true | Report extracted sequences in a tab-delimited BED format instead of FASTA format. Output format: chr<tab>start<tab>end<tab>name<tab>sequence. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.0 \
-main-script target/nextflow/bedtools/bedtools_getfasta/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
bedtools/bedtools_getfastabiobox v0.4.0
Uses
0 relationships
No component dependencies found.