bedtools/bedtools_intersect
feature intersection
BAM
BED
GFF
VCF
overlap
Description
Find overlaps between genomic features from two sets of intervals.
bedtools intersect allows one to screen for overlaps between two sets of genomic features.
Moreover, it allows one to have fine control as to how the intersections are reported.
bedtools intersect works with both BED/GFF/VCF and BAM files as input.
Input arguments
Name | Type & Properties | Description |
|---|---|---|
--input_a -a | file required | The input file (BED/GFF/VCF/BAM) to be used as the -a file. |
--input_b -b | file required multiple | The input file(s) (BED/GFF/VCF/BAM) to be used as the -b file(s). |
Output arguments
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The output BED file. |
Output format options
Name | Type & Properties | Description |
|---|---|---|
--write_a -wa | boolean_true | Write the original A entry for each overlap. |
--write_b -wb | boolean_true | Write the original B entry for each overlap. Useful for knowing _what_ A overlaps. Restricted by -f and -r. |
--left_outer_join -loj | boolean_true | Perform a "left outer join". That is, for each feature in A report each overlap with B. If no overlaps are found, report a NULL feature for B. |
--write_overlap -wo | boolean_true | Write the original A and B entries plus the number of base pairs of overlap between the two features. Overlaps restricted by -f and -r. Only A features with overlap are reported. |
--write_overlap_plus -wao | boolean_true | Write the original A and B entries plus the number of base pairs of overlap between the two features. Overlaps restricted by -f and -r. However, A features w/o overlap are also reported with a NULL B feature and overlap = 0. |
--report_A_if_no_overlap -u | boolean_true | Write the original A entry _if_ no overlap is found. In other words, just report the fact >=1 hit was found. Overlaps restricted by -f and -r. |
--number_of_overlaps_A -c | boolean_true | For each entry in A, report the number of overlaps with B. Reports 0 for A entries that have no overlap with B. Overlaps restricted by -f and -r. |
--report_no_overlaps_A -v | boolean_true | Only report those entries in A that have _no overlaps_ with B. Similar to "grep -v" (an homage). |
--uncompressed_bam -ubam | boolean_true | Write uncompressed BAM output. Default writes compressed BAM. |
Filtering options
Name | Type & Properties | Description |
|---|---|---|
--same_strand -s | boolean_true | Require same strandedness. That is, only report hits in B that overlap A on the _same_ strand. By default, overlaps are reported without respect to strand. |
--opposite_strand -S | boolean_true | Require different strandedness. That is, only report hits in B that overlap A on the _opposite_ strand. By default, overlaps are reported without respect to strand. |
--min_overlap_A -f | double | Minimum overlap required as a fraction of A. Default is 1E-9 (i.e., 1bp). |
--min_overlap_B -F | double | Minimum overlap required as a fraction of B. Default is 1E-9 (i.e., 1bp). |
--reciprocal_overlap -r | boolean_true | Require that the fraction overlap be reciprocal for A AND B. - In other words, if -f is 0.90 and -r is used, this requires that B overlap 90% of A and A _also_ overlaps 90% of B. |
--either_overlap -e | boolean_true | Require that the minimum fraction be satisfied for A OR B. - In other words, if -e is used with -f 0.90 and -F 0.10 this requires that either 90% of A is covered OR 10% of B is covered. Without -e, both fractions would have to be satisfied. |
--split | boolean_true | Treat "split" BAM or BED12 entries as distinct BED intervals. |
--genome -g | file | Provide a genome file to enforce consistent chromosome sort order across input files. Only applies when used with -sorted option. |
--nonamecheck | boolean_true | For sorted data, don't throw an error if the file has different naming conventions for the same chromosome (e.g., "chr1" vs "chr01"). |
--sorted | boolean_true | Use the "chromsweep" algorithm for sorted (-k1,1 -k2,2n) input. |
--names | string | When using multiple databases, provide an alias for each that will appear instead of a fileId when also printing the DB record. |
--filenames | boolean_true | When using multiple databases, show each complete filename instead of a fileId when also printing the DB record. |
--sortout | boolean_true | When using multiple databases, sort the output DB hits for each record. |
--bed | boolean_true | If using BAM input, write output as BED. |
--header | boolean_true | Print the header from the A file prior to results. |
--no_buffer_output --nobuf | boolean_true | Disable buffered output. Using this option will cause each line of output to be printed as it is generated, rather than saved in a buffer. This will make printing large output files noticeably slower, but can be useful in conjunction with other software tools and scripts that need to process one line of bedtools output at a time. |
--io_buffer_size --iobuf | integer | Specify amount of memory to use for input buffer. Takes an integer argument. Optional suffixes K/M/G supported. Note: currently has no effect with compressed files. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/bedtools/bedtools_intersect/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
bedtools/bedtools_intersectbiobox v0.4.2
Uses
0 relationships
No component dependencies found.