bedtools/bedtools_spacing
Description
Calculate gaps between adjacent genomic intervals within each chromosome.
bedtools spacing analyzes sorted genomic intervals and reports the gap length
between each interval and its predecessor on the same chromosome. The gap distances
are added as an additional column to the output, providing insight into the spatial
distribution of features.
This tool is commonly used for:
Analyzing spacing patterns in genomic features
Quality control of interval datasets and their density
Identifying clustering or regular spacing in genomic data
Preprocessing for downstream spatial analysis
Detecting overlapping or adjacent intervals in datasets
Statistical analysis of genomic feature distribution
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input file containing genomic intervals for spacing analysis. **Format:** BED, GFF, VCF, or BAM file **Content:** Genomic intervals to analyze for gap spacing **Requirements:** Must be sorted by chromosome and start coordinate (sort -k1,1 -k2,2n) **Output:** Original intervals with gap distances appended as last column |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output file with gap spacing information appended. **Format:** Same format as input with additional spacing column **Content:** Original intervals plus gap distances in the last column **Gap values:** - "." for first interval on each chromosome - "0" for adjacent intervals - "-1" for overlapping intervals - Positive integers for gap sizes |
Format Options
Name | Type & Properties | Description |
|---|---|---|
--output_bed -bed | boolean_true | Convert BAM input to BED format output. **Usage:** Only applicable when input is BAM format **Effect:** Output genomic coordinates in BED format instead of maintaining BAM format **Applications:** Converting alignment data to interval format with spacing information **Default:** false (maintain input format) |
--include_header -header | boolean_true | Include the original file header in output. **Usage:** Preserves metadata and format information from input **Applications:** Maintaining file structure and annotations **Formats:** Particularly relevant for VCF and GFF files **Default:** false (no header included) |
Performance Options
Name | Type & Properties | Description |
|---|---|---|
--no_buffer -nobuf | boolean_true | Disable output buffering for real-time processing. **Effect:** Each line printed immediately instead of buffered **Trade-off:** Slower output but enables real-time processing **Applications:** Pipeline integration, streaming analysis **Default:** false (buffered output for performance) |
--input_buffer -iobuf | string | Amount of memory to allocate for input buffer. **Format:** Integer with optional K/M/G suffix **Examples:** "1G", "512M", "2048K" **Usage:** Larger buffers can improve I/O performance for large files **Note:** Currently has no effect with compressed files |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output.bed"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/bedtools/bedtools_spacing/main.nf \
-params-file params.yaml Relationships
Used by
No components use this component.
Current component
Uses
No component dependencies found.