bedtools/bedtools_split
Description
Split a BED file into multiple output files using different algorithms.
bedtools split divides a single BED file into multiple smaller BED files
based on the specified number of output files and splitting algorithm. This
is useful for parallelizing analysis workflows, creating balanced datasets,
or distributing genomic intervals across multiple processing units.
This tool is commonly used for:
Parallelizing genomic analysis workflows by splitting input data
Creating balanced datasets for distributed computing
Dividing large BED files for memory-efficient processing
Preparing input files for parallel execution frameworks
Load balancing genomic intervals across multiple processes
Creating subsets of genomic data for testing or validation
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input BED file to split into multiple files. **Format:** BED file with genomic intervals **Content:** Genomic intervals to be distributed across output files **Requirements:** Standard BED format with at least 3 columns (chr, start, end) **Memory:** File contents are loaded into memory during processing **Size limits:** Available system memory determines maximum input size |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--number -n | integer required | Number of output files to create from the input BED file. **Range:** Positive integer (typically 2 or more) **Distribution:** Input intervals will be distributed across this many files **Naming:** Output files numbered sequentially with 5-digit padding (prefix.00001.bed, prefix.00002.bed, etc.) **Algorithm dependency:** Distribution method depends on selected algorithm |
--prefix -p | string | Prefix for output BED filenames. **Default:** "split" (creates split.00001.bed, split.00002.bed, etc.) **Pattern:** Files named as "{prefix}.{5-digit-number}.bed" **Directory:** Files created in current working directory unless path specified **Examples:** "sample" → sample.00001.bed, sample.00002.bed |
--output_dir | string required | Directory where output files should be created. **Usage:** Specify target directory for all split output files **Creation:** Directory will be created if it doesn't exist **Path:** Can be relative or absolute path **Required:** Must be specified for output file placement |
Algorithm Options
Name | Type & Properties | Description |
|---|---|---|
--algorithm -a | string | Algorithm used to split the BED file data. **size (default):** Uses heuristic algorithm to group intervals so all output files contain approximately the same total number of base pairs. Best for balanced genomic coverage across files. **simple:** Routes records so each file has approximately equal number of intervals (like Unix split command). Best for balanced record counts regardless of interval sizes. **Trade-offs:** - size: Balanced coverage, variable record counts - simple: Balanced record counts, variable coverage |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/bedtools/bedtools_split/main.nf \
-params-file params.yaml Relationships
Used by
No components use this component.
Current component
Uses
No component dependencies found.