bedtools/bedtools_sample
Description
Take a random sample of records from BED/GFF/VCF/BAM files using reservoir sampling algorithm.
bedtools sample uses the reservoir sampling algorithm to randomly select a specified number
of records from genomic interval files. This is particularly useful for creating representative
subsets of large datasets for testing, quality control, or downstream analysis.
This tool is commonly used for:
Creating representative subsets of large genomic datasets
Quality control and validation with smaller sample sizes
Testing pipelines with manageable data volumes
Generating training datasets for machine learning applications
Reducing file sizes while maintaining statistical properties
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input file to sample from. **Format:** BED, GFF, VCF, or BAM file **Content:** Genomic intervals or alignments to sample from **Usage:** Records will be randomly selected from this file **Requirements:** File must be in valid format for the specified type |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output file containing the sampled records. **Format:** Same format as input file **Content:** Randomly selected subset of input records **Size:** Contains exactly n records (unless input has fewer records) |
Sampling Options
Name | Type & Properties | Description |
|---|---|---|
--number -n | integer | Number of records to randomly sample. **Default:** 1,000,000 **Range:** 1 to total number of records in input **Usage:** If input has fewer records than requested, all records are returned **Memory:** All selected records held in memory before output |
--seed | integer | Integer seed for random number generation. **Usage:** Ensures reproducible random sampling **Range:** Any integer value **Default:** Automatically chosen seed (non-reproducible) **Applications:** Reproducible research, testing, validation |
Strand Options
Name | Type & Properties | Description |
|---|---|---|
--strand_requirement -s | string | Require records from specific strand orientation. **Values:** "forward", "reverse", or unspecified for both strands **Usage:** Only sample records from the specified strand **Requirements:** Input must contain strand information **Applications:** Strand-specific analyses, RNA-seq processing |
Output Format Options
Name | Type & Properties | Description |
|---|---|---|
--output_bed -bed | boolean_true | Convert BAM input to BED format output. **Usage:** Only applicable when input is BAM format **Effect:** Output genomic coordinates in BED format instead of BAM **Applications:** Converting alignment data to interval format **Default:** false (maintain input format) |
--uncompressed_bam -ubam | boolean_true | Write uncompressed BAM output. **Usage:** Only applicable when input is BAM format **Effect:** Output BAM file without compression **Trade-off:** Faster writing but larger file size **Default:** false (compressed BAM output) |
--include_header -header | boolean_true | Include the original file header in output. **Usage:** Preserves metadata from input file **Applications:** Maintaining file structure and annotations **Formats:** Particularly relevant for VCF and GFF files **Default:** false (no header included) |
Performance Options
Name | Type & Properties | Description |
|---|---|---|
--no_buffer -nobuf | boolean_true | Disable output buffering for real-time processing. **Effect:** Each line printed immediately instead of buffered **Trade-off:** Slower output but enables real-time processing **Applications:** Pipeline integration, streaming processing **Default:** false (buffered output for performance) |
--input_buffer -iobuf | string | Amount of memory to allocate for input buffer. **Format:** Integer with optional K/M/G suffix **Examples:** "1G", "512M", "2048K" **Usage:** Larger buffers can improve I/O performance **Note:** Currently has no effect with compressed files |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output.bed"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/bedtools/bedtools_sample/main.nf \
-params-file params.yaml Relationships
Used by
No components use this component.
Current component
Uses
No component dependencies found.