bedtools/bedtools_groupby
groupby
BED
Description
Summarizes a dataset column based upon common column groupings.
Akin to the SQL "group by" command.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | The input BED file to be used. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The output groupby BED file. |
Options
Name | Type & Properties | Description |
|---|---|---|
--groupby -g -grp | string required | Specify the columns (1-based) for the grouping. The columns must be comma separated. - Default: 1,2,3 |
--column -c -opCols | integer required | Specify the column (1-based) that should be summarized. |
--operation -o -ops | string | Specify the operation that should be applied to opCol. Valid operations: sum, count, count_distinct, min, max, mean, median, mode, antimode, stdev, sstdev (sample standard dev.), collapse (i.e., print a comma separated list (duplicates allowed)), distinct (i.e., print a comma separated list (NO duplicates allowed)), distinct_sort_num (as distinct, but sorted numerically, ascending), distinct_sort_num_desc (as distinct, but sorted numerically, descending), concat (i.e., merge values into a single, non-delimited string), freqdesc (i.e., print desc. list of values:freq) freqasc (i.e., print asc. list of values:freq) first (i.e., print first value) last (i.e., print last value) Default value: sum If there is only column, but multiple operations, all operations will be applied on that column. Likewise, if there is only one operation, but multiple columns, that operation will be applied to all columns. Otherwise, the number of columns must match the the number of operations, and will be applied in respective order. E.g., "-c 5,4,6 -o sum,mean,count" will give the sum of column 5, the mean of column 4, and the count of column 6. The order of output columns will match the ordering given in the command. |
--full | boolean_true | Print all columns from input file. The first line in the group is used. Default: print only grouped columns. |
--inheader | boolean_true | Input file has a header line - the first line will be ignored. |
--outheader | boolean_true | Print header line in the output, detailing the column names. If the input file has headers (-inheader), the output file will use the input's column names. If the input file has no headers, the output file will use "col_1", "col_2", etc. as the column names. |
--header | boolean_true | same as '-inheader -outheader'. |
--ignorecase | boolean_true | Group values regardless of upper/lower case. |
--precision -prec | integer | Sets the decimal precision for output. |
--delimiter -delim | string | Specify a custom delimiter for the collapse operations. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output.bed"
precision: [ 5 ]
delimiter: [ "," ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.3.0 \
-main-script target/nextflow/bedtools/bedtools_groupby/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
bedtools/bedtools_groupbybiobox v0.3.0
Uses
0 relationships
No component dependencies found.