umi_tools/umi_tools_extract
extract
umi-tools
umi
fastq
Description
Flexible removal of UMI sequences from fastq reads.
UMIs are removed and appended to the read name. Any other barcode, for example a library barcode,
is left on the read. Can also filter reads by quality or against a whitelist.
Input
Name | Type & Properties | Description |
|---|---|---|
--input | file required | File containing the input data. |
--read2_in | file | File containing the input data for the R2 reads (if paired). If provided, a <list of other required arguments> need to be provided. |
--bc_pattern -p | string | The UMI barcode pattern to use e.g. 'NNNNNN' indicates that the first 6 nucleotides of the read are from the UMI. |
--bc_pattern2 | string | The UMI barcode pattern to use for read 2. |
Output
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output file for read 1. |
--read2_out | file output | Output file for read 2. |
--filtered_out | file | Write out reads not matching regex pattern or cell barcode whitelist to this file. |
--filtered_out2 | file | Write out read pairs not matching regex pattern or cell barcode whitelist to this file. |
Extract Options
Name | Type & Properties | Description |
|---|---|---|
--extract_method | string | UMI pattern to use. Default: `string`. |
--error_correct_cell | boolean_true | Error correct cell barcodes to the whitelist. |
--whitelist | file | Whitelist of accepted cell barcodes tab-separated format, where column 1 is the whitelisted cell barcodes and column 2 is the list (comma-separated) of other cell barcodes which should be corrected to the barcode in column 1. If the --error_correct_cell option is not used, this column will be ignored. |
--blacklist | file | BlackWhitelist of cell barcodes to discard. |
--subset_reads | integer | Only parse the first N reads. |
--quality_filter_threshold | integer | Remove reads where any UMI base quality score falls below this threshold. |
--quality_filter_mask | string | If a UMI base has a quality below this threshold, replace the base with 'N'. |
--quality_encoding | string | Quality score encoding. Choose from: * phred33 [33-77] * phred64 [64-106] * solexa [59-106] |
--reconcile_pairs | boolean_true | Allow read 2 infile to contain reads not in read 1 infile. This enables support for upstream protocols where read one contains cell barcodes, and the read pairs have been filtered and corrected without regard to the read2. |
--three_prime --3prime | boolean_true | By default the barcode is assumed to be on the 5' end of the read, but use this option to sepecify that it is on the 3' end instead. This option only works with --extract_method=string since 3' encoding can be specified explicitly with a regex, e.g `.*(?P<umi_1>.{5})$`. |
--ignore_read_pair_suffixes | boolean_true | Ignore "/1" and "/2" read name suffixes. Note that this options is required if the suffixes are not whitespace separated from the rest of the read name. arguments: |
--umi_separator | string | The character that separates the UMI in the read name. Most likely a colon if you skipped the extraction with UMI-tools and used other software. Default: `_` |
--grouping_method | string | Method to use to determine read groups by subsuming those with similar UMIs. All methods start by identifying the reads with the same mapping position, but treat similar yet nonidentical UMIs differently. Default: `directional` |
Common Options
Name | Type & Properties | Description |
|---|---|---|
--log | file output | File with logging information. |
--log2stderr | boolean_true | Send logging information to stderr. |
--verbose | integer | Log level. The higher, the more output. |
--error | file output | File with error information. |
--temp_dir | string | Directory for temporary files. If not set, the bash environmental variable TMPDIR is used. |
--compresslevel | integer | Level of Gzip compression to use. Default=6 matches GNU gzip rather than python gzip default (which is 9). Default `6`. |
--timeit | file output | Store timing information in file. |
--timeit_name | string | Name in timing file for this class of jobs. |
--timeit_header | boolean_true | Add header for timing information. |
--random_seed | integer | Random seed to initialize number generator with. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output"
read2_out: "$id.$key.read2_out"
log: "$id.$key.log"
error: "$id.$key.error"
timeit: "$id.$key.timeit"
timeit_name: [ "all" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.1 \
-main-script target/nextflow/umi_tools/umi_tools_extract/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
umi_tools/umi_tools_extractbiobox v0.4.1
Uses
0 relationships
No component dependencies found.