cutadapt
RNA-seq
scRNA-seq
high-throughput
Description
Cutadapt removes adapter sequences from high-throughput sequencing reads.
Specify Adapters for R1
Name | Type & Properties | Description |
|---|---|---|
--adapter -a | string multiple | Sequence of an adapter ligated to the 3' end (paired data: of the first read). The adapter and subsequent bases are trimmed. If a '$' character is appended ('anchoring'), the adapter is only found if it is a suffix of the read. |
--front -g | string multiple | Sequence of an adapter ligated to the 5' end (paired data: of the first read). The adapter and any preceding bases are trimmed. Partial matches at the 5' end are allowed. If a '^' character is prepended ('anchoring'), the adapter is only found if it is a prefix of the read. |
--anywhere -b | string multiple | Sequence of an adapter that may be ligated to the 5' or 3' end (paired data: of the first read). Both types of matches as described under -a and -g are allowed. If the first base of the read is part of the match, the behavior is as with -g, otherwise as with -a. This option is mostly for rescuing failed library preparations - do not use if you know which end your adapter was ligated to! |
Specify Adapters using Fasta files for R1
Name | Type & Properties | Description |
|---|---|---|
--adapter_fasta | file multiple | Fasta file containing sequences of an adapter ligated to the 3' end (paired data: of the first read). The adapter and subsequent bases are trimmed. If a '$' character is appended ('anchoring'), the adapter is only found if it is a suffix of the read. |
--front_fasta | file | Fasta file containing sequences of an adapter ligated to the 5' end (paired data: of the first read). The adapter and any preceding bases are trimmed. Partial matches at the 5' end are allowed. If a '^' character is prepended ('anchoring'), the adapter is only found if it is a prefix of the read. |
--anywhere_fasta | file | Fasta file containing sequences of an adapter that may be ligated to the 5' or 3' end (paired data: of the first read). Both types of matches as described under -a and -g are allowed. If the first base of the read is part of the match, the behavior is as with -g, otherwise as with -a. This option is mostly for rescuing failed library preparations - do not use if you know which end your adapter was ligated to! |
Specify Adapters for R2
Name | Type & Properties | Description |
|---|---|---|
--adapter_r2 -A | string multiple | Sequence of an adapter ligated to the 3' end (paired data: of the first read). The adapter and subsequent bases are trimmed. If a '$' character is appended ('anchoring'), the adapter is only found if it is a suffix of the read. |
--front_r2 -G | string multiple | Sequence of an adapter ligated to the 5' end (paired data: of the first read). The adapter and any preceding bases are trimmed. Partial matches at the 5' end are allowed. If a '^' character is prepended ('anchoring'), the adapter is only found if it is a prefix of the read. |
--anywhere_r2 -B | string multiple | Sequence of an adapter that may be ligated to the 5' or 3' end (paired data: of the first read). Both types of matches as described under -a and -g are allowed. If the first base of the read is part of the match, the behavior is as with -g, otherwise as with -a. This option is mostly for rescuing failed library preparations - do not use if you know which end your adapter was ligated to! |
Specify Adapters using Fasta files for R2
Name | Type & Properties | Description |
|---|---|---|
--adapter_r2_fasta | file | Fasta file containing sequences of an adapter ligated to the 3' end (paired data: of the first read). The adapter and subsequent bases are trimmed. If a '$' character is appended ('anchoring'), the adapter is only found if it is a suffix of the read. |
--front_r2_fasta | file | Fasta file containing sequences of an adapter ligated to the 5' end (paired data: of the first read). The adapter and any preceding bases are trimmed. Partial matches at the 5' end are allowed. If a '^' character is prepended ('anchoring'), the adapter is only found if it is a prefix of the read. |
--anywhere_r2_fasta | file | Fasta file containing sequences of an adapter that may be ligated to the 5' or 3' end (paired data: of the first read). Both types of matches as described under -a and -g are allowed. If the first base of the read is part of the match, the behavior is as with -g, otherwise as with -a. This option is mostly for rescuing failed library preparations - do not use if you know which end your adapter was ligated to! |
Paired-end options
Name | Type & Properties | Description |
|---|---|---|
--pair_adapters | boolean_true | Treat adapters given with -a/-A etc. as pairs. Either both or none are removed from each read pair. |
--pair_filter | string | Which of the reads in a paired-end read have to match the filtering criterion in order for the pair to be filtered. |
--interleaved | boolean_true | Read and/or write interleaved paired-end reads. |
Input parameters
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input fastq file for single-end reads or R1 for paired-end reads. |
--input_r2 | file | Input fastq file for R2 in the case of paired-end reads. |
--error_rate -E --errors | double | Maximum allowed error rate (if 0 <= E < 1), or absolute number of errors for full-length adapter match (if E is an integer >= 1). Error rate = no. of errors divided by length of matching region. Default: 0.1 (10%). |
--no_indels | boolean_true | Allow only mismatches in alignments. |
--times -n | integer | Remove up to COUNT adapters from each read. Default: 1. |
--overlap -O | integer | Require MINLENGTH overlap between read and adapter for an adapter to be found. The default is 3. |
--match_read_wildcards | boolean_true | Interpret IUPAC wildcards in reads. |
--no_match_adapter_wildcards | boolean_true | Do not interpret IUPAC wildcards in adapters. |
--action | string | What to do if a match was found. trim: trim adapter and up- or downstream sequence; retain: trim, but retain adapter; mask: replace with 'N' characters; lowercase: convert to lowercase; none: leave unchanged. The default is trim. |
--revcomp --rc | boolean_true | Check both the read and its reverse complement for adapter matches. If match is on reverse-complemented version, output that one. |
Demultiplexing options
Name | Type & Properties | Description |
|---|---|---|
--demultiplex_mode | string | Enable demultiplexing and set the mode for it. With mode 'unique_dual', adapters from the first and second read are used, and the indexes from the reads are only used in pairs. This implies --pair_adapters. Enabling mode 'combinatorial_dual' allows all combinations of the sets of indexes on R1 and R2. It is necessary to write each read pair to an output file depending on the adapters found on both R1 and R2. Mode 'single', uses indexes or barcodes located at the 5' end of the R1 read (single). |
Read modifications
Name | Type & Properties | Description |
|---|---|---|
--cut -u | integer multiple | Remove LEN bases from each read (or R1 if paired; use --cut_r2 option for R2). If LEN is positive, remove bases from the beginning. If LEN is negative, remove bases from the end. Can be used twice if LENs have different signs. Applied *before* adapter trimming. |
--cut_r2 | integer multiple | Remove LEN bases from each read (for R2). If LEN is positive, remove bases from the beginning. If LEN is negative, remove bases from the end. Can be used twice if LENs have different signs. Applied *before* adapter trimming. |
--nextseq_trim | string | NextSeq-specific quality trimming (each read). Trims also dark cycles appearing as high-quality G bases. |
--quality_cutoff -q | string | Trim low-quality bases from 5' and/or 3' ends of each read before adapter removal. Applied to both reads if data is paired. If one value is given, only the 3' end is trimmed. If two comma-separated cutoffs are given, the 5' end is trimmed with the first cutoff, the 3' end with the second. |
--quality_cutoff_r2 -Q | string | Quality-trimming cutoff for R2. Default: same as for R1 |
--quality_base | integer | Assume that quality values in FASTQ are encoded as ascii(quality + N). This needs to be set to 64 for some old Illumina FASTQ files. The default is 33. |
--poly_a | boolean_true | Trim poly-A tails |
--length -l | integer | Shorten reads to LENGTH. Positive values remove bases at the end while negative ones remove bases at the beginning. This and the following modifications are applied after adapter trimming. |
--trim_n | boolean_true | Trim N's on ends of reads. |
--length_tag | string | Search for TAG followed by a decimal number in the description field of the read. Replace the decimal number with the correct length of the trimmed read. For example, use --length-tag 'length=' to correct fields like 'length=123'. |
--strip_suffix | string | Remove this suffix from read names if present. Can be given multiple times. |
--prefix -x | string | Add this prefix to read names. Use {name} to insert the name of the matching adapter. |
--suffix -y | string | Add this suffix to read names; can also include {name} |
--rename | string | Rename reads using TEMPLATE containing variables such as {id}, {adapter_name} etc. (see documentation) |
--zero_cap -z | boolean_true | Change negative quality values to zero. |
Filtering of processed reads
Name | Type & Properties | Description |
|---|---|---|
--minimum_length -m | string | Discard reads shorter than LEN. Default is 0. When trimming paired-end reads, the minimum lengths for R1 and R2 can be specified separately by separating them with a colon (:). If the colon syntax is not used, the same minimum length applies to both reads, as discussed above. Also, one of the values can be omitted to impose no restrictions. For example, with -m 17:, the length of R1 must be at least 17, but the length of R2 is ignored. |
--maximum_length -M | string | Discard reads longer than LEN. Default: no limit. For paired reads, see the remark for --minimum_length |
--max_n | string | Discard reads with more than COUNT 'N' bases. If COUNT is a number between 0 and 1, it is interpreted as a fraction of the read length. |
--max_expected_errors --max_ee | long | Discard reads whose expected number of errors (computed from quality values) exceeds ERRORS. |
--max_average_error_rate --max_aer | long | as --max_expected_errors (see above), but divided by length to account for reads of varying length. |
--discard_trimmed --discard | boolean_true | Discard reads that contain an adapter. Use also -O to avoid discarding too many randomly matching reads. |
--discard_untrimmed --trimmed_only | boolean_true | Discard reads that do not contain an adapter. |
--discard_casava | boolean_true | Discard reads that did not pass CASAVA filtering (header has :Y:). |
Output parameters
Name | Type & Properties | Description |
|---|---|---|
--report | string | Which type of report to print: 'full' (default) or 'minimal'. |
--json | boolean_true | Write report in JSON format to this file. |
--output | file required multiple output | Glob pattern for matching the expected output files. Should include `$output_dir`. |
--fasta | boolean_true | Output FASTA to standard output even on FASTQ input. |
--info_file | boolean_true | Write information about each read and its adapter matches into info.txt in the output directory. See the documentation for the file format. |
Debug
Name | Type & Properties | Description |
|---|---|---|
--debug | boolean_true | Print debug information |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output._*.fast[a,q]"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.3.1 \
-main-script target/nextflow/cutadapt/main.nf \
-params-file params.yaml Relationships
Used by
12 relationships, 1 components
Current component
cutadaptbiobox v0.3.1
Uses
0 relationships
No component dependencies found.