fastp
Description
An ultra-fast all-in-one FASTQ preprocessor (QC/adapters/trimming/filtering/splitting/merging...).
Features:
comprehensive quality profiling for both before and after filtering data (quality curves, base contents, KMER, Q20/Q30, GC Ratio, duplication, adapter contents...)
filter out bad reads (too low quality, too short, or too many N...)
cut low quality bases for per read in its 5' and 3' by evaluating the mean quality from a sliding window (like Trimmomatic but faster).
trim all reads in front and tail
cut adapters. Adapter sequences can be automatically detected, which means you don't have to input the adapter sequences to trim them.
correct mismatched base pairs in overlapped regions of paired end reads, if one base is with high quality while the other is with ultra low quality
trim polyG in 3' ends, which is commonly seen in NovaSeq/NextSeq data. Trim polyX in 3' ends to remove unwanted polyX tailing (i.e. polyA tailing for mRNA-Seq data)
preprocess unique molecular identifier (UMI) enabled data, shift UMI to sequence name.
report JSON format result for further interpreting.
visualize quality control and filtering results on a single HTML page (like FASTQC but faster and more informative).
split the output to multiple files (0001.R1.gz, 0002.R1.gz...) to support parallel processing. Two modes can be used, limiting the total split file number, or limitting the lines of each split file.
support long reads (data from PacBio / Nanopore devices).
support reading from STDIN and writing to STDOUT
support interleaved input
support ultra-fast FASTQ-level deduplication
Inputs
Name | Type & Properties | Description |
|---|---|---|
--in1 -i | file required | Input FastQ file. Must be single-end or paired-end R1. Can be gzipped. |
--in2 -I | file | Input FastQ file. Must be paired-end R2. Can be gzipped. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--out1 -o | file required output | The single-end or paired-end R1 reads that pass QC. Will be gzipped if its file name ends with `.gz`. |
--out2 -O | file output | The paired-end R2 reads that pass QC. Will be gzipped if its file name ends with `.gz`. |
--unpaired1 | file output | Store the reads that `read1` passes filters but its paired `read2` doesn't. |
--unpaired2 | file output | Store the reads that `read2` passes filters but its paired `read1` doesn't. |
--failed_out | file output | Store the reads that fail filters. If one read failed and is written to --failed_out, its failure reason will be appended to its read name. For example, failed_quality_filter, failed_too_short etc. For PE data, if unpaired reads are not stored (by giving --unpaired1 or --unpaired2), the failed pair of reads will be put together. If one read passes the filters but its pair doesn't, the failure reason will be paired_read_is_failing. |
--overlapped_out | file output | For each read pair, output the overlapped region if it has no any mismatched base. |
Report output arguments
Name | Type & Properties | Description |
|---|---|---|
--json -j | file output | The json format report file name |
--html | file output | The html format report file name |
--report_title | string | The title of the html report, default is "fastp report". |
Adapter trimming
Name | Type & Properties | Description |
|---|---|---|
--disable_adapter_trimming -A | boolean_true | Disable adapter trimming. |
--detect_adapter_for_pe | boolean_true | By default, the auto-detection for adapter is for SE data input only, turn on this option to enable it for PE data. |
--adapter_sequence -a | string | The adapter sequences to be trimmed. For SE data, if not specified, the adapters will be auto-detected. For PE data, this is used if R1/R2 are found not overlapped |
--adapter_sequence_r2 | string | The adapter sequences to be trimmed for R2. This is used for PE data if R1/R2 are found overlapped. |
--adapter_fasta | file | A FASTA file containing all the adapter sequences to be trimmed. For SE data, if not specified, the adapters will be auto-detected. For PE data, this is used if R1/R2 are found not overlapped. |
Base trimming
Name | Type & Properties | Description |
|---|---|---|
--trim_front1 -f | integer | Trimming how many bases in front for read1, default is 0. |
--trim_tail1 -t | integer | Trimming how many bases in tail for read1, default is 0. |
--max_len1 -b | integer | If read1 is longer than max_len1, then trim read1 at its tail to make it as long as max_len1. Default 0 means no limitation. |
--trim_front2 -F | integer | Trimming how many bases in front for read2, default is 0. |
--trim_tail2 -T | integer | Trimming how many bases in tail for read2, default is 0. |
--max_len2 -B | integer | If read2 is longer than max_len2, then trim read2 at its tail to make it as long as max_len2. Default 0 means no limitation. |
Merging mode
Name | Type & Properties | Description |
|---|---|---|
--merge -m | boolean_true | For paired-end input, merge each pair of reads into a single read if they are overlapped. The merged reads will be written to the file given by --merged_out, the unmerged reads will be written to the files specified by --out1 and --out2. The merging mode is disabled by default. |
--merged_out | file output | In the merging mode, specify the file name to store merged output, or specify --stdout to stream the merged output. |
--include_unmerged | boolean_true | In the merging mode, write the unmerged or unpaired reads to the file specified by --merge. Disabled by default. |
Additional input arguments
Name | Type & Properties | Description |
|---|---|---|
--interleaved_in | boolean_true | Indicate that <in1> is an interleaved FASTQ which contains both read1 and read2. Disabled by default. |
--fix_mgi_id | boolean_true | The MGI FASTQ ID format is not compatible with many BAM operation tools, enable this option to fix it. |
--phred64 -6 | boolean_true | Indicate the input is using phred64 scoring (it'll be converted to phred33, so the output will still be phred33) |
Additional output arguments
Name | Type & Properties | Description |
|---|---|---|
--compression -z | integer | Compression level for gzip output (1 ~ 9). 1 is fastest, 9 is smallest, default is 4. |
--dont_overwrite | boolean_true | Don't overwrite existing files. Overwritting is allowed by default. |
Logging arguments
Name | Type & Properties | Description |
|---|---|---|
--verbose -V | boolean_true | Output verbose log information (i.e. when every 1M reads are processed). |
Processing arguments
Name | Type & Properties | Description |
|---|---|---|
--reads_to_process | long | Specify how many reads/pairs to be processed. Default 0 means process all reads. |
Deduplication arguments
Name | Type & Properties | Description |
|---|---|---|
--dedup | boolean_true | Enable deduplication to drop the duplicated reads/pairs |
--dup_calc_accuracy | integer | Accuracy level to calculate duplication (1~6). Higher level uses more memory (1G, 2G, 4G, 8G, 16G, 24G). Default 1 for no-dedup mode, and 3 for dedup mode. |
--dont_eval_duplication | boolean_true | Don't evaluate duplication rate to save time and use less memory. |
PolyG tail trimming arguments
Name | Type & Properties | Description |
|---|---|---|
--trim_poly_g -g | boolean_true | Force polyG tail trimming, by default trimming is automatically enabled for Illumina NextSeq/NovaSeq data |
--poly_g_min_len | integer | The minimum length to detect polyG in the read tail. 10 by default. |
--disable_trim_poly_g -G | boolean_true | Disable polyG tail trimming, by default trimming is automatically enabled for Illumina NextSeq/NovaSeq data |
PolyX tail trimming arguments
Name | Type & Properties | Description |
|---|---|---|
--trim_poly_x -x | boolean_true | Enable polyX trimming in 3' ends. |
--poly_x_min_len | integer | The minimum length to detect polyX in the read tail. 10 by default. |
Cut arguments
Name | Type & Properties | Description |
|---|---|---|
--cut_front -5 | integer | Move a sliding window from front (5') to tail, drop the bases in the window if its mean quality < threshold, stop otherwise. |
--cut_tail -3 | integer | Move a sliding window from tail (3') to front, drop the bases in the window if its mean quality < threshold, stop otherwise. |
--cut_right -r | integer | Move a sliding window from front to tail, if meet one window with mean quality < threshold, drop the bases in the window and the right part, and then stop. |
--cut_window_size -W | integer | The window size option shared by cut_front, cut_tail or cut_sliding. Range: 1~1000, default: 4. |
--cut_mean_quality -M | integer | The mean quality requirement option shared by cut_front, cut_tail or cut_sliding. Range: 1~36 default: 20 (Q20) |
--cut_front_window_size | integer | The window size option of cut_front, default to cut_window_size if not specified. |
--cut_front_mean_quality | integer | The mean quality requirement option of cut_front, default to cut_mean_quality if not specified. |
--cut_tail_window_size | integer | The window size option of cut_tail, default to cut_window_size if not specified. |
--cut_tail_mean_quality | integer | The mean quality requirement option of cut_tail, default to cut_mean_quality if not specified. |
--cut_right_window_size | integer | The window size option of cut_right, default to cut_window_size if not specified. |
--cut_right_mean_quality | integer | The mean quality requirement option of cut_right, default to cut_mean_quality if not specified. |
Quality filtering arguments
Name | Type & Properties | Description |
|---|---|---|
--disable_quality_filtering -Q | boolean_true | Quality filtering is enabled by default. If this option is specified, quality filtering is disabled. |
--qualified_quality_phred -q | integer | The quality value that a base is qualified. Default 15 means phred quality >=Q15 is qualified. |
--unqualified_percent_limit -u | integer | How many percents of bases are allowed to be unqualified (0~100). Default 40 means 40%. |
--n_base_limit -n | integer | If one read's number of N base is >n_base_limit, then this read/pair is discarded. Default is 5. |
--average_qual -e | integer | If one read's average quality score <avg_qual, then this read/pair is discarded. Default 0 means no requirement. |
Length filtering arguments
Name | Type & Properties | Description |
|---|---|---|
--disable_length_filtering -L | boolean_true | Length filtering is enabled by default. If this option is specified, length filtering is disabled. |
--length_required -l | integer | Reads shorter than length_required will be discarded, default is 15. |
--length_limit | integer | Reads longer than length_limit will be discarded, default 0 means no limitation. |
Low complexity filtering arguments
Name | Type & Properties | Description |
|---|---|---|
--low_complexity_filter -y | boolean_true | Enable low complexity filter. The complexity is defined as the percentage of base that is different from its next base (base[i] != base[i+1]). |
--complexity_threshold -Y | integer | The threshold for low complexity filter (0~100). Default is 30, which means 30% complexity is required. |
Index filtering arguments
Name | Type & Properties | Description |
|---|---|---|
--filter_by_index1 | file | Specify a file contains a list of barcodes of index1 to be filtered out, one barcode per line. |
--filter_by_index2 | file | Specify a file contains a list of barcodes of index2 to be filtered out, one barcode per line. |
--filter_by_index_threshold | integer | The allowed difference of index barcode for index filtering, default 0 means completely identical. |
Overlapped region correction
Name | Type & Properties | Description |
|---|---|---|
--correction -c | boolean_true | Enable base correction in overlapped regions (only for PE data), default is disabled. |
--overlap_len_require | integer | The minimum length to detect overlapped region of PE reads. This will affect overlap analysis based PE merge, adapter trimming and correction. 30 by default. |
--overlap_diff_limit | integer | The maximum number of mismatched bases to detect overlapped region of PE reads. This will affect overlap analysis based PE merge, adapter trimming and correction. 5 by default. |
--overlap_diff_percent_limit | integer | The maximum percentage of mismatched bases to detect overlapped region of PE reads. This will affect overlap analysis based PE merge, adapter trimming and correction. Default 20 means 20%. |
UMI arguments
Name | Type & Properties | Description |
|---|---|---|
--umi -U | boolean_true | Enable unique molecular identifier (UMI) preprocessing. |
--umi_loc | string | Specify the location of UMI, can be (index1/index2/read1/read2/per_index/per_read, default is none. |
--umi_len | integer | If the UMI is in read1/read2, its length should be provided. |
--umi_prefix | string | If specified, an underline will be used to connect prefix and UMI (i.e. prefix=UMI, UMI=AATTCG, final=UMI_AATTCG). No prefix by default. |
--umi_skip | integer | If the UMI is in read1/read2, fastp can skip several bases following UMI, default is 0. |
--umi_delim | string | If the UMI is in index1/index2, fastp can use a delimiter to separate UMI from the read sequence, default is none. |
Overrepresentation analysis arguments
Name | Type & Properties | Description |
|---|---|---|
--overrepresentation_analysis -p | boolean_true | Enable overrepresentation analysis. |
--overrepresentation_sampling | integer | One in (--overrepresentation_sampling) reads will be computed for overrepresentation analysis (1~10000), smaller is slower, default is 20. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
out1: "$id.$key.out1.gz"
out2: "$id.$key.out2.gz"
unpaired1: "$id.$key.unpaired1.gz"
unpaired2: "$id.$key.unpaired2.gz"
failed_out: "$id.$key.failed_out.gz"
overlapped_out: "$id.$key.overlapped_out"
json: "$id.$key.json.json"
html: "$id.$key.html.html"
merged_out: "$id.$key.merged_out.gz"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.3.0 \
-main-script target/nextflow/fastp/main.nf \
-params-file params.yaml Relationships
Used by
No components use this component.
Current component
Uses
No component dependencies found.