fastqc
Quality control
BAM
SAM
FASTQ
Description
FastQC - A high throughput sequence QC analysis tool.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required multiple | FASTQ file(s) to be analyzed. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--outdir | file output | Output directory where the results will be saved. |
--html | file multiple output | Create the HTML report of the results. '*' wild card must be provided in the output file name. Wild card will be replaced by the input file basename. e.g. --input "sample_1.fq" --html "*.html" would create an output html file named sample_1.html |
--zip | file multiple output | Create the zip file(s) containing: html report, data, images, icons, summary, etc. '*' wild card must be provided in the output file name. Wild card will be replaced by the input basename. e.g. --input "sample_1.fq" --html "*.zip" would create an output zip file named sample_1.zip |
--summary | file multiple output | Create the summary file(s). '*' wild card must be provided in the output file name. Wild card will be replaced by the input basename. e.g. --input "sample_1.fq" --summary "*_summary.txt" would create an output summary.txt file named sample_1_summary.txt |
--data | file multiple output | Create the data file(s). '*' wild card must be provided in the output file name. Wild card will be replaced by the input basename. e.g. --input "sample_1.fq" --summary "*_data.txt" would create an output data.txt file named sample_1_data.txt |
Options
Name | Type & Properties | Description |
|---|---|---|
--casava | boolean_true | Files come from raw casava output. Files in the same sample group (differing only by the group number) will be analysed as a set rather than individually. Sequences with the filter flag set in the header will be excluded from the analysis. Files must have the same names given to them by casava (including being gzipped and ending with .gz) otherwise they won't be grouped together correctly. |
--nano | boolean_true | Files come from nanopore sequences and are in fast5 format. In this mode you can pass in directories to process and the program will take in all fast5 files within those directories and produce a single output file from the sequences found in all files. |
--nofilter | boolean_true | If running with --casava then don't remove read flagged by casava as poor quality when performing the QC analysis. |
--nogroup | boolean_true | Disable grouping of bases for reads >50bp. All reports will show data for every base in the read. WARNING: Using this option will cause fastqc to crash and burn if you use it on really long reads, and your plots may end up a ridiculous size. You have been warned! |
--min_length | integer | Sets an artificial lower limit on the length of the sequence to be shown in the report. As long as you set this to a value greater or equal to your longest read length then this will be the sequence length used to create your read groups. This can be useful for making directly comparable statistics from datasets with somewhat variable read lengths. |
--format -f | string | Bypasses the normal sequence file format detection and forces the program to use the specified format. Valid formats are bam, sam, bam_mapped, sam_mapped, and fastq. |
--contaminants -c | file | Specifies a non-default file which contains the list of contaminants to screen overrepresented sequences against. The file must contain sets of named contaminants in the form name[tab]sequence. Lines prefixed with a hash will be ignored. |
--adapters -a | file | Specifies a non-default file which contains the list of adapter sequences which will be explicitly searched against the library. The file must contain sets of named adapters in the form name[tab]sequence. Lines prefixed with a hash will be ignored. |
--limits -l | file | Specifies a non-default file which contains a set of criteria which will be used to determine the warn/error limits for the various modules. This file can also be used to selectively remove some modules from the output altogether. The format needs to mirror the default limits.txt file found in the Configuration folder. |
--kmers -k | integer | Specifies the length of Kmer to look for in the Kmer content module. Specified Kmer length must be between 2 and 10. Default length is 7 if not specified. |
--quiet -q | boolean_true | Suppress all progress messages on stdout and only report errors. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
outdir: "$id.$key.outdir.results"
html: "$id.$key.html._*.html"
zip: "$id.$key.zip._*.zip"
summary: "$id.$key.summary._*.txt"
data: "$id.$key.data._*.txt"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.1 \
-main-script target/nextflow/fastqc/main.nf \
-params-file params.yaml Relationships
Used by
3 relationships, 1 components
Current component
fastqcbiobox v0.4.1
Uses
0 relationships
No component dependencies found.