sortmerna
sort
mRNA
rRNA
alignment
filtering
mapping
clustering
Description
Local sequence alignment tool for filtering, mapping and clustering. The main
application of SortMeRNA is filtering rRNA from metatranscriptomic data.
Input
Name | Type & Properties | Description |
|---|---|---|
--paired | boolean_true | Reads are paired-end. If a single reads file is provided, use this option to indicate the file contains interleaved paired reads when neither 'paired_in' | 'paired_out' | 'out2' | 'sout' are specified. |
--input | file multiple | Input fastq |
--ref | file multiple | Reference fasta file(s) for rRNA database. |
--ribo_database_manifest | file | Text file containing paths to fasta files (one per line) that will be used to create the database for SortMeRNA. |
Output
Name | Type & Properties | Description |
|---|---|---|
--log | file output | Sortmerna log file. |
--output --aligned | file output | Directory and file prefix for aligned output. The appropriate extension: (fasta|fastq|blast|sam|etc) is automatically added. If 'dir' is not specified, the output is created in the WORKDIR/out/. If 'pfx' is not specified, the prefix 'aligned' is used. |
--other | file output | Create Non-aligned reads output file with this path/prefix. Must be used with fastx. |
Options
Name | Type & Properties | Description |
|---|---|---|
--kvdb | string | Path to directory of the key-value database file, used for storing the alignment results. |
--idx_dir | string | Path to the directory for storing the reference index files. |
--readb | string | Path to the directory for storing pre-processed reads. |
--fastx | boolean_true | Output aligned reads into FASTA/FASTQ file |
--sam | boolean_true | Output SAM alignment for aligned reads. |
--sq | boolean_true | Add SQ tags to the SAM file |
--blast | string | Blast options: * '0' - pairwise * '1' - tabular(Blast - m 8 format) * '1 cigar' - tabular + column for CIGAR * '1 cigar qcov' - tabular + columns for CIGAR and query coverage * '1 cigar qcov qstrand' - tabular + columns for CIGAR, query coverage and strand |
--num_alignments | integer | Report first INT alignments per read reaching E-value. If Int = 0, all alignments will be output. Default: '0' |
--min_lis | integer | search all alignments having the first INT longest LIS. LIS stands for Longest Increasing Subsequence, it is computed using seeds' positions to expand hits into longer matches prior to Smith-Waterman alignment. Default: '2'. |
--print_all_reads | boolean_true | output null alignment strings for non-aligned reads to SAM and/or BLAST tabular files. |
--paired_in | boolean_true | In the case where a pair of reads is aligned with a score above the threshold, the output of the reads is controlled by the following options: * --paired_in and --paired_out are both false: Only one read per pair is output to the aligned fasta file. * --paired_in is true and --paired_out is false: Both reads of the pair are output to the aligned fasta file. * --paired_in is false and --paired_out is true: Both reads are output the the other fasta file (if it is specified). |
--paired_out | boolean_true | See description of --paired_in. |
--out2 | boolean_true | Output paired reads into separate files. Must be used with '--fastx'. If a single reads file is provided, this options implies interleaved paired reads. When used with 'sout', four (4) output files for aligned reads will be generated: 'aligned-paired-fwd, aligned-paired-rev, aligned-singleton-fwd, aligned-singleton-rev'. If 'other' option is also used, eight (8) output files will be generated. |
--sout | boolean_true | Separate paired and singleton aligned reads. Must be used with '--fastx'. If a single reads file is provided, this options implies interleaved paired reads. Cannot be used with '--paired_in' or '--paired_out'. |
--zip_out | string | Compress the output files. The possible values are: * '1/true/t/yes/y' * '0/false/f/no/n' *'-1' (the same format as input - default) The values are Not case sensitive. |
--match | integer | Smith-Waterman score for a match (positive integer). Default: '2'. |
--mismatch | integer | Smith-Waterman penalty for a mismatch (negative integer). Default: '-3'. |
--gap_open | integer | Smith-Waterman penalty for introducing a gap (positive integer). Default: '5'. |
--gap_ext | integer | Smith-Waterman penalty for extending a gap (positive integer). Default: '2'. |
--N | integer | Smith-Waterman penalty for ambiguous letters (N's) scored as --mismatch. Default: '-1'. |
--a | integer | Number of threads to use. Default: '1'. |
--e | double | E-value threshold. Default: '1'. |
--F | boolean_true | Search only the forward strand. |
--R | boolean_true | Search only the reverse-complementary strand. |
--num_alignment | integer | Report first INT alignments per read reaching E-value (--num_alignments 0 signifies all alignments will be output). Default: '-1' |
--best | integer | Report INT best alignments per read reaching E-value by searching --min_lis INT candidate alignments (--best 0 signifies all candidate alignments will be searched) Default: '1'. |
--verbose -v | boolean_true | Verbose output. |
OTU picking options
Name | Type & Properties | Description |
|---|---|---|
--id | double | %id similarity threshold (the alignment must still pass the E-value threshold). Default: '0.97'. |
--coverage | double | %query coverage threshold (the alignment must still pass the E-value threshold). Default: '0.97'. |
--de_novo | boolean_true | FASTA/FASTQ file for reads matching database < %id off (set using --id) and < %cov (set using --coverage) (alignment must still pass the E-value threshold). |
--otu_map | boolean_true | Output OTU map (input to QIIME's make_otu_table.py). |
Advanced options
Name | Type & Properties | Description |
|---|---|---|
--num_seed | integer | Number of seeds matched before searching for candidate LIS. Default: '2'. |
--passes | integer multiple | Three intervals at which to place the seed on the read L,L/2,3 (L is the seed length set in ./indexdb_rna). |
--edge | string | The number (or percentage if followed by %) of nucleotides to add to each edge of the alignment region on the reference sequence before performing Smith-Waterman alignment. Default: '4'. |
--full_search | boolean_true | Search for all 0-error and 1-error seed off matches in the index rather than stopping after finding a 0-error match (<1% gain in sensitivity with up four-fold decrease in speed). |
Indexing Options
Name | Type & Properties | Description |
|---|---|---|
--index | integer | Create index files for the reference database. By default when this option is not used, the program checks the reference index and builds it if not already existing. This can be changed by using '-index' as follows: * '-index 0' - skip indexing. If the index does not exist, the program will terminate and warn to build the index prior performing the alignment * '-index 1' - only perform the indexing and terminate * '-index 2' - the default behaviour, the same as when not using this option at all |
-L | double | Indexing seed length. Default: '18' |
--interval | integer | Index every Nth L-mer in the reference database. Default: '1' |
--max_pos | integer | Maximum number of positions to store for each unique L-mer. Set to 0 to store all positions. Default: '1000' |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
log: "$id.$key.log.log"
output: "$id.$key.output"
other: "$id.$key.other"
id: .nan
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.0 \
-main-script target/nextflow/sortmerna/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
sortmernabiobox v0.4.0
Uses
0 relationships
No component dependencies found.