sgdemux
demultiplex
fastq
Description
Demultiplex sequence data generated on Singular Genomics' sequencing instruments.
Input
Name | Type & Properties | Description |
|---|---|---|
--fastqs -f | file required multiple | Path to the input FASTQs, or path prefix if not a file |
--sample_metadata -s | file required | Path to the sample metadata CSV file including sample names and barcode sequences |
Output
Name | Type & Properties | Description |
|---|---|---|
--sample_fastq | file required output | The directory containing demultiplexed sample FASTQ files. |
--metrics | file output | Demultiplexing summary statisitcs: - control_reads_omitted: The number of reads that were omitted for being control reads. - failing_reads_omitted: The number of reads that were omitted for having failed QC. - total_templates: The total number of template reads that were output. |
--most_frequent_unmatched | file output | It contains the (approximate) counts of the most prevelant observed barcode sequences that did not match to one of the expected barcodes. Can only be created when 'most_unmatched_to_output' is not set to 0. |
--sample_barcode_hop_metrics | file output | File containing the frequently observed barcodes that are unexpected combinations of expected barcodes in a dual-indexed run. |
--per_project_metrics | file output | Aggregates the metrics by project (aggregates the metrics across samples with the same project) and has the same columns as `--metrics`. In this case, sample_ID will contain the project name (or None if no project is given). THe barcode will contain all Ns. The undetermined sample will not be aggregated with any other sample. |
--per_sample_metrics | file output | Tab-separated file containing statistics per sample. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--read_structures -r | string multiple | Read structures, one per input FASTQ. Do not provide when using a path prefix for FASTQs |
--allowed_mismatches -m | integer | Number of allowed mismatches between the observed barcode and the expected barcode |
--min_delta -d | integer | The minimum allowed difference between an observed barcode and the second closest expected barcode |
--free_ns -F | integer | Number of N's to allow in a barcode without counting against the allowed_mismatches |
--max_no_calls -N | integer | Max no-calls (N's) in a barcode before it is considered unmatchable. A barcode with total N's greater than 'max_no_call' will be considered unmatchable. |
--quality_mask_threshold -M | integer multiple | Mask template bases with quality scores less than specified value(s). Sample barcode/index and UMI bases are never masked. If provided either a single value, or one value per FASTQ must be provided. |
--filter_control_reads -C | boolean_true | Filter out control reads |
--filter_failing_quality -Q | boolean_true | Filter reads failing quality filter |
--output_types -T | string multiple | The types of output FASTQs to write. For each read structure, all segment types listed will be output to a FASTQ file. These may be any of the following: - `T` - Template bases - `B` - Sample barcode bases - `M` - Molecular barcode bases - `S` - Skip bases |
--undetermined_sample_name -u | string | The sample name for undetermined reads (reads that do not match an expected barcode) |
--most_unmatched_to_output -U | integer | Output the most frequent "unmatched" barcodes up to this number. If set to 0 unmatched barcodes will not be collected, improving overall performance. |
--override_matcher | string | If the sample barcodes are > 12 bp long, a cached hamming distance matcher is used. If the barcodes are less than or equal to 12 bp long, all possible matches are precomputed. This option allows for overriding that heuristic. |
--skip_read_name_check | boolean_true | If this is true, then all the read names across FASTQs will not be enforced to be the same. This may be useful when the read names are known to be the same and performance matters. Regardless, the first read name in each FASTQ will always be checked. |
--sample_barcode_in_fastq_header | boolean_true | If this is true, then the sample barcode is expected to be in the FASTQ read header. For dual indexed data, the barcodes must be `+` (plus) delimited. Additionally, if true, then neither index FASTQ files nor sample barcode segments in the read structure may be specified. |
--metric_prefix | string | Prepend this prefix to all output metric file names |
--lane -l | integer multiple | Select a subset of lanes to demultiplex. Will cause only samples and input FASTQs with the given `Lane`(s) to be demultiplexed. Samples without a lane will be ignored, and FASTQs without lane information will be ignored |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
sample_fastq: "$id.$key.sample_fastq.output"
metrics: "$id.$key.metrics.tsv"
most_frequent_unmatched: "$id.$key.most_frequent_unmatched.tsv"
sample_barcode_hop_metrics: "$id.$key.sample_barcode_hop_metrics.tsv"
per_project_metrics: "$id.$key.per_project_metrics.tsv"
per_sample_metrics: "$id.$key.per_sample_metrics.tsv"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.3.0 \
-main-script target/nextflow/sgdemux/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
sgdemuxbiobox v0.3.0
Uses
0 relationships
No component dependencies found.