mapping/bd_rhapsody
Description
BD Rhapsody Sequence Analysis CWL pipeline v2.2.1
This pipeline performs analysis of single-cell multiomic sequence read (FASTQ) data. The supported
sequencing libraries are those generated by the BD Rhapsody assay kits, including: Whole Transcriptome
mRNA, Targeted mRNA, AbSeq Antibody-Oligonucleotides, Single-Cell Multiplexing, TCR/BCR, and
ATAC-Seq
The CWL pipeline file is obtained by cloning 'https://bitbucket.org/CRSwDev/cwl' and removing all objects with class 'DockerRequirement' from the YAML.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--reads | file multiple | Reads (optional) - Path to your FASTQ.GZ formatted read files from libraries that may include: - WTA mRNA - Targeted mRNA - AbSeq - Sample Multiplexing - VDJ You may specify as many R1/R2 read pairs as you want. |
--reads_atac | file multiple | Path to your FASTQ.GZ formatted read files from ATAC-Seq libraries. You may specify as many R1/R2/I2 files as you want. |
References
Name | Type & Properties | Description |
|---|---|---|
--reference_archive | file | Path to Rhapsody WTA Reference in the tar.gz format. Structure of the reference archive: - `BD_Rhapsody_Reference_Files/`: top level folder - `star_index/`: sub-folder containing STAR index, that is files created with `STAR --runMode genomeGenerate` - GTF for gene-transcript-annotation e.g. "gencode.v43.primary_assembly.annotation.gtf" |
--targeted_reference | file multiple | Path to the targeted reference file in FASTA format. |
--abseq_reference | file multiple | Path to the AbSeq reference file in FASTA format. Only needed if BD AbSeq Ab-Oligos are used. |
--supplemental_reference -s | file multiple | Path to the supplemental reference file in FASTA format. Only needed if there are additional transgene sequences to be aligned against in a WTA assay experiment. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output_dir -o | file required output | The unprocessed output directory containing all the outputs from the pipeline. |
--output_seurat | file output | Single-cell analysis tool inputs. Seurat (.rds) input file containing RSEC molecules data table and all cell annotation metadata. |
--output_mudata | file output | |
--metrics_summary | file output | Metrics Summary. Report containing sequencing, molecules, and cell metrics. |
--pipeline_report | file output | Pipeline Report. Summary report containing the results from the sequencing analysis pipeline run. |
--rsec_mols_per_cell | file output | Molecules per bioproduct per cell bassed on RSEC |
--dbec_mols_per_cell | file output | Molecules per bioproduct per cell bassed on DBEC. DBEC data table is only output if the experiment includes targeted mRNA or AbSeq bioproducts. |
--rsec_mols_per_cell_unfiltered | file output | Unfiltered tables containing all cell labels with 10 reads. |
--bam | file output | Alignment file of R2 with associated R1 annotations for Bioproduct. |
--bam_index | file output | Index file for the alignment file. |
--bioproduct_stats | file output | Bioproduct Stats. Metrics from RSEC and DBEC Unique Molecular Identifier adjustment algorithms on a per-bioproduct basis. |
--dimred_tsne | file output | t-SNE dimensionality reduction coordinates per cell index |
--dimred_umap | file output | UMAP dimensionality reduction coordinates per cell index |
--immune_cell_classification | file output | Immune Cell Classification. Cell type classification based on the expression of immune cell markers. |
Multiplex outputs
Name | Type & Properties | Description |
|---|---|---|
--sample_tag_metrics | file output | Sample Tag Metrics. Metrics from the sample determination algorithm. |
--sample_tag_calls | file output | Sample Tag Calls. Assigned Sample Tag for each putative cell |
--sample_tag_counts | file multiple output | Sample Tag Counts. Separate data tables and metric summary for cells assigned to each sample tag. Note: For putative cells that could not be assigned a specific Sample Tag, a Multiplet_and_Undetermined.zip file is also output. |
--sample_tag_counts_unassigned | file output | Sample Tag Counts Unassigned. Data table and metric summary for cells that could not be assigned a specific Sample Tag. |
VDJ Outputs
Name | Type & Properties | Description |
|---|---|---|
--vdj_metrics | file output | VDJ Metrics. Overall metrics from the VDJ analysis. |
--vdj_per_cell | file output | VDJ Per Cell. Cell specific read and molecule counts, VDJ gene segments, CDR3 sequences, paired chains, and cell type. |
--vdj_per_cell_uncorrected | file output | VDJ Per Cell Uncorrected. Cell specific read and molecule counts, VDJ gene segments, CDR3 sequences, paired chains, and cell type. |
--vdj_dominant_contigs | file output | VDJ Dominant Contigs. Dominant contig for each cell label chain type combination (putative cells only). |
--vdj_unfiltered_contigs | file output | VDJ Unfiltered Contigs. All contigs that were assembled and annotated successfully (all cells). |
ATAC-Seq outputs
Name | Type & Properties | Description |
|---|---|---|
--atac_metrics | file output | ATAC Metrics. Overall metrics from the ATAC-Seq analysis. |
--atac_metrics_json | file output | ATAC Metrics JSON. Overall metrics from the ATAC-Seq analysis in JSON format. |
--atac_fragments | file output | ATAC Fragments. Chromosomal location, cell index, and read support for each fragment detected |
--atac_fragments_index | file output | Index of ATAC Fragments. |
--atac_transposase_sites | file output | ATAC Transposase Sites. Chromosomal location, cell index, and read support for each transposase site detected |
--atac_transposase_sites_index | file output | Index of ATAC Transposase Sites. |
--atac_peaks | file output | ATAC Peaks. Peak regions of transposase activity |
--atac_peaks_index | file output | Index of ATAC Peaks. |
--atac_peak_annotation | file output | ATAC Peak Annotation. Estimated annotation of peak-to-gene connections |
--atac_cell_by_peak | file output | ATAC Cell by Peak. Peak regions of transposase activity per cell |
--atac_cell_by_peak_unfiltered | file output | ATAC Cell by Peak Unfiltered. Unfiltered file containing all cell labels with >=1 transposase sites in peaks. |
--atac_bam | file output | ATAC BAM. Alignment file for R1 and R2 with associated I2 annotations for ATAC-Seq. Only output if the BAM generation flag is set to true. |
--atac_bam_index | file output | Index of ATAC BAM. |
AbSeq Cell Calling outputs
Name | Type & Properties | Description |
|---|---|---|
--protein_aggregates_experimental | file output | Protein Aggregates Experimental |
Putative Cell Calling Settings
Name | Type & Properties | Description |
|---|---|---|
--cell_calling_data | string | Specify the dataset to be used for putative cell calling: mRNA, AbSeq, ATAC, mRNA_and_ATAC For putative cell calling using an AbSeq dataset, please provide an AbSeq_Reference fasta file above. For putative cell calling using an ATAC dataset, please provide a WTA+ATAC-Seq Reference_Archive file above. The default data for putative cell calling, will be determined the following way: - If mRNA Reads and ATAC Reads exist: mRNA_and_ATAC - If only ATAC Reads exist: ATAC - Otherwise: mRNA |
--cell_calling_bioproduct_algorithm | string | Specify the bioproduct algorithm to be used for putative cell calling: Basic or Refined By default, the Basic algorithm will be used for putative cell calling. |
--cell_calling_atac_algorithm | string | Specify the ATAC-seq algorithm to be used for putative cell calling: Basic or Refined By default, the Basic algorithm will be used for putative cell calling. |
--exact_cell_count | integer | Set a specific number of cells as putative, based on those with the highest error-corrected read count |
--expected_cell_count | integer | Guide the basic putative cell calling algorithm by providing an estimate of the number of cells expected. Usually this can be the number of cells loaded into the Rhapsody cartridge. If there are multiple inflection points on the second derivative cumulative curve, this will ensure the one selected is near the expected. |
Intronic Reads Settings
Name | Type & Properties | Description |
|---|---|---|
--exclude_intronic_reads | boolean | By default, the flag is false, and reads aligned to exons and introns are considered and represented in molecule counts. When the flag is set to true, intronic reads will be excluded. The value can be true or false. |
Multiplex Settings
Name | Type & Properties | Description |
|---|---|---|
--sample_tags_version | string | Specify the version of the Sample Tags used in the run: * If Sample Tag Multiplexing was done, specify the appropriate version: human, mouse, flex, nuclei_includes_mrna, nuclei_atac_only * If this is an SMK + Nuclei mRNA run or an SMK + Multiomic ATAC-Seq (WTA+ATAC-Seq) run (and not an SMK + ATAC-Seq only run), choose the "nuclei_includes_mrna" option. * If this is an SMK + ATAC-Seq only run (and not SMK + Multiomic ATAC-Seq (WTA+ATAC-Seq)), choose the "nuclei_atac_only" option. |
--tag_names | string multiple | Specify the tag number followed by '-' and the desired sample name to appear in Sample_Tag_Metrics.csv Do not use the special characters. |
VDJ arguments
Name | Type & Properties | Description |
|---|---|---|
--vdj_version | string | If VDJ was done, specify the appropriate option: human, mouse, humanBCR, humanTCR, mouseBCR, mouseTCR |
ATAC options
Name | Type & Properties | Description |
|---|---|---|
--predefined_atac_peaks | file | An optional BED file containing pre-established chromatin accessibility peak regions for generating the ATAC cell-by-peak matrix. |
Additional options
Name | Type & Properties | Description |
|---|---|---|
--run_name | string | Specify a run name to use as the output file base name. Use only letters, numbers, or hyphens. Do not use special characters or spaces. |
--generate_bam | boolean | Specify whether to create the BAM file output |
--long_reads | boolean | Use STARlong (default: undefined - i.e. autodetects based on read lengths) - Specify if the STARlong aligner should be used instead of STAR. Set to true if the reads are longer than 650bp. |
Advanced options
Name | Type & Properties | Description |
|---|---|---|
--custom_star_params | string | Modify STAR alignment parameters - Set this parameter to fully override default STAR mapping parameters used in the pipeline. For reference this is the default that is used: Short Reads: `--outFilterScoreMinOverLread 0 --outFilterMatchNminOverLread 0 --outFilterMultimapScoreRange 0 --clip3pAdapterSeq AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA --seedSearchStartLmax 50 --outFilterMatchNmin 25 --limitOutSJcollapsed 2000000` Long Reads: Same as Short Reads + `--seedPerReadNmax 10000` This applies to fastqs provided in the Reads user input Do NOT set any non-mapping related params like `--genomeDir`, `--outSAMtype`, `--outSAMunmapped`, `--readFilesIn`, `--runThreadN`, etc. We use STAR version 2.7.10b |
--custom_bwa_mem2_params | string | Modify bwa-mem2 alignment parameters - Set this parameter to fully override bwa-mem2 mapping parameters used in the pipeline The pipeline does not specify any custom mapping params to bwa-mem2 so program default values are used This applies to fastqs provided in the Reads_ATAC user input Do NOT set any non-mapping related params like `-C`, `-t`, etc. We use bwa-mem2 version 2.2.1 |
CWL-runner arguments
Name | Type & Properties | Description |
|---|---|---|
--parallel | boolean | Run jobs in parallel. |
--timestamps | boolean_true | Add timestamps to the errors, warnings, and notifications. |
Undocumented arguments
Name | Type & Properties | Description |
|---|---|---|
--abseq_umi | integer | |
--target_analysis | boolean | |
--vdj_jgene_evalue | double | e-value threshold for J gene. The e-value threshold for J gene call by IgBlast/PyIR, default is set as 0.001 |
--vdj_vgene_evalue | double | e-value threshold for V gene. The e-value threshold for V gene call by IgBlast/PyIR, default is set as 0.001 |
--write_filtered_reads | boolean | Output processed FASTQ reads. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output_dir: "$id.$key.output_dir.output_dir"
output_seurat: "$id.$key.output_seurat.rds"
output_mudata: "$id.$key.output_mudata.h5mu"
metrics_summary: "$id.$key.metrics_summary.csv"
pipeline_report: "$id.$key.pipeline_report.html"
rsec_mols_per_cell: "$id.$key.rsec_mols_per_cell.zip"
dbec_mols_per_cell: "$id.$key.dbec_mols_per_cell.zip"
rsec_mols_per_cell_unfiltered: "$id.$key.rsec_mols_per_cell_unfiltered.zip"
bam: "$id.$key.bam.bam"
bam_index: "$id.$key.bam_index.bai"
bioproduct_stats: "$id.$key.bioproduct_stats.csv"
dimred_tsne: "$id.$key.dimred_tsne.csv"
dimred_umap: "$id.$key.dimred_umap.csv"
immune_cell_classification: "$id.$key.immune_cell_classification.csv"
sample_tag_metrics: "$id.$key.sample_tag_metrics.csv"
sample_tag_calls: "$id.$key.sample_tag_calls.csv"
sample_tag_counts: "$id.$key.sample_tag_counts._*.zip"
sample_tag_counts_unassigned: "$id.$key.sample_tag_counts_unassigned.zip"
vdj_metrics: "$id.$key.vdj_metrics.csv"
vdj_per_cell: "$id.$key.vdj_per_cell.csv"
vdj_per_cell_uncorrected: "$id.$key.vdj_per_cell_uncorrected.csv"
vdj_dominant_contigs: "$id.$key.vdj_dominant_contigs.csv"
vdj_unfiltered_contigs: "$id.$key.vdj_unfiltered_contigs.csv"
atac_metrics: "$id.$key.atac_metrics.csv"
atac_metrics_json: "$id.$key.atac_metrics_json.json"
atac_fragments: "$id.$key.atac_fragments.gz"
atac_fragments_index: "$id.$key.atac_fragments_index.tbi"
atac_transposase_sites: "$id.$key.atac_transposase_sites.gz"
atac_transposase_sites_index: "$id.$key.atac_transposase_sites_index.tbi"
atac_peaks: "$id.$key.atac_peaks.gz"
atac_peaks_index: "$id.$key.atac_peaks_index.tbi"
atac_peak_annotation: "$id.$key.atac_peak_annotation.gz"
atac_cell_by_peak: "$id.$key.atac_cell_by_peak.zip"
atac_cell_by_peak_unfiltered: "$id.$key.atac_cell_by_peak_unfiltered.zip"
atac_bam: "$id.$key.atac_bam.bam"
atac_bam_index: "$id.$key.atac_bam_index.bai"
protein_aggregates_experimental: "$id.$key.protein_aggregates_experimental.csv"
run_name: [ "sample" ]
generate_bam: [ false ]
parallel: [ true ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/mapping/bd_rhapsody/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
mapping/bd_rhapsodyopenpipeline v4.2.0
Uses
0 relationships
No component dependencies found.