workflows/ingestion/cellranger_multi
Description
A pipeline for running Cell Ranger multi.
Input files
Name | Type & Properties | Description |
|---|---|---|
--input | file multiple | The FASTQ files to be analyzed. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
Feature type-specific input files
Name | Type & Properties | Description |
|---|---|---|
--gex_input | file multiple | The FASTQ files to be analyzed for Gene Expression. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--abc_input | file multiple | The FASTQ files to be analyzed for Antibody Capture. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--cgc_input | file multiple | The FASTQ files to be analyzed for CRISPR Guide Capture. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--mux_input | file multiple | The FASTQ files to be analyzed for Multiplexing Capture. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--vdj_input | file multiple | The FASTQ files to be analyzed for VDJ. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--vdj_t_input | file multiple | The FASTQ files to be analyzed for VDJ-T. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--vdj_t_gd_input | file multiple | The FASTQ files to be analyzed for VDJ-T-GD. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--vdj_b_input | file multiple | The FASTQ files to be analyzed for VDJ-B. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
--agc_input | file multiple | The FASTQ files to be analyzed for Antigen Capture. FASTQ files should conform to the naming conventions of bcl2fastq and mkfastq: `[Sample Name]_S[Sample Index]_L00[Lane Number]_[Read Type]_001.fastq.gz` |
Library arguments
Name | Type & Properties | Description |
|---|---|---|
--library_id | string multiple | The Illumina sample name to analyze. This must exactly match the 'Sample Name'part of the FASTQ files specified in the `--input` argument. |
--library_type | string multiple | The underlying feature type of the library. |
--library_subsample | string multiple | The rate at which reads from the provided FASTQ files are sampled. Must be strictly greater than 0 and less than or equal to 1. |
--library_lanes | string multiple | Lanes associated with this sample. Defaults to using all lanes. |
--library_chemistry | string | Only applicable to FRP. Library-specific assay configuration. By default, the assay configuration is detected automatically. Typically, users will not need to specify a chemistry. |
Sample parameters
Name | Type & Properties | Description |
|---|---|---|
--sample_ids --cell_multiplex_sample_id | string multiple | A name to identify a multiplexed sample. Must be alphanumeric with hyphens and/or underscores, and less than 64 characters. Required for Cell Multiplexing libraries. |
--sample_description --cell_multiplex_description | string multiple | A description for the sample. |
--sample_expect_cells | integer multiple | Expected number of recovered cells, used as input to cell calling algorithm. |
--sample_force_cells | integer multiple | Force pipeline to use this number of cells, bypassing cell detection. |
Feature Barcode library specific arguments
Name | Type & Properties | Description |
|---|---|---|
--feature_reference | file | Path to the Feature reference CSV file, declaring Feature Barcode constructs and associated barcodes. Required only for Antibody Capture or CRISPR Guide Capture libraries. See https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/using/feature-bc-analysis#feature-ref for more information." |
--feature_r1_length | integer | Limit the length of the input Read 1 sequence of V(D)J libraries to the first N bases, where N is the user-supplied value. Note that the length includes the Barcode and UMI sequences so do not set this below 26. |
--feature_r2_length | integer | Limit the length of the input Read 2 sequence of V(D)J libraries to the first N bases, where N is a user-supplied value. Trimming occurs before sequencing metrics are computed and therefore, limiting the length of Read 2 may affect Q30 scores. |
--min_crispr_umi | integer | Set the minimum number of CRISPR guide RNA UMIs required for protospacer detection. If a lower or higher sensitivity is desired for detection, this value can be customized according to specific experimental needs. Applicable only to datasets that include a CRISPR Guide Capture library. |
Gene expression arguments
Name | Type & Properties | Description |
|---|---|---|
--gex_reference | file required | Genome refence index built by Cell Ranger mkref. |
--gex_secondary_analysis | boolean | Whether or not to run the secondary analysis e.g. clustering. |
--gex_generate_bam | boolean | Whether to generate a BAM file. |
--tenx_cloud_token_path | file | The 10x Cloud Analysis user token used to enable cell annotation. |
--cell_annotation_model | string | "Cell annotation model to use. If auto, uses the default model for the species. If not given, does not run cell annotation." |
--gex_expect_cells | integer | Expected number of recovered cells, used as input to cell calling algorithm. |
--gex_force_cells | integer | Force pipeline to use this number of cells, bypassing cell detection. |
--gex_include_introns | boolean | Whether or not to include intronic reads in counts. This option does not apply to Fixed RNA Profiling analysis. |
--gex_r1_length | integer | Limit the length of the input Read 1 sequence of V(D)J libraries to the first N bases, where N is the user-supplied value. Note that the length includes the Barcode and UMI sequences so do not set this below 26. |
--gex_r2_length | integer | Limit the length of the input Read 2 sequence of V(D)J libraries to the first N bases, where N is a user-supplied value. Trimming occurs before sequencing metrics are computed and therefore, limiting the length of Read 2 may affect Q30 scores. |
--gex_chemistry | string | Assay configuration. Either specify a single value which will be applied to all libraries, or a number of values that is equal to the number of libararies. The latter is only applicable to only applicable to Fixed RNA Profiling. - auto: Chemistry autodetection (default) - threeprime: Single Cell 3' - SC3Pv1, SC3Pv2, SC3Pv3(-polyA), SC3Pv4(-polyA): Single Cell 3' v1, v2, v3, or v4 - SC3Pv3HT(-polyA): Single Cell 3' v3.1 HT - SC-FB: Single Cell Antibody-only 3' v2 or 5' - fiveprime: Single Cell 5' - SC5P-PE: Paired-end Single Cell 5' - SC5P-PE-v3: Paired-end Single Cell 5' v3 - SC5P-R2: R2-only Single Cell 5' - SC5P-R2-v3: R2-only Single Cell 5' v3 - SCP5-PE-v3: Single Cell 5' paired-end v3 (GEM-X) - SC5PHT : Single Cell 5' v2 HT - SFRP: Fixed RNA Profiling (Singleplex) - MFRP: Fixed RNA Profiling (Multiplex, Probe Barcode on R2) - MFRP-R1: Fixed RNA Profiling (Multiplex, Probe Barcode on R1) - MFRP-RNA: Fixed RNA Profiling (Multiplex, RNA, Probe Barcode on R2) - MFRP-Ab: Fixed RNA Profiling (Multiplex, Antibody, Probe Barcode at R2:69) - MFRP-Ab-R2pos50: Fixed RNA Profiling (Multiplex, Antibody, Probe Barcode at R2:50) - MFRP-RNA-R1: Fixed RNA Profiling (Multiplex, RNA, Probe Barcode on R1) - MFRP-Ab-R1: Fixed RNA Profiling (Multiplex, Antibody, Probe Barcode on R1) - ARC-v1 for analyzing the Gene Expression portion of Multiome data. If Cell Ranger auto-detects ARC-v1 chemistry, an error is triggered. See https://kb.10xgenomics.com/hc/en-us/articles/115003764132-How-does-Cell-Ranger-auto-detect-chemistry- for more information. |
VDJ related parameters
Name | Type & Properties | Description |
|---|---|---|
--vdj_reference | file | VDJ refence index built by Cell Ranger mkref. |
--vdj_inner_enrichment_primers | file | V(D)J Immune Profiling libraries: if inner enrichment primers other than those provided in the 10x Genomics kits are used, they need to be specified here as a text file with one primer per line. |
--vdj_r1_length | integer | Limit the length of the input Read 1 sequence of V(D)J libraries to the first N bases, where N is the user-supplied value. Note that the length includes the Barcode and UMI sequences so do not set this below 26. |
--vdj_r2_length | integer | Limit the length of the input Read 2 sequence of V(D)J libraries to the first N bases, where N is a user-supplied value. Trimming occurs before sequencing metrics are computed and therefore, limiting the length of Read 2 may affect Q30 scores |
--vdj_denovo | boolean | Run in reference-free mode (i.e., do not use annotations). This option is not supported for multiplexed experiments. |
3' Cell multiplexing parameters (CellPlex Multiplexing)
Name | Type & Properties | Description |
|---|---|---|
--cell_multiplex_oligo_ids --cmo_ids | string multiple | The Cell Multiplexing oligo IDs used to multiplex this sample. If multiple CMOs were used for a sample, separate IDs with a pipe (e.g., CMO301|CMO302). Required for Cell Multiplexing libraries. |
--min_assignment_confidence | double | The minimum estimated likelihood to call a sample as tagged with a Cell Multiplexing Oligo (CMO) instead of "Unassigned". Users may wish to tolerate a higher rate of mis-assignment in order to obtain more singlets to include in their analysis, or a lower rate of mis-assignment at the cost of obtaining fewer singlets. |
--cmo_set | file | Path to a custom CMO set CSV file, declaring CMO constructs and associated barcodes. If the default CMO reference IDs that are built into the Cell Ranger software are required, this option does not need to be used. |
--barcode_sample_assignment | file | Path to a barcode-sample assignment CSV file that specifies the barcodes that belong to each sample. |
Hashtag multiplexing parameters
Name | Type & Properties | Description |
|---|---|---|
--hashtag_ids | string multiple | The hashtag IDs used to multiplex this sample. If multiple antibody hashtags were used for the same sample, you can separate IDs with a pipe. |
On-chip multiplexing parameters
Name | Type & Properties | Description |
|---|---|---|
--ocm_barcode_ids | string multiple | The OCM barcode IDs used to multiplex this sample. Must be one of OB1, OB2, OB3, OB4. If multiple OCM Barcodes were used for the same sample, you can separate IDs with a pipe (e.g., OB1|OB2). |
Flex multiplexing paramaters
Name | Type & Properties | Description |
|---|---|---|
--probe_set | file | A probe set reference CSV file. It specifies the sequences used as a reference for probe alignment and the gene ID associated with each probe. It must include 4 columns (probe file format 1.0.0): gene_id,probe_seq,probe_id,included,region and an optional 5th column (probe file format 1.0.1). - gene_id: The Ensembl gene identifier targeted by the probe. - probe_seq: The nucleotide sequence of the probe, which is complementary to the transcript sequence. - probe_id: The probe identifier, whose format is described in Probe identifiers. - included: A TRUE or FALSE flag specifying whether the probe is included in the filtered counts matrix output or excluded by the probe filter. See filter-probes option of cellranger multi. All probes of a gene must be marked TRUE in the included column for that gene to be included. - region: Present only in v1.0.1 probe set reference CSV. The gene boundary targeted by the probe. Accepted values are spliced or unspliced. The file also contains a number of required metadata fields in the header in the format #key=value: - panel_name: The name of the probe set. - panel_type: Always predesigned for predesigned probe sets. - reference_genome: The reference genome build used for probe design. - reference_version: The version of the Cell Ranger reference transcriptome used for probe design. - probe_set_file_format: The version of the probe set file format specification that this file conforms to. |
--filter_probes | boolean | If 'false', include all non-deprecated probes listed in the probe set reference CSV file. If 'true' or not set, probes that are predicted to have off-target activity to homologous genes are excluded from analysis. Not filtering will result in UMI counts from all non-deprecated probes, including those with predicted off-target activity, to be used in the analysis. Probes whose ID is prefixed with DEPRECATED are always excluded from the analysis. |
--probe_barcode_ids | string multiple | The Fixed RNA Probe Barcode ID used for this sample, and for multiplex GEX + Antibody Capture libraries, the corresponding Antibody Multiplexing Barcode IDs. 10x recommends specifying both barcodes (e.g., BC001+AB001) when an Antibody Capture library is present. The barcode pair order is BC+AB and they are separated with a "+" (no spaces). Alternatively, you can specify the Probe Barcode ID alone and Cell Ranger's barcode pairing auto-detection algorithm will automatically match to the corresponding Antibody Multiplexing Barcode. |
--emptydrops_minimum_umis | integer | For singleplex Flex experiments, use this option to adjust the UMI cutoff during the second step of cell calling. Cell Ranger will still perform the full cell calling process but will only evaluate barcodes with UMIs above the threshold you specify. |
Antigen Capture (BEAM) libary arguments
Name | Type & Properties | Description |
|---|---|---|
--control_id | string multiple | A user-defined ID for any negative controls used in the T/BCR Antigen Capture assay. Must match id specified in the feature reference CSV. May only include ASCII characters and must not use whitespace, slash, quote, or comma characters. Each ID must be unique and must not collide with a gene identifier from the transcriptome. |
--mhc_allele | string multiple | The MHC allele for TCR Antigen Capture libraries. Must match mhc_allele name specified in the Feature Reference CSV. |
General arguments
Name | Type & Properties | Description |
|---|---|---|
--check_library_compatibility | boolean | Optional. This option allows users to disable the check that evaluates 10x Barcode overlap between ibraries when multiple libraries are specified (e.g., Gene Expression + Antibody Capture). Setting this option to false will disable the check across all library combinations. We recommend running this check (default), however if the pipeline errors out, users can bypass the check to generate outputs for troubleshooting. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output_raw | file required output | The raw output folder. |
--output_h5mu | file required output | Locations for the output files. Must contain a wildcard (*) character, which will be replaced with the sample name. |
--uns_metrics | string | Name of the .uns slot under which to QC metrics (if any). |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
gex_secondary_analysis: [ false ]
gex_generate_bam: [ false ]
gex_include_introns: [ true ]
gex_chemistry: [ "auto" ]
check_library_compatibility: [ true ]
output_raw: "$id.$key.output_raw.output_dir"
output_h5mu: "$id.$key.output_h5mu.h5mu"
uns_metrics: [ "metrics_cellranger" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/workflows/ingestion/cellranger_multi/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/ingestion/cellranger_multiopenpipeline v4.2.0