workflows/multiomics/process_samples
Description
A pipeline to analyse multiple multiomics samples.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input -i | file required | Path to the sample. |
--rna_layer | string | Input layer for the gene expression modality. If not specified, .X is used. |
--prot_layer | string | Input layer for the antibody capture modality. If not specified, .X is used. |
--gdo_layer | string | Input layer for the guide-derived oligonucleotide (GDO) data. If not specified, .X is used. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
Sample ID options
Name | Type & Properties | Description |
|---|---|---|
--add_id_to_obs | boolean | Add the value passed with --id to .obs. |
--add_id_obs_output | string | .Obs column to add the sample IDs to. Required and only used when --add_id_to_obs is set to 'true' |
--add_id_make_observation_keys_unique | boolean | Join the id to the .obs index (.obs_names). Only used when --add_id_to_obs is set to 'true'. |
RNA filtering options
Name | Type & Properties | Description |
|---|---|---|
--rna_min_counts | integer | Minimum number of counts captured per cell. |
--rna_max_counts | integer | Maximum number of counts captured per cell. |
--rna_min_genes_per_cell | integer | Minimum of non-zero values per cell. |
--rna_max_genes_per_cell | integer | Maximum of non-zero values per cell. |
--rna_min_cells_per_gene | integer | Minimum of non-zero values per gene. |
--rna_min_fraction_mito | double | Minimum fraction of UMIs that are mitochondrial. |
--rna_max_fraction_mito | double | Maximum fraction of UMIs that are mitochondrial. |
--rna_min_fraction_ribo | double | Minimum fraction of UMIs that are mitochondrial. |
--rna_max_fraction_ribo | double | Maximum fraction of UMIs that are mitochondrial. |
--skip_scrublet_doublet_detection | boolean_true | Skip the scrublet doublet detection step. |
CITE-seq filtering options
Name | Type & Properties | Description |
|---|---|---|
--prot_min_counts | integer | Minimum number of counts per cell. |
--prot_max_counts | integer | Minimum number of counts per cell. |
--prot_min_proteins_per_cell | integer | Minimum of non-zero values per cell. |
--prot_max_proteins_per_cell | integer | Maximum of non-zero values per cell. |
--prot_min_cells_per_protein | integer | Minimum of non-zero values per protein. |
GDO filtering options
Name | Type & Properties | Description |
|---|---|---|
--gdo_min_counts | integer | Minimum number of counts per cell. |
--gdo_max_counts | integer | Minimum number of counts per cell. |
--gdo_min_guides_per_cell | integer | Minimum of non-zero values per cell. |
--gdo_max_guides_per_cell | integer | Maximum of non-zero values per cell. |
--gdo_min_cells_per_guide | integer | Minimum of non-zero values per guide. |
Highly variable features detection
Name | Type & Properties | Description |
|---|---|---|
--highly_variable_features_var_output --filter_with_hvg_var_output | string | In which .var slot to store a boolean array corresponding to the highly variable genes. |
--highly_variable_features_obs_batch_key --filter_with_hvg_obs_batch_key | string | If specified, highly-variable genes are selected within each batch separately and merged. This simple process avoids the selection of batch-specific genes and acts as a lightweight batch correction method. |
Mitochondrial & Ribosomal Gene Detection
Name | Type & Properties | Description |
|---|---|---|
--var_gene_names | string | .var column name to be used to detect mitochondrial/ribosomal genes instead of .var_names (default if not set). Gene names matching with the regex value from --mitochondrial_gene_regex or --ribosomal_gene_regex will be identified as mitochondrial or ribosomal genes, respectively. |
--var_name_mitochondrial_genes | string | In which .var slot to store a boolean array corresponding the mitochondrial genes. |
--obs_name_mitochondrial_fraction | string | When specified, write the fraction of counts originating from mitochondrial genes (based on --mitochondrial_gene_regex) to an .obs column with the specified name. Requires --var_name_mitochondrial_genes. |
--mitochondrial_gene_regex | string | Regex string that identifies mitochondrial genes from --var_gene_names. By default will detect human and mouse mitochondrial genes from a gene symbol. |
--var_name_ribosomal_genes | string | In which .var slot to store a boolean array corresponding the ribosomal genes. |
--obs_name_ribosomal_fraction | string | When specified, write the fraction of counts originating from ribosomal genes (based on --ribosomal_gene_regex) to an .obs column with the specified name. Requires --var_name_ribosomal_genes. |
--ribosomal_gene_regex | string | Regex string that identifies ribosomal genes from --var_gene_names. By default will detect human and mouse ribosomal genes from a gene symbol. |
QC metrics calculation options
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. Defaults to the combined values specified for --var_name_mitochondrial_genes and --highly_variable_features_var_output. |
--top_n_vars | integer multiple | Number of top vars to be used to calculate cumulative proportions. If not specified, proportions are not calculated. `--top_n_vars 20,50` finds cumulative proportion to the 20th and 50th most expressed vars. |
PCA options
Name | Type & Properties | Description |
|---|---|---|
--pca_overwrite | boolean_true | Allow overwriting slots for PCA output. |
CLR options
Name | Type & Properties | Description |
|---|---|---|
--clr_axis | integer | Axis to perform the CLR transformation on. |
RNA Scaling options
Name | Type & Properties | Description |
|---|---|---|
--rna_enable_scaling | boolean_true | Enable scaling for the RNA modality. |
--rna_scaling_output_layer | string | Output layer where the scaled log-normalized data will be stored. |
--rna_scaling_pca_obsm_output | string | Name of the .obsm key where the PCA representation of the log-normalized and scaled data is stored. |
--rna_scaling_pca_loadings_varm_output | string | Name of the .varm key where the PCA loadings of the log-normalized and scaled data is stored. |
--rna_scaling_pca_variance_uns_output | string | Name of the .uns key where the variance and variance ratio will be stored as a map. The map will contain two keys: variance and variance_ratio respectively. |
--rna_scaling_umap_obsm_output | string | Name of the .obsm key where the UMAP representation of the log-normalized and scaled data is stored. |
--rna_scaling_max_value | double | Clip (truncate) data to this value after scaling. If not specified, do not clip. |
--rna_scaling_zero_center | boolean_false | If set, omit zero-centering variables, which allows to handle sparse input efficiently." |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
output: "$id.$key.output.h5mu"
add_id_to_obs: [ true ]
add_id_obs_output: [ "sample_id" ]
add_id_make_observation_keys_unique: [ true ]
highly_variable_features_var_output: [ "filter_with_hvg" ]
highly_variable_features_obs_batch_key: [ "sample_id" ]
mitochondrial_gene_regex: [ "^[mM][tT]-" ]
ribosomal_gene_regex: [ "^[Mm]?[Rr][Pp][LlSs]" ]
top_n_vars: [ 50, 100, 200, 500 ]
clr_axis: [ 0 ]
rna_scaling_output_layer: [ "scaled" ]
rna_scaling_pca_obsm_output: [ "scaled_pca" ]
rna_scaling_pca_loadings_varm_output: [ "scaled_pca_loadings" ]
rna_scaling_pca_variance_uns_output: [ "scaled_pca_variance" ]
rna_scaling_umap_obsm_output: [ "scaled_umap" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/multiomics/process_samples/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
workflows/multiomics/process_samplesopenpipeline v4.0.0