workflows/multiomics/process_samples
Description
A pipeline to analyse multiple multiomics samples.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input -i | file required | Path to the sample. |
--rna_layer | string | Input layer for the gene expression modality. If not specified, .X is used. |
--prot_layer | string | Input layer for the antibody capture modality. If not specified, .X is used. |
--gdo_layer | string | Input layer for the guide-derived oligonucleotide (GDO) data. If not specified, .X is used. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
Sample ID options
Name | Type & Properties | Description |
|---|---|---|
--add_id_to_obs | boolean | Add the value passed with --id to .obs. |
--add_id_obs_output | string | .Obs column to add the sample IDs to. Required and only used when --add_id_to_obs is set to 'true' |
--add_id_make_observation_keys_unique | boolean | Join the id to the .obs index (.obs_names). Only used when --add_id_to_obs is set to 'true'. |
RNA filtering options
Name | Type & Properties | Description |
|---|---|---|
--rna_min_counts | integer | Minimum number of counts captured per cell. |
--rna_max_counts | integer | Maximum number of counts captured per cell. |
--rna_min_genes_per_cell | integer | Minimum of non-zero values per cell. |
--rna_max_genes_per_cell | integer | Maximum of non-zero values per cell. |
--rna_min_cells_per_gene | integer | Minimum of non-zero values per gene. |
--rna_min_fraction_mito | double | Minimum fraction of UMIs that are mitochondrial. |
--rna_max_fraction_mito | double | Maximum fraction of UMIs that are mitochondrial. |
CITE-seq filtering options
Name | Type & Properties | Description |
|---|---|---|
--prot_min_counts | integer | Minimum number of counts per cell. |
--prot_max_counts | integer | Minimum number of counts per cell. |
--prot_min_proteins_per_cell | integer | Minimum of non-zero values per cell. |
--prot_max_proteins_per_cell | integer | Maximum of non-zero values per cell. |
--prot_min_cells_per_protein | integer | Minimum of non-zero values per protein. |
GDO filtering options
Name | Type & Properties | Description |
|---|---|---|
--gdo_min_counts | integer | Minimum number of counts per cell. |
--gdo_max_counts | integer | Minimum number of counts per cell. |
--gdo_min_guides_per_cell | integer | Minimum of non-zero values per cell. |
--gdo_max_guides_per_cell | integer | Maximum of non-zero values per cell. |
--gdo_min_cells_per_guide | integer | Minimum of non-zero values per guide. |
Highly variable features detection
Name | Type & Properties | Description |
|---|---|---|
--highly_variable_features_var_output --filter_with_hvg_var_output | string | In which .var slot to store a boolean array corresponding to the highly variable genes. |
--highly_variable_features_obs_batch_key --filter_with_hvg_obs_batch_key | string | If specified, highly-variable genes are selected within each batch separately and merged. This simple process avoids the selection of batch-specific genes and acts as a lightweight batch correction method. |
Mitochondrial Gene Detection
Name | Type & Properties | Description |
|---|---|---|
--var_name_mitochondrial_genes | string | In which .var slot to store a boolean array corresponding the mitochondrial genes. |
--obs_name_mitochondrial_fraction | string | When specified, write the fraction of counts originating from mitochondrial genes (based on --mitochondrial_gene_regex) to an .obs column with the specified name. Requires --var_name_mitochondrial_genes. |
--var_gene_names | string | .var column name to be used to detect mitochondrial genes instead of .var_names (default if not set). Gene names matching with the regex value from --mitochondrial_gene_regex will be identified as a mitochondrial gene. |
--mitochondrial_gene_regex | string | Regex string that identifies mitochondrial genes from --var_gene_names. By default will detect human and mouse mitochondrial genes from a gene symbol. |
QC metrics calculation options
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. Defaults to the combined values specified for --var_name_mitochondrial_genes and --highly_variable_features_var_output. |
--top_n_vars | integer multiple | Number of top vars to be used to calculate cumulative proportions. If not specified, proportions are not calculated. `--top_n_vars 20,50` finds cumulative proportion to the 20th and 50th most expressed vars. |
PCA options
Name | Type & Properties | Description |
|---|---|---|
--pca_overwrite | boolean_true | Allow overwriting slots for PCA output. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
output: "$id.$key.output.h5mu"
add_id_to_obs: [ true ]
add_id_obs_output: [ "sample_id" ]
add_id_make_observation_keys_unique: [ true ]
highly_variable_features_var_output: [ "filter_with_hvg" ]
highly_variable_features_obs_batch_key: [ "sample_id" ]
mitochondrial_gene_regex: [ "^[mM][tT]-" ]
top_n_vars: [ 50, 100, 200, 500 ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision 1.0.0 \
-main-script target/nextflow/workflows/multiomics/process_samples/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/multiomics/process_samplesopenpipeline 1.0.0