single_cell/process_integrate_annotate
Description
A pipeline to process, integrate and annotate single cell (multi-)omics data.
Available integration methods:
Harmony
scVI
Available annotation methods:CellTypist
scANVI (with scArches)
Input (query) data arguments
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input query dataset(s) to be annotated |
--modality | string | Modality to be processed. Should match the modality in the --reference dataset, if provided. |
--input_layer | string | The layer in the input data containing the raw counts, if .X is not to be used. |
--input_var_gene_names | string | The name of the adata var column containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
Reference data arguments
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference dataset in .h5mu format to be used as a reference mapper and to train annotation algorithms on. |
--reference_layer_raw_counts | string | The layer in the reference dataset containing the raw counts, if .X is not to be used. |
--reference_layer_lognormalized_counts | string | The layer in the reference dataset containing the log-normalized counts, if .X is not to be used. |
--reference_var_gene_names | string | The name of the adata .var column containing gene names if the .var index is not to be used. |
--reference_obs_batch | string | The .obs column of the reference dataset containing the batch information. |
--reference_obs_label | string | The `.obs` key of the target labels to tranfer. |
--reference_obs_label_unlabeled_category | string | Value in the --reference_obs_label field that indicates unlabeled observations |
--reference_var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
Methods
Name | Type & Properties | Description |
|---|---|---|
--integration_methods | string multiple | Integration methods to be executed. |
--annotation_methods | string multiple | Annotation methods to be executed. |
Pre-processing options: RNA filtering
Name | Type & Properties | Description |
|---|---|---|
--rna_min_counts | integer | Minimum number of counts captured per cell. |
--rna_max_counts | integer | Maximum number of counts captured per cell. |
--rna_min_genes_per_cell | integer | Minimum of non-zero values per cell. |
--rna_max_genes_per_cell | integer | Maximum of non-zero values per cell. |
--rna_min_cells_per_gene | integer | Minimum of non-zero values per gene. |
--rna_min_fraction_mito | double | Minimum fraction of UMIs that are mitochondrial. |
--rna_max_fraction_mito | double | Maximum fraction of UMIs that are mitochondrial. |
Pre-processing options: Highly variable features detection
Name | Type & Properties | Description |
|---|---|---|
--n_hvg | integer | Number of highly-variable features to keep. Only relevant if HVG need to be calculated across query and reference datasets (e.g. for --annotation_methods scvi_knn and harmony_knn). For reference mapping-based methods, the HVG's specified in --reference_var_input will be used. |
Pre-processing options: Mitochondrial & Ribosomal Gene Detection
Name | Type & Properties | Description |
|---|---|---|
--var_name_mitochondrial_genes | string | In which .var slot to store a boolean array corresponding the mitochondrial genes. |
--var_name_ribosomal_genes | string | In which .var slot to store a boolean array corresponding the ribosomal genes. |
--obs_name_mitochondrial_fraction | string | When specified, write the fraction of counts originating from mitochondrial genes (based on --mitochondrial_gene_regex) to an .obs column with the specified name. Requires --var_name_mitochondrial_genes. |
--obs_name_ribosomal_fraction | string | When specified, write the fraction of counts originating from ribosomal genes (based on --ribosomal_gene_regex) to an .obs column with the specified name. Requires --var_name_ribosomal_genes. |
--mitochondrial_gene_regex | string | Regex string that identifies mitochondrial genes from --var_gene_names. By default will detect human and mouse mitochondrial genes from a gene symbol. |
--ribosomal_gene_regex | string | Regex string that identifies ribosomal genes from --var_gene_names. By default will detect human and mouse ribosomal genes from a gene symbol. |
Pre-processing options: QC metrics calculation options
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. Defaults to the combined values specified for --var_name_mitochondrial_genes and --highly_variable_features_var_output. |
Harmony integration options
Name | Type & Properties | Description |
|---|---|---|
--harmony_theta | double multiple | Diversity clustering penalty parameter. Specify for each variable in group.by.vars. theta=0 does not encourage any diversity. Larger values of theta result in more diverse clusters." |
--harmony_obs_covariates | string required multiple | The .obs field(s) that define the covariate(s) to regress out. |
scVI, scANVI and scArches training options
Name | Type & Properties | Description |
|---|---|---|
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
CellTypist reference model
Name | Type & Properties | Description |
|---|---|---|
--celltypist_model | file | Pretrained model in pkl format. If not provided, the model will be trained on the reference data and --reference should be provided. |
CellTypist annotation options
Name | Type & Properties | Description |
|---|---|---|
--celltypist_feature_selection | boolean | Whether to perform feature selection. |
--celltypist_majority_voting | boolean | Whether to refine the predicted labels by running the majority voting classifier after over-clustering. |
--celltypist_C | double | Inverse of regularization strength in logistic regression. |
--celltypist_max_iter | integer | Maximum number of iterations before reaching the minimum of the cost function. |
--celltypist_use_SGD | boolean_true | Whether to use the stochastic gradient descent algorithm. |
--celltypist_min_prop | double | "For the dominant cell type within a subcluster, the minimum proportion of cells required to support naming of the subcluster by this cell type. Ignored if majority_voting is set to False. Subcluster that fails to pass this proportion threshold will be assigned 'Heterogeneous'." |
Clustering options
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Neighbor classifier arguments
Name | Type & Properties | Description |
|---|---|---|
--knn_weights | string | Weight function used in prediction. Possible values are: `uniform` (all points in each neighborhood are weighted equally) or `distance` (weight points by the inverse of their distance) |
--knn_n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The output file. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
reference_layer_lognormalized_counts: [ "log_normalized" ]
reference_obs_label_unlabeled_category: [ "Unkown" ]
n_hvg: [ 2000 ]
mitochondrial_gene_regex: [ "^[mM][tT]-" ]
ribosomal_gene_regex: [ "^[Mm]?[Rr][Pp][LlSs]" ]
harmony_theta: [ 2 ]
harmony_obs_covariates: [ "sample_id" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
celltypist_feature_selection: [ false ]
celltypist_majority_voting: [ false ]
celltypist_C: [ 1 ]
celltypist_max_iter: [ 1000 ]
celltypist_min_prop: [ 0 ]
leiden_resolution: [ 1 ]
knn_weights: [ "uniform" ]
knn_n_neighbors: [ 15 ]
output: "$id.$key.output.h5mu"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_composed.git \
-revision v0.1.0 \
-main-script target/nextflow/single_cell/process_integrate_annotate/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
single_cell/process_integrate_annotateopenpipeline_composed v0.1.0