workflows/integration/totalvi_leiden
Description
Run totalVI integration, followed by neighbour calculations, leiden clustering and UMAP.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input h5mu file to be integrated. |
--rna_modality | string | |
--prot_modality | string | Name of the modality in the input (query) h5mu file containing protein data |
--input_layer_rna | string | Input layer to use from the rna modality for gene counts. If None, X is used |
--input_layer_protein | string | Input layer to use from the protein modality for protein counts. If None, X is used |
--obs_batch | string | Column name discriminating between your batches. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--var_gene_names | string | .var column in the rna modality containing gene names. By default, use mod['rna'].var_names. |
--var_protein_names | string | .var column in the protein modality containing protein names. By default, use mod['prot'].var_names. |
--var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
--output_model | file output | Directory where the trained reference model will be stored. |
TotalVI Training Options
Name | Type & Properties | Description |
|---|---|---|
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
TotalVI Integration Options
Name | Type & Properties | Description |
|---|---|---|
--obsm_integrated | string | In which .obsm slot to store the resulting integrated embedding. |
--obsm_normalized_rna_output | string | In which .obsm slot to store the normalized RNA data from TOTALVI. |
--obsm_normalized_protein_output | string | In which .obsm slot to store the normalized protein data from TOTALVI. |
Neighbour calculation
Name | Type & Properties | Description |
|---|---|---|
--uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Clustering options
Name | Type & Properties | Description |
|---|---|---|
--obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions resolutions specified in '--leiden_resolution'. |
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Umap options
Name | Type & Properties | Description |
|---|---|---|
--obsm_umap | string | In which .obsm slot to store the resulting UMAP embedding. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
rna_modality: [ "rna" ]
prot_modality: [ "prot" ]
obs_batch: [ "sample_id" ]
output: "$id.$key.output.h5mu"
output_model: "$id.$key.output_model"
early_stopping: [ true ]
obsm_integrated: [ "X_integrated_totalvi" ]
obsm_normalized_rna_output: [ "X_totalvi_normalized_rna" ]
obsm_normalized_protein_output: [ "X_totalvi_normalized_protein" ]
uns_neighbors: [ "totalvi_integration_neighbors" ]
obsp_neighbor_distances: [ "totalvi_integration_distances" ]
obsp_neighbor_connectivities: [ "totalvi_integration_connectivities" ]
obs_cluster: [ "totalvi_integration_leiden" ]
leiden_resolution: [ 1 ]
obsm_umap: [ "X_totalvi_umap" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/integration/totalvi_leiden/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/integration/totalvi_leidenopenpipeline v4.0.0