workflows/integration/totalvi_scarches_leiden
Description
Run totalVI integration by mapping the query dataset to a reference, followed by neighbour calculations, leiden clustering and run umap on the result.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Path to the sample. |
--layer | string | use specified layer for expression values instead of the .X object from the modality. |
--modality | string | Which modality to process. |
--prot_modality | string | Which modality to process. |
--reference -r | file required | Input h5mu file with reference data to train the TOTALVI model. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
--reference_model_path | file output | Directory with the reference model. If not exists, trained model will be saved there |
--query_model_path | file output | Directory, where the query model will be saved |
General TotalVI Options
Name | Type & Properties | Description |
|---|---|---|
--obs_batch | string | .Obs column name discriminating between your batches. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--max_epochs | integer | Number of passes through the dataset |
--max_query_epochs | integer | Number of passes through the dataset, when fine-tuning model for query |
--weight_decay | double | Weight decay, when fine-tuning model for query |
--force_retrain | boolean_true | If true, retrain the model and save it to reference_model_path |
--var_input | string | Boolean .var column to subset data with (e.g. containing highly variable genes). By default, do not subset genes. |
TotalVI integration options RNA
Name | Type & Properties | Description |
|---|---|---|
--rna_reference_modality | string | |
--rna_obsm_output | string | In which .obsm slot to store the normalized RNA from TOTALVI. |
TotalVI integration options ADT
Name | Type & Properties | Description |
|---|---|---|
--prot_reference_modality | string | Name of the modality containing proteins in the reference |
--prot_obsm_output | string | In which .obsm slot to store the normalized protein data from TOTALVI. |
Neighbour calculation RNA
Name | Type & Properties | Description |
|---|---|---|
--rna_uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--rna_obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--rna_obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Neighbour calculation ADT
Name | Type & Properties | Description |
|---|---|---|
--prot_uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--prot_obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--prot_obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Clustering options RNA
Name | Type & Properties | Description |
|---|---|---|
--rna_obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions resolutions specified in '--leiden_resolution'. |
--rna_leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Clustering options ADT
Name | Type & Properties | Description |
|---|---|---|
--prot_obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions resolutions specified in '--leiden_resolution'. |
--prot_leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Umap options
Name | Type & Properties | Description |
|---|---|---|
--obsm_umap | string | In which .obsm slot to store the resulting UMAP embedding. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
prot_modality: [ "prot" ]
output: "$id.$key.output.h5mu"
reference_model_path: "$id.$key.reference_model_path"
query_model_path: "$id.$key.query_model_path"
obs_batch: [ "sample_id" ]
max_epochs: [ 400 ]
max_query_epochs: [ 200 ]
weight_decay: [ 0 ]
rna_reference_modality: [ "rna" ]
rna_obsm_output: [ "X_totalvi" ]
prot_reference_modality: [ "prot" ]
prot_obsm_output: [ "X_totalvi" ]
rna_uns_neighbors: [ "totalvi_integration_neighbors" ]
rna_obsp_neighbor_distances: [ "totalvi_integration_distances" ]
rna_obsp_neighbor_connectivities: [ "totalvi_integration_connectivities" ]
prot_uns_neighbors: [ "totalvi_integration_neighbors" ]
prot_obsp_neighbor_distances: [ "totalvi_integration_distances" ]
prot_obsp_neighbor_connectivities: [ "totalvi_integration_connectivities" ]
rna_obs_cluster: [ "totalvi_integration_leiden" ]
rna_leiden_resolution: [ 1 ]
prot_obs_cluster: [ "totalvi_integration_leiden" ]
prot_leiden_resolution: [ 1 ]
obsm_umap: [ "X_totalvi_umap" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/integration/totalvi_scarches_leiden/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/integration/totalvi_scarches_leidenopenpipeline v4.0.0