workflows/integration/scvi_leiden
Description
Run scvi integration followed by neighbour calculations, leiden clustering and run umap on the result.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file | Path to the sample. Mutually exclusive with the 'input_uri' argument. |
--tiledb_input_uri | string | A URI containing the TileDB-SOMA objects. Mutually exclusive with the 'input' argument. |
--tiledb_s3_region | string | Region where the TileDB-SOMA database is hosted. |
--tiledb_endpoint | string | Custom endpoint to use to connect to S3 |
--tiledb_s3_no_sign_request | boolean | Do not sign S3 requests. Credentials will not be loaded if this argument is provided. |
Input slots
Name | Type & Properties | Description |
|---|---|---|
--layer | string | use specified layer for expression values instead of the .X object from the modality. |
--modality | string | Which modality to process. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
--output_model | file required output | Folder where the state of the trained model will be saved to. |
--output_tiledb | file output | Output the TileDB database to the specified directory instead of adding it to the existing database. |
Neighbour calculation
Name | Type & Properties | Description |
|---|---|---|
--uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Scvi integration options
Name | Type & Properties | Description |
|---|---|---|
--obs_batch | string required | Column name discriminating between your batches. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Clustering options
Name | Type & Properties | Description |
|---|---|---|
--obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions resolutions specified in '--leiden_resolution'. |
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Umap options
Name | Type & Properties | Description |
|---|---|---|
--obsm_umap | string | In which .obsm slot to store the resulting UMAP embedding. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
tiledb_s3_no_sign_request: [ false ]
modality: [ "rna" ]
sanitize_ensembl_ids: [ true ]
output: "$id.$key.output.h5mu"
output_model: "$id.$key.output_model.output_dir"
output_tiledb: "$id.$key.output_tiledb"
uns_neighbors: [ "scvi_integration_neighbors" ]
obsp_neighbor_distances: [ "scvi_integration_distances" ]
obsp_neighbor_connectivities: [ "scvi_integration_connectivities" ]
obsm_output: [ "X_scvi_integrated" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
obs_cluster: [ "scvi_integration_leiden" ]
leiden_resolution: [ 1 ]
obsm_umap: [ "X_scvi_umap" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/integration/scvi_leiden/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
workflows/integration/scvi_leidenopenpipeline v4.0.0