integrate/scvi
Description
Performs scvi integration as done in the human lung cell atlas https://github.com/LungCellAtlas/HLCA
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | Input layer to use. If None, X is used |
--obs_batch | string | Column name discriminating between your batches. |
--var_gene_names | string | .var column containing gene names. By default, use the index. |
--var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
--obs_labels | string | Key in adata.obs for label information. Categories will automatically be converted into integer categories and saved to adata.obs['_scvi_labels']. If None, assigns the same label to all the data. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--output_model | file output | Folder where the state of the trained model will be saved to. |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
SCVI options
Name | Type & Properties | Description |
|---|---|---|
--n_hidden_nodes | integer | Number of nodes per hidden layer. |
--n_dimensions_latent_space | integer | Dimensionality of the latent space. |
--n_hidden_layers | integer | Number of hidden layers used for encoder and decoder neural-networks. |
--dropout_rate | double | Dropout rate for the neural networks. |
--dispersion | string | Set the behavior for the dispersion for negative binomial distributions: - gene: dispersion parameter of negative binomial is constant per gene across cells - gene-batch: dispersion can differ between different batches - gene-label: dispersion can differ between different labels - gene-cell: dispersion can differ for every gene in every cell |
--gene_likelihood | string | Model used to generate the expression data from a count-based likelihood distribution. - nb: Negative binomial distribution - zinb: Zero-inflated negative binomial distribution - poisson: Poisson distribution |
Variational auto-encoder model options
Name | Type & Properties | Description |
|---|---|---|
--use_layer_normalization | string | Neural networks for which to enable layer normalization. |
--use_batch_normalization | string | Neural networks for which to enable batch normalization. |
--encode_covariates | boolean_false | Whether to concatenate covariates to expression in encoder |
--deeply_inject_covariates | boolean_true | Whether to concatenate covariates into output of hidden layers in encoder/decoder. This option only applies when n_layers > 1. The covariates are concatenated to the input of subsequent hidden layers. |
--use_observed_lib_size | boolean_true | Use observed library size for RNA as scaling factor in mean of conditional distribution. |
Early stopping arguments
Name | Type & Properties | Description |
|---|---|---|
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
Learning parameters
Name | Type & Properties | Description |
|---|---|---|
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Data validition
Name | Type & Properties | Description |
|---|---|---|
--n_obs_min_count | integer | Minimum number of cells threshold ensuring that every obs_batch category has sufficient observations (cells) for model training. |
--n_var_min_count | integer | Minimum number of genes threshold ensuring that every var_input filter has sufficient observations (genes) for model training. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
obs_batch: [ "sample_id" ]
sanitize_ensembl_ids: [ true ]
output: "$id.$key.output"
output_model: "$id.$key.output_model"
obsm_output: [ "X_scvi_integrated" ]
n_hidden_nodes: [ 128 ]
n_dimensions_latent_space: [ 30 ]
n_hidden_layers: [ 2 ]
dropout_rate: [ 0.1 ]
dispersion: [ "gene" ]
gene_likelihood: [ "nb" ]
use_layer_normalization: [ "both" ]
use_batch_normalization: [ "none" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
n_obs_min_count: [ 0 ]
n_var_min_count: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/integrate/scvi/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
integrate/scviopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.