annotate/scanvi
Description
scANVI () is a semi-supervised model for single-cell transcriptomics data. scANVI is an scVI extension that can leverage the cell type knowledge for a subset of the cells present in the data sets to infer the states of the rest of the cells.
This component will instantiate a scANVI model from a pre-trained scVI model, integrate the data and perform label prediction.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file. Note that this needs to be the exact same dataset as the --scvi_model was trained on. |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | Input layer to use. If None, X is used |
--var_input | string | .var column containing highly variable genes that were used to train the scVi model. By default, do not subset genes. |
--var_gene_names | string | .var column containing gene names. By default, use the index. |
--obs_labels | string required | .obs field containing the labels |
--unlabeled_category | string | Value in the --obs_labels field that indicates unlabeled observations |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
scVI Model
Name | Type & Properties | Description |
|---|---|---|
--scvi_model | file required | Pretrained SCVI reference model to initialize the SCANVI model with. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--output_model | file output | Folder where the state of the trained model will be saved to. |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--obs_output_predictions | string | In which .obs slot to store the predicted labels. |
--obs_output_probabilities | string | In which. obs slot to store the probabilities of the predicted labels. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
scANVI training arguments
Name | Type & Properties | Description |
|---|---|---|
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
unlabeled_category: [ "Unknown" ]
sanitize_ensembl_ids: [ true ]
output: "$id.$key.output"
output_model: "$id.$key.output_model"
obsm_output: [ "X_scanvi_integrated" ]
obs_output_predictions: [ "scanvi_pred" ]
obs_output_probabilities: [ "scanvi_proba" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/annotate/scanvi/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
annotate/scanviopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.