workflows/annotation/scanvi_scarches
Description
Cell type annotation workflow using ScanVI with scArches for reference mapping.
Query Input
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input dataset consisting of the (unlabeled) query observations. The dataset is expected to be pre-processed in the same way as --reference. |
--modality | string | Which modality to process. Should match the modality of the --reference dataset. |
--layer | string | Which layer to use for integration if .X is not to be used. Should match the layer of the --reference dataset. |
--input_obs_batch_label | string required | The .obs field in the input (query) dataset containing the batch labels. |
--input_obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--input_obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the categorical covariates matters and should match the order of the covariates in the reference data. |
--input_obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the continuous covariates matters and should match the order of the covariates in the reference data. |
--input_var_gene_names | string | .var column containing gene names. By default, use the index. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Reference input
Name | Type & Properties | Description |
|---|---|---|
--reference | file required | Reference dataset consisting of the labeled observations to train the KNN classifier on. The dataset is expected to be pre-processed in the same way as the --input query dataset. |
--reference_obs_target | string required | The `.obs` key containing the target labels. |
--reference_obs_batch_label | string required | The .obs field in the reference dataset containing the batch labels. |
--reference_obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--reference_obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the categorical covariates matters and should match the order of the covariates in the query data. |
--reference_obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the continuous covariates matters and should match the order of the covariates in the query data. |
--unlabeled_category | string | Value in the --reference_obs_batch_label field that indicates unlabeled observations |
--reference_var_hvg | string | .var column containing highly variable genes. If not provided, genes will not be subset. |
--reference_var_gene_names | string | .var column containing gene names. By default, use the index. |
scVI, scANVI and scArches training options
Name | Type & Properties | Description |
|---|---|---|
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Leiden clustering options
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Neighbor classifier arguments
Name | Type & Properties | Description |
|---|---|---|
--knn_weights | string | Weight function used in prediction. Possible values are: `uniform` (all points in each neighborhood are weighted equally) or `distance` (weight points by the inverse of their distance) |
--knn_n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels predicted from the classifier trained on the reference. |
--output_obs_predictions | string | In which .obs slot to store the predicted labels. |
--output_obs_probability | string | In which. obs slot to store the probabilities of the predicted labels. |
--output_obsm_integrated | string | In which .obsm slot to store the integrated embedding. |
--output_compression | string | The compression format to be used on the output h5mu object. |
--output_model | file output | Path to the resulting scANVI model that was updated with query data. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
sanitize_ensembl_ids: [ true ]
unlabeled_category: [ "Unknown" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
leiden_resolution: [ 1 ]
knn_weights: [ "uniform" ]
knn_n_neighbors: [ 15 ]
output: "$id.$key.output.h5mu"
output_obs_predictions: [ "scanvi_pred" ]
output_obs_probability: [ "scanvi_probabilities" ]
output_obsm_integrated: [ "X_integrated_scanvi" ]
output_model: "$id.$key.output_model"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/annotation/scanvi_scarches/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/annotation/scanvi_scarchesopenpipeline v4.0.0