workflows/annotation/scvi_knn
Description
"Cell type annotation workflow that performs scVI integration of reference and query dataset followed by KNN label transfer.
The query and reference datasets are expected to be pre-processed in the same way, for example with the process_samples workflow of OpenPipeline.
Note that this workflow will integrate the reference dataset from scratch and integrate the query dataset in the same embedding space.
The workflow does not currently output the trained SCVI reference model."
Query Input
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input dataset consisting of the (unlabeled) query observations. |
--modality | string | Which modality to process. Should match the modality of the --reference dataset. |
--input_layer | string | The layer of the input dataset containing the raw counts if .X is not to be used. |
--input_layer_lognormalized | string | The layer of the input dataset containing the lognormalized counts if .X is not to be used. |
--input_obs_batch_label | string required | The .obs field in the input (query) dataset containing the batch labels. |
--input_var_gene_names | string | The .var field in the input (query) dataset containing gene names; if not provided, the .var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
--overwrite_existing_key | boolean_true | If provided, will overwrite existing fields in the input dataset when data are copied during the reference alignment process. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Reference input
Name | Type & Properties | Description |
|---|---|---|
--reference | file required | Reference dataset consisting of the labeled observations. |
--reference_layer | string | The layer of the reference dataset containing the raw counts if .X is not to be used. |
--reference_layer_lognormalized | string | The layer of the reference dataset containing the lognormalized counts if .X is not to be used. |
--reference_obs_target | string required | The `.obs` key(s) of the target labels to transfer. |
--reference_var_gene_names | string | The .var field in the reference dataset containing gene names; if not provided, the .var index will be used. |
--reference_obs_batch_label | string required | The .obs field in the reference dataset containing the batch labels. |
HVG subset arguments
Name | Type & Properties | Description |
|---|---|---|
--n_hvg | integer | Number of highly variable genes to subset for. |
scVI integration options
Name | Type & Properties | Description |
|---|---|---|
--scvi_early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--scvi_early_stopping_monitor | string | Metric logged during validation set epoch. |
--scvi_early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--scvi_early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
--scvi_max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--scvi_reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--scvi_lr_factor | double | Factor to reduce learning rate. |
--scvi_lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Leiden clustering options
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Neighbor classifier arguments
Name | Type & Properties | Description |
|---|---|---|
--knn_weights | string | Weight function used in prediction. Possible values are: `uniform` (all points in each neighborhood are weighted equally) or `distance` (weight points by the inverse of their distance) |
--knn_n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels predicted from the classifier trained on the reference. |
--output_obs_predictions | string multiple | In which `.obs` slots to store the predicted cell labels. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_pred"` suffix. |
--output_obs_probability | string multiple | In which `.obs` slots to store the probability of the predictions. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_probability"` suffix. |
--output_obsm_integrated | string | In which .obsm slot to store the integrated embedding. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
sanitize_ensembl_ids: [ true ]
n_hvg: [ 2000 ]
scvi_early_stopping_monitor: [ "elbo_validation" ]
scvi_early_stopping_patience: [ 45 ]
scvi_early_stopping_min_delta: [ 0 ]
scvi_reduce_lr_on_plateau: [ true ]
scvi_lr_factor: [ 0.6 ]
scvi_lr_patience: [ 30 ]
leiden_resolution: [ 1 ]
knn_weights: [ "uniform" ]
knn_n_neighbors: [ 15 ]
output: "$id.$key.output.h5mu"
output_obsm_integrated: [ "X_integrated_scvi" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/annotation/scvi_knn/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/annotation/scvi_knnopenpipeline v4.0.0