workflows/annotation/harmony_knn
Description
Cell type annotation workflow by performing harmony integration of reference and query dataset followed by KNN label transfer.
Query Input
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input dataset consisting of the (unlabeled) query observations. The dataset is expected to be pre-processed in the same way as --reference. |
--modality | string | Which modality to process. Should match the modality of the --reference dataset. |
--input_layer | string | The layer of the input dataset to process if .X is not to be used. Should contain log normalized counts. |
--input_obs_batch_label | string required | The .obs field in the input (query) dataset containing the batch labels. |
--input_var_gene_names | string | The .var field in the input (query) dataset containing gene names; if not provided, the .var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
--overwrite_existing_key | boolean_true | If provided, will overwrite existing fields in the input dataset when data are copied during the reference alignment process. |
Reference input
Name | Type & Properties | Description |
|---|---|---|
--reference | file required | Reference dataset consisting of the labeled observations to train the KNN classifier on. The dataset is expected to be pre-processed in the same way as the --input query dataset. |
--reference_layer | string | The layer of the reference dataset to process if .X is not to be used. Should contain log normalized counts. |
--reference_obs_target | string required | The `.obs` key of the target cell type labels to transfer. |
--reference_var_gene_names | string | The .var field in the reference dataset containing gene names; if not provided, the .var index will be used. |
--reference_obs_batch_label | string required | The .obs field in the reference dataset containing the batch labels. |
HVG subset arguments
Name | Type & Properties | Description |
|---|---|---|
--n_hvg | integer | Number of highly variable genes to subset for. |
PCA options
Name | Type & Properties | Description |
|---|---|---|
--pca_num_components | integer | Number of principal components to compute. Defaults to 50, or 1 - minimum dimension size of selected representation. |
Harmony integration options
Name | Type & Properties | Description |
|---|---|---|
--harmony_theta | double multiple | Diversity clustering penalty parameter. Can be set as a single value for all batch observations or as multiple values, one for each observation in the batches defined by --input_obs_batch_label. theta=0 does not encourage any diversity. Larger values of theta result in more diverse clusters." |
Leiden clustering options
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Neighbor classifier arguments
Name | Type & Properties | Description |
|---|---|---|
--knn_weights | string | Weight function used in prediction. Possible values are: `uniform` (all points in each neighborhood are weighted equally) or `distance` (weight points by the inverse of their distance) |
--knn_n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels predicted from the classifier trained on the reference. |
--output_obs_predictions | string multiple | In which `.obs` slots to store the predicted cell labels. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_pred"` suffix. |
--output_obs_probability | string multiple | In which `.obs` slots to store the probability of the predictions. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_probability"` suffix. |
--output_obsm_integrated | string | In which .obsm slot to store the integrated embedding. |
--output_compression | string | The compression format to be used on the output h5mu object. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
n_hvg: [ 2000 ]
harmony_theta: [ 2 ]
leiden_resolution: [ 1 ]
knn_weights: [ "uniform" ]
knn_n_neighbors: [ 15 ]
output: "$id.$key.output.h5mu"
output_obsm_integrated: [ "X_integrated_harmony" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/annotation/harmony_knn/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/annotation/harmony_knnopenpipeline v4.0.0