labels_transfer/knn
Description
This component performs label transfer from reference to query using a K-Neirest Neighbors classifier.
Input dataset (query) arguments
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The query data to transfer the labels to. Should be a .h5mu file. |
--modality | string | Which modality to use. |
--input_obsm_features | string | The `.obsm` key of the embedding to use for the classifier's inference. If not provided, the `.X` slot will be used instead. Make sure that embedding was obtained in the same way as the reference embedding (e.g. by the same model or preprocessing). |
Reference dataset arguments
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train classifiers on. |
--reference_obsm_features | string | The `.obsm` key of the embedding to use for the classifier's training. If not provided, the `.X` slot will be used instead. Make sure that embedding was obtained in the same way as the query embedding (e.g. by the same model or preprocessing). |
--reference_obs_targets | string multiple | The `.obs` key(s) of the target labels to tranfer. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels transfered from the reference. |
--output_obs_predictions | string multiple | In which `.obs` slots to store the predicted information. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_pred"` suffix. |
--output_obs_probability | string multiple | In which `.obs` slots to store the probability of the predictions. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_probability"` suffix. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Input dataset (query) arguments
Name | Type & Properties | Description |
|---|---|---|
--input_obsm_distances | string | The `.obsm` key of the input (query) dataset containing pre-calculated distances. If not provided, the distances will be calculated using PyNNDescent. Make sure the distance matrix contains distances relative to the reference dataset and were obtained in the same way as the reference embedding. |
Reference dataset arguments
Name | Type & Properties | Description |
|---|---|---|
--reference_obsm_distances | string | The `.obsm` key of the reference dataset containing pre-calculated distances. If not provided, the distances will be calculated using PyNNDescent. |
KNN label transfer arguments
Name | Type & Properties | Description |
|---|---|---|
--weights | string | Weight function used in prediction. Possible values are: - `uniform` - all points in each neighborhood are weighted equally - `distance` - weight points by the inverse of their distance - `gaussian` - weight points by the sum of their Gaussian kernel similarities to each sample |
--n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
reference_obs_targets:
[
"ann_level_1",
"ann_level_2",
"ann_level_3",
"ann_level_4",
"ann_level_5",
"ann_finest_level"
]
output: "$id.$key.output"
weights: [ "uniform" ]
n_neighbors: [ 15 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/labels_transfer/knn/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
labels_transfer/knnopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.