workflows/annotation/scgpt_integration_knn
Description
Cell type annotation workflow that performs scGPT integration of reference and query dataset followed by KNN label transfer.
Query Input
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Input dataset consisting of the (unlabeled) query observations. The dataset is expected to be pre-processed in the same way as --reference. |
--modality | string | Which modality to process. Should match the modality of the --reference dataset. |
--input_layer | string | Mudata layer (key from layers) to use as input data for scGPT integration; if not specified, X is used. Should match the layer name of the reference dataset. |
--input_var_gene_names | string | The .var field in the input (query) dataset containing gene names; if not provided, the .var index will be used. |
--input_obs_batch_label | string required | The .obs field in the input (query) dataset containing the batch labels. |
--overwrite_existing_key | boolean_true | If provided, will overwrite existing fields in the input dataset when data are copied during the reference alignment process. |
Reference input
Name | Type & Properties | Description |
|---|---|---|
--reference | file required | Reference dataset consisting of observations with cell type labels present in the .obs --reference_obs_batch_label column to train the classifier on. The dataset is expected to be pre-processed in the same way as the --input query dataset(s). |
--reference_obs_targets | string required multiple | The `.obs` key(s) of the target labels to tranfer. |
--reference_var_gene_names | string | The .var field in the reference dataset containing gene names; if not provided, the .var index will be used. |
--reference_obs_batch_label | string required | The .obs field in the reference dataset containing the batch labels. |
scGPT model
Name | Type & Properties | Description |
|---|---|---|
--model | file required | The scGPT model file. Can either be a foundation model or a fine-tuned model. If the model file is a fine-tuned model, it must contain a key for the checkpoints (--finetuned_checkpoints_key). |
--model_vocab | file required | The scGPT model vocabulary file. |
--model_config | file required | The scGPT model configuration file. |
--finetuned_checkpoints_key | string | Key in the model file containing the pretrained checkpoints. Must be provided when `--model` is a fine-tuned model. |
Padding arguments
Name | Type & Properties | Description |
|---|---|---|
--pad_token | string | Token used for padding. |
--pad_value | integer | The value of the padding token. |
HVG subset arguments
Name | Type & Properties | Description |
|---|---|---|
--n_hvg | integer | Number of highly variable genes to subset for. |
Tokenization arguments
Name | Type & Properties | Description |
|---|---|---|
--max_seq_len | integer | The maximum sequence length of the tokenized data. Defaults to the number of features if not provided. |
Embedding arguments
Name | Type & Properties | Description |
|---|---|---|
--dsbn | boolean | Apply domain-specific batch normalization |
--batch_size | integer | The batch size to be used for embedding inference. |
Binning arguments
Name | Type & Properties | Description |
|---|---|---|
--n_input_bins | integer | The number of bins to discretize the data into; When no value is provided, data won't be binned. |
--seed | integer | Seed for random number generation used for binning. If not set, no seed is used. |
Leiden clustering options
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Neighbor classifier arguments
Name | Type & Properties | Description |
|---|---|---|
--weights | string | Weight function used in prediction. Possible values are: `uniform` (all points in each neighborhood are weighted equally) or `distance` (weight points by the inverse of their distance) |
--n_neighbors | integer | The number of neighbors to use in k-neighbor graph structure used for fast approximate nearest neighbor search with PyNNDescent. Larger values will result in more accurate search results at the cost of computation time. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels predicted from the classifier trained on the reference. |
--output_obs_predictions | string multiple | In which `.obs` slots to store the predicted information. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_pred"` suffix. |
--output_obs_probability | string multiple | In which `.obs` slots to store the probability of the predictions. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_probability"` suffix. |
--output_obsm_integrated | string | In which .obsm slot to store the integrated embedding. |
--output_compression | string | The compression format to be used on the output h5mu object. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
pad_token: [ "<pad>" ]
pad_value: [ -2 ]
n_hvg: [ 1200 ]
dsbn: [ true ]
batch_size: [ 64 ]
n_input_bins: [ 51 ]
leiden_resolution: [ 1 ]
weights: [ "uniform" ]
n_neighbors: [ 15 ]
output: "$id.$key.output.h5mu"
output_obsm_integrated: [ "X_integrated_scgpt" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision 2.0.0 \
-main-script target/nextflow/workflows/annotation/scgpt_integration_knn/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/annotation/scgpt_integration_knnopenpipeline 2.0.0