scgpt/cell_type_annotation
Description
Annotate gene expression data with cell type classes through the scGPT model.
Model input
Name | Type & Properties | Description |
|---|---|---|
--model | file required | The model file containing checkpoints and cell type label mapper. |
--model_config | file required | The model configuration file. |
--model_vocab | file required | Model vocabulary file directory. |
--finetuned_checkpoints_key | string | Key in the model file containing the pretrained checkpoints. |
--label_mapper_key | string | Key in the model file containing the cell type class to label mapper dictionary. |
Query input
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The input h5mu file containing of data that have been pre-processed (normalized, binned, genes cross-checked and tokenized). |
--modality | string | |
--obs_batch_label | string | The name of the adata.obs column containing the batch labels. Required if dsbn is set to true. |
--obsm_gene_tokens | string | The key of the .obsm array containing the gene token ids |
--obsm_tokenized_values | string | The key of the .obsm array containing the count values of the tokenized genes |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The output mudata file. |
--output_compression | string | The compression algorithm to use for the output h5mu file. |
--output_obs_predictions | string | The name of the adata.obs column to write predicted cell type labels to. |
--output_obs_probability | string | The name of the adata.obs column to write the probabilities of the predicted cell type labels to. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--pad_token | string | The padding token used in the model. |
--pad_value | integer | The value of the padding. |
--n_input_bins | integer | The number of input bins. |
--batch_size | integer | The batch size. |
--dsbn | boolean | Whether to use domain-specific batch normalization. |
--seed | integer | Seed for random number generation. If not specified, no seed is used. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
finetuned_checkpoints_key: [ "model_state_dict" ]
label_mapper_key: [ "id_to_class" ]
modality: [ "rna" ]
obsm_gene_tokens: [ "gene_id_tokens" ]
obsm_tokenized_values: [ "values_tokenized" ]
output: "$id.$key.output.h5mu"
output_compression: [ "gzip" ]
output_obs_predictions: [ "scgpt_pred" ]
output_obs_probability: [ "scgpt_probability" ]
pad_token: [ "<pad>" ]
pad_value: [ -2 ]
n_input_bins: [ 51 ]
batch_size: [ 64 ]
dsbn: [ true ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision 2.0.0 \
-main-script target/nextflow/scgpt/cell_type_annotation/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
scgpt/cell_type_annotationopenpipeline 2.0.0
Uses
0 relationships
No component dependencies found.