scgpt/embedding
Description
Generation of cell embeddings for the integration of single cell transcriptomic count data using scGPT.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The input h5mu file containing tokenized gene and count data. |
--modality | string | Which modality from the input MuData file to process. |
--model | file required | Path to scGPT model file. |
--model_vocab | file required | Path to scGPT model vocabulary file. |
--model_config | file required | Path to scGPT model config file. |
--obsm_gene_tokens | string | The key of the .obsm array containing the gene token ids |
--obsm_tokenized_values | string | The key of the .obsm array containing the count values of the tokenized genes |
--obsm_padding_mask | string | The key of the .obsm array containing the padding mask. |
--var_gene_names | string | The name of the .var column containing gene names. When no gene_name_layer is provided, the .var index will be used. |
--obs_batch_label | string | The name of the adata.obs column containing the batch labels. Must be provided when 'dsbn' is set to True. |
--finetuned_checkpoints_key | string | Key in the model file containing the pretrained checkpoints. Only relevant for fine-tuned models. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Path to output anndata file containing pre-processed data as well as scGPT embeddings. |
--obsm_embeddings | string | The name of the adata.obsm array to which scGPT embeddings will be written. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--pad_token | string | The token to be used for padding. |
--pad_value | integer | The value of the padding token. |
--dsbn | boolean | Whether to apply domain-specific batch normalization for generating embeddings. When set to True, 'obs_batch_labels' must be set as well. |
--batch_size | integer | The batch size to be used for inference |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
obsm_gene_tokens: [ "gene_id_tokens" ]
obsm_tokenized_values: [ "values_tokenized" ]
obsm_padding_mask: [ "padding_mask" ]
output: "$id.$key.output.h5mu"
obsm_embeddings: [ "X_scGPT" ]
pad_token: [ "<pad>" ]
pad_value: [ -2 ]
dsbn: [ true ]
batch_size: [ 64 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v3.0.2 \
-main-script target/nextflow/scgpt/embedding/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
scgpt/embeddingopenpipeline v3.0.2
Uses
0 relationships
No component dependencies found.