workflows/integration/scgpt_leiden
Description
Run scGPT integration (cell embedding generation) followed by neighbour calculations, leiden clustering and run umap on the result.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Path to the input file. |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | The layer of the input dataset to process if .X is not to be used. Should contain log normalized counts. |
--var_gene_names | string | The name of the adata var column containing gene names; when no gene_name_layer is provided, the var index will be used. |
--obs_batch_label | string | The name of the adata obs column containing the batch labels. |
Model
Name | Type & Properties | Description |
|---|---|---|
--model | file required | Path to scGPT model file. |
--model_vocab | file required | Path to scGPT model vocabulary file. |
--model_config | file required | Path to scGPT model config file. |
--finetuned_checkpoints_key | string | Key in the model file containing the pretrained checkpoints. Only relevant for fine-tuned models. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output file path |
--obsm_integrated | string | In which .obsm slot to store the resulting integrated embedding. |
Padding arguments
Name | Type & Properties | Description |
|---|---|---|
--pad_token | string | Token used for padding. |
--pad_value | integer | The value of the padding token. |
HVG subset arguments
Name | Type & Properties | Description |
|---|---|---|
--n_hvg | integer | Number of highly variable genes to subset for. |
--hvg_flavor | string | Method to be used for identifying highly variable genes. Note that the default for this workflow (`cell_ranger`) is not the default method for scanpy hvg detection (`seurat`). |
Tokenization arguments
Name | Type & Properties | Description |
|---|---|---|
--max_seq_len | integer | The maximum sequence length of the tokenized data. Defaults to the number of features if not provided. |
Embedding arguments
Name | Type & Properties | Description |
|---|---|---|
--dsbn | boolean | Apply domain-specific batch normalization |
--batch_size | integer | The batch size to be used for embedding inference. |
Binning arguments
Name | Type & Properties | Description |
|---|---|---|
--n_input_bins | integer | The number of bins to discretize the data into; When no value is provided, data won't be binned. |
--seed | integer | Seed for random number generation used for binning. If not set, no seed is used. |
Clustering arguments
Name | Type & Properties | Description |
|---|---|---|
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
obsm_integrated: [ "X_scgpt" ]
pad_token: [ "<pad>" ]
pad_value: [ -2 ]
n_hvg: [ 1200 ]
hvg_flavor: [ "cell_ranger" ]
dsbn: [ true ]
batch_size: [ 64 ]
n_input_bins: [ 51 ]
leiden_resolution: [ 1 ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v3.0.0 \
-main-script target/nextflow/workflows/integration/scgpt_leiden/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/integration/scgpt_leidenopenpipeline v3.0.0