annotate/celltypist
Description
Automated cell type annotation tool for scRNA-seq datasets on the basis of logistic regression classifiers optimised by the stochastic gradient descent algorithm.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. |
--input_layer | string | The layer in the input data containing counts that are lognormalized to 10000, .X is not to be used. |
--input_var_gene_names | string | The name of the adata var column in the input data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train the CellTypist classifiers on. Only required if a pre-trained --model is not provided. |
--reference_layer | string | The layer in the reference data containing counts that are lognormalized to 10000, if .X is not to be used. |
--reference_obs_target | string | The name of the adata obs column in the reference data containing cell type annotations. |
--reference_var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
--reference_var_gene_names | string | The name of the adata var column in the reference data containing gene names; when no gene_name_layer is provided, the var index will be used. |
Model arguments
Name | Type & Properties | Description |
|---|---|---|
--model | file | Pretrained model in pkl format. If not provided, the model will be trained on the reference data and --reference should be provided. |
--feature_selection | boolean | Whether to perform feature selection. |
--majority_voting | boolean | Whether to refine the predicted labels by running the majority voting classifier after over-clustering. |
--C | double | Inverse of regularization strength in logistic regression. |
--max_iter | integer | Maximum number of iterations before reaching the minimum of the cost function. |
--use_SGD | boolean_true | Whether to use the stochastic gradient descent algorithm. |
--min_prop | double | "For the dominant cell type within a subcluster, the minimum proportion of cells required to support naming of the subcluster by this cell type. Ignored if majority_voting is set to False. Subcluster that fails to pass this proportion threshold will be assigned 'Heterogeneous'." |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file output | Output h5mu file. |
--output_obs_predictions | string | In which `.obs` slots to store the predicted information. |
--output_obs_probability | string | In which `.obs` slots to store the probability of the predictions. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
sanitize_ensembl_ids: [ true ]
reference_obs_target: [ "cell_ontology_class" ]
feature_selection: [ false ]
majority_voting: [ false ]
C: [ 1 ]
max_iter: [ 1000 ]
min_prop: [ 0 ]
output: "$id.$key.output.h5mu"
output_obs_predictions: [ "celltypist_pred" ]
output_obs_probability: [ "celltypist_probability" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/annotate/celltypist/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
annotate/celltypistopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.