annotate/celltypist
Description
Automated cell type annotation tool for scRNA-seq datasets on the basis of logistic regression classifiers optimised by the stochastic gradient descent algorithm.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. |
--input_layer | string | The layer in the input data containing log normalized counts to be used for cell type annotation if .X is not to be used. |
--input_var_gene_names | string | The name of the adata var column in the input data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train the CellTypist classifiers on. Only required if a pre-trained --model is not provided. |
--reference_layer | string | The layer in the reference data to be used for cell type annotation if .X is not to be used. Data are expected to be processed in the same way as the --input query dataset. |
--reference_obs_target | string | The name of the adata obs column in the reference data containing cell type annotations. |
--reference_var_gene_names | string | The name of the adata var column in the reference data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--reference_var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
Model arguments
Name | Type & Properties | Description |
|---|---|---|
--model | file | Pretrained model in pkl format. If not provided, the model will be trained on the reference data and --reference should be provided. |
--feature_selection | boolean | Whether to perform feature selection. |
--majority_voting | boolean | Whether to refine the predicted labels by running the majority voting classifier after over-clustering. |
--C | double | Inverse of regularization strength in logistic regression. |
--max_iter | integer | Maximum number of iterations before reaching the minimum of the cost function. |
--use_SGD | boolean_true | Whether to use the stochastic gradient descent algorithm. |
--min_prop | double | "For the dominant cell type within a subcluster, the minimum proportion of cells required to support naming of the subcluster by this cell type. Ignored if majority_voting is set to False. Subcluster that fails to pass this proportion threshold will be assigned 'Heterogeneous'." |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file output | Output h5mu file. |
--output_obs_predictions | string | In which `.obs` slots to store the predicted information. |
--output_obs_probability | string | In which `.obs` slots to store the probability of the predictions. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
reference_obs_target: [ "cell_ontology_class" ]
feature_selection: [ false ]
majority_voting: [ false ]
C: [ 1 ]
max_iter: [ 1000 ]
min_prop: [ 0 ]
output: "$id.$key.output.h5mu"
output_obs_predictions: [ "celltypist_pred" ]
output_obs_probability: [ "celltypist_probability" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v3.0.0 \
-main-script target/nextflow/annotate/celltypist/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships, 1 components
Current component
annotate/celltypistopenpipeline v3.0.0
Uses
0 relationships
No component dependencies found.