annotate/onclass
Description
OnClass is a python package for single-cell cell type annotation. It uses the Cell Ontology to capture the cell type similarity.
These similarities enable OnClass to annotate cell types that are never seen in the training data.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. |
--input_layer | string | The layer in the input data to be used for cell type annotation if .X is not to be used. |
--input_var_gene_names | string | The name of the adata var column in the input data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Ontology
Name | Type & Properties | Description |
|---|---|---|
--cl_nlp_emb_file | file required | The .nlp.emb file with the cell type embeddings. |
--cl_ontology_file | file required | The .ontology file with the cell type ontology. |
--cl_obo_file | file required | The .obo file with the cell type ontology. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train the CellTypist classifiers on. Only required if a pre-trained --model is not provided. |
--reference_layer | string | The layer in the reference data to be used for cell type annotation if .X is not to be used. |
--reference_obs_target | string required | The name of the adata obs column in the reference data containing cell type annotations. |
--reference_var_gene_names | string | The name of the adata var column in the reference data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--reference_var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
--unknown_celltype | string | Label for unknown cell types. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file output | Output h5mu file. |
--output_obs_predictions | string | In which `.obs` slots to store the predicted information. |
--output_obs_probability | string | In which `.obs` slots to store the probability of the predictions. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata H5 files. By default no compression is applied. |
Model arguments
Name | Type & Properties | Description |
|---|---|---|
--model | string | "Pretrained model path without a file extension. If not provided, the model will be trained on the reference data and --reference should be provided. The path namespace should contain: - a .npz or .pkl file - a .data file - a .meta file - a .index file e.g. /path/to/model/pretrained_model_target1 as saved by OnClass." |
--max_iter | integer | Maximum number of iterations for training the model. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
sanitize_ensembl_ids: [ true ]
unknown_celltype: [ "Unknown" ]
output: "$id.$key.output.h5mu"
output_obs_predictions: [ "onclass_pred" ]
output_obs_probability: [ "onclass_prob" ]
max_iter: [ 30 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/annotate/onclass/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
annotate/onclassopenpipeline v4.2.0
Uses
0 relationships
No component dependencies found.