annotate/singler
Description
SingleR performs reference-based cell type annotation for single-cell RNA-seq data
by computing Spearman correlations between test cells and reference samples with known labels,
using marker genes to assign the most similar cell type label to each new cell.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. |
--input_layer | string | The layer in the input data containing log normalized counts to be used for cell type annotation if .X is not to be used. |
--input_var_gene_names | string | The name of the adata .var column in the input data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_obs_clusters | string | The name of the adata .obs column containing cluster identities of the observations. If provided, annoation is performed on the aggregated cluster profiles, otherwise it defaults to annotation per observation. |
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference | file required | The reference data to train the CellTypist classifiers on. Only required if a pre-trained --model is not provided. |
--reference_layer | string | The layer in the reference data containing lognormalized couns to be used for cell type annotation if .X is not to be used. Data are expected to be processed in the same way as the --input query dataset. |
--reference_obs_target | string required | The name of the adata obs column in the reference data containing cell type annotations. |
--reference_var_gene_names | string | The name of the adata var column in the reference data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--reference_var_input | string | .var column containing a boolean mask corresponding to genes to be used for marker selection. By default, do not subset genes. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--de_n_genes | integer | The number of differentially expressed genes across labels to be calculated from the reference. Defaults to 500 * (2/3) ^ log2(N) where N is the number of unique labels when if `--de_method` is set to `classic`, otherwise, defaults to 10. |
--de_method | string | Method to detect differentially expressed genes between pairs of labels. |
--quantile | double | The quantile of the correlation distribution to use to compute the score per label. |
--fine_tune | boolean | Whether finetuning should be performed to improve the resolution. If set to True, an additional finetuning step is performed after initial classification, new marker genes are calculated based on all cells with a score higher then the maximum score minus `--fine_tuning_thershold`, and the calculation of the scores is repeated. |
--fine_tuning_threshold | double | The maximum difference from the maximum correlation to use in fine-tuning |
--prune | boolean | Whether label pruning should be performed. If set to True, an additional output .obs field `--output_obs_pruned_predictions` will be added to the `--output`, containing labels where 'low-quality' labels are replaced with NA's. Labels are considered 'low-quality' when their delta score (stored in `--output_obs_delta_next`) fall more than 3 median absolute deviations below the median for that label type. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file output | Output h5mu file. |
--output_obs_predictions | string | In which `.obs` slot to store the predicted labels. If `--fine_tune False`, this is based only on the maximum entry in `--output_obsm_scores`. |
--output_obs_probability | string | In which `.obs` slots to store the probability of the predicted labels. |
--output_obs_delta_next | string | In which `.obs` slot to store the delta between the best and next-best score. If `--fine_tune True`, this is reported for scores after fine-tuning. |
--output_obs_pruned_predictions | string | In which `.obs` slot to store the pruned labels, where low-quality labels are replaced with NA's. Only added if `--prune True`. |
--output_obsm_scores | string | In which `.obsm` slot to store the matrix of prediction correlations at the specified quantile for each label (column) in each cell (row). |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
input_reference_gene_overlap: [ 100 ]
de_method: [ "classic" ]
quantile: [ 0.8 ]
fine_tune: [ true ]
fine_tuning_threshold: [ 0.05 ]
prune: [ true ]
sanitize_ensembl_ids: [ true ]
output: "$id.$key.output.h5mu"
output_obs_predictions: [ "singler_pred" ]
output_obs_probability: [ "singler_probability" ]
output_obs_delta_next: [ "singler_delta_next" ]
output_obs_pruned_predictions: [ "singler_pruned_labels" ]
output_obsm_scores: [ "singler_scores" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.2 \
-main-script target/nextflow/annotate/singler/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
annotate/singleropenpipeline v4.0.2
Uses
0 relationships
No component dependencies found.