feature_annotation/highly_variable_features_scanpy
Description
Annotate highly variable features [Satija15] [Zheng17] [Stuart19].
Expects logarithmized data, except when flavor='seurat_v3' in which count data is expected.
Depending on flavor, this reproduces the R-implementations of Seurat [Satija15], Cell Ranger [Zheng17], and Seurat v3 [Stuart19].
For the dispersion-based methods ([Satija15] and [Zheng17]), the normalized dispersion is obtained by scaling with the mean and standard deviation of the dispersions for features falling into a given bin for mean expression of features. This means that for each bin of mean expression, highly variable features are selected.
For [Stuart19], a normalized variance for each feature is computed. First, the data are standardized (i.e., z-score normalization per feature) with a regularized standard deviation. Next, the normalized variance is computed as the variance of each feature after the transformation. Features are ranked by the normalized variance.
Arguments
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string | Which modality from the input MuData file to process. |
--layer | string | use adata.layers[layer] for expression values instead of adata.X. |
--var_input | string | If specified, use boolean array in adata.var[var_input] to calculate hvg on subset of vars. |
--output | file output | Output h5mu file. |
--var_name_filter | string | In which .var slot to store a boolean array corresponding to which observations should be filtered out. |
--varm_name | string | In which .varm slot to store additional metadata. |
--flavor | string | Choose the flavor for identifying highly variable features. For the dispersion based methods in their default workflows, Seurat passes the cutoffs whereas Cell Ranger passes n_top_features. |
--n_top_features | integer | Number of highly-variable features to keep. Mandatory if flavor='seurat_v3'. |
--min_mean | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Ignored if flavor='seurat_v3'. |
--max_mean | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Ignored if flavor='seurat_v3'. |
--min_disp | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Ignored if flavor='seurat_v3'. |
--max_disp | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Ignored if flavor='seurat_v3'. Default is +inf. |
--span | double | The fraction of the data (cells) used when estimating the variance in the loess model fit if flavor='seurat_v3'. |
--n_bins | integer | Number of bins for binning the mean feature expression. Normalization is done with respect to each bin. If just a single feature falls into a bin, the normalized dispersion is artificially set to 1. |
--obs_batch_key | string | If specified, highly-variable features are selected within each batch separately and merged. This simple process avoids the selection of batch-specific features and acts as a lightweight batch correction method. For all flavors, features are first sorted by how many batches they are a HVG. For dispersion-based flavors ties are broken by normalized dispersion. If flavor = 'seurat_v3', ties are broken by the median (across batches) rank based on within-batch normalized variance. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
var_name_filter: [ "filter_with_hvg" ]
varm_name: [ "hvg" ]
flavor: [ "seurat" ]
min_mean: [ 0.0125 ]
max_mean: [ 3 ]
min_disp: [ 0.5 ]
span: [ 0.3 ]
n_bins: [ 20 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/feature_annotation/highly_variable_features_scanpy/main.nf \
-params-file params.yaml Relationships
Used by
Current component
Uses
No component dependencies found.