wrappers/preprocessing/highly_variable_genes
Description
Annotate highly variable genes, running either the GPU (rapids-singlecell) or
the CPU (scanpy) variant of highly_variable_genes, selected with --device_type.
Expects log-normalized data for the default 'seurat' flavor. The GPU variant
exposes additional flavors ('seurat_v3_paper', 'pearson_residuals',
'poisson_gene_selection') that the CPU variant does not support.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | Input layer to use for expression values. By default, X is used. |
Compute
Name | Type & Properties | Description |
|---|---|---|
--device_type | string | Which implementation to run: the GPU (rapids-singlecell) variant or the CPU (scanpy/squidpy) variant of the component. Selecting `gpu` requires a CUDA-capable GPU; the component errors out if none is available (there is no automatic fallback to CPU). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output h5mu file. |
--var_name_filter | string | In which .var slot to store a boolean array indicating which features are highly variable. |
--varm_name | string | In which .varm slot to store the per-gene HVG metrics. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--flavor | string | Choose the flavor for identifying highly variable genes. For the dispersion-based methods in their default workflows, Seurat passes the cutoffs whereas Cell Ranger passes n_top_features. The 'seurat_v3_paper', 'pearson_residuals' and 'poisson_gene_selection' flavors are only available on the GPU (rapids-singlecell) variant. |
--n_top_features | integer | Number of highly-variable features to keep. Mandatory if flavor is 'seurat_v3', 'seurat_v3_paper', 'pearson_residuals' or 'poisson_gene_selection'. |
--min_mean | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Only used for dispersion-based flavors. |
--max_mean | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Only used for dispersion-based flavors. |
--min_disp | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Only used for dispersion-based flavors. |
--max_disp | double | If n_top_features is defined, this and all other cutoffs for the means and the normalized dispersions are ignored. Only used for dispersion-based flavors. Default is +inf. |
--span | double | The fraction of the data (cells) used when estimating the variance in the loess model fit if flavor is 'seurat_v3' or 'seurat_v3_paper'. |
--n_bins | integer | Number of bins for binning the mean gene expression. Normalization is done with respect to each bin. If just a single gene falls into a bin, the normalized dispersion is artificially set to 1. |
--obs_batch_key | string | If specified, highly-variable features are selected within each batch separately and merged. |
rapids-singlecell options
Name | Type & Properties | Description |
|---|---|---|
--theta | integer | The negative binomial overdispersion parameter for Pearson residuals. Higher values correspond to less overdispersion. Only used if flavor is 'pearson_residuals'. |
--clip | double | Determines if and how Pearson residuals are clipped. If unset, residuals are clipped to [-sqrt(n_obs), sqrt(n_obs)]. If a scalar c is given, residuals are clipped to [-c, c]. Only used if flavor is 'pearson_residuals'. |
--chunksize | integer | If flavor is 'poisson_gene_selection', this determines how many genes are processed at once. Choosing a smaller value will reduce the required memory. |
--n_samples | integer | The number of Binomial samples used to estimate the posterior probability of enrichment of zeros for each gene. Only used if flavor is 'poisson_gene_selection'. |
--check_values | boolean | Check if counts in the selected layer are integers. A warning is emitted otherwise. Only used if flavor is 'seurat_v3', 'seurat_v3_paper', 'pearson_residuals' or 'poisson_gene_selection'. |
Scanpy options
Name | Type & Properties | Description |
|---|---|---|
--var_input | string | If specified, use boolean array in adata.var[var_input] to calculate hvg on subset of vars. |
--features_to_exclude | string multiple | User-defined list of feature names to exclude before HVG calculation. These features will be excluded from HVG selection but will remain in the output data. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
device_type: [ "gpu" ]
output: "$id.$key.output.h5mu"
var_name_filter: [ "highly_variable" ]
varm_name: [ "hvg" ]
flavor: [ "seurat" ]
min_mean: [ 0.0125 ]
max_mean: [ 3 ]
min_disp: [ 0.5 ]
span: [ 0.3 ]
n_bins: [ 20 ]
theta: [ 100 ]
chunksize: [ 1000 ]
n_samples: [ 10000 ]
check_values: [ true ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/wrappers/preprocessing/highly_variable_genes/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
wrappers/preprocessing/highly_variable_genesopenpipeline_rapids v0.1.3