preprocessing/normalize_total
Description
Normalize counts per cell.
Normalize each cell by total counts over all genes, so that every cell has
the same total count after normalization. If choosing target_sum=1e6, this
is CPM normalization.
If exclude_highly_expressed=True, very highly expressed genes are excluded
from the computation of the normalization factor (size factor) for each
cell. This is meaningful as these can strongly influence the resulting
normalized values for all other genes [Weinreb17].
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | Input layer to use. By default, X is normalized. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--output_layer | string | Output layer to use. By default, use X. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--target_sum | integer | If None, after normalization, each observation (cell) has a total count equal to the median of total counts for observations (cells) before normalization. |
--exclude_highly_expressed | boolean_true | Exclude (very) highly expressed genes for the computation of the normalization factor (size factor) for each cell. A gene is considered highly expressed if it has more than max_fraction of the total counts in at least one cell. The not-excluded genes will sum up to target_sum. |
--max_fraction | double | If exclude_highly_expressed=True, consider cells as highly expressed that have more counts than max_fraction of the original total counts in at least one cell. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output: "$id.$key.output"
max_fraction: [ 0.05 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/preprocessing/normalize_total/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
preprocessing/normalize_totalopenpipeline_rapids v0.1.3
Uses
0 relationships
No component dependencies found.