qc/calculate_qc_metrics
Description
Add basic quality control metrics to an .h5mu file.
The metrics are comparable to what scanpy.pp.calculate_qc_metrics output,
although they have slightly different names:
Var metrics (name in this component -> name in scanpy):
pct_dropout -> pct_dropout_by_{expr_type}
num_nonzero_obs -> n_cells_by_{expr_type}
obs_mean -> mean_{expr_type}
total_counts -> total_{expr_type}
Obs metrics:
- num_nonzero_vars -> n_genes_by_{expr_type}
- pct_{var_qc_metrics} -> pct_{expr_type}{qc_var}
- total_counts{var_qc_metrics} -> total_{expr_type}{qc_var}
- pct_of_counts_in_top{top_n_vars}vars -> pct{expr_type}in_top{n}{var_type}
- total_counts -> total{expr_type}
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string | Which modality from the input MuData file to process. |
--layer | string | Layer from modality to use as input data. If not provided the .X attribute is used. |
Metrics added to .obs
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. |
--var_qc_metrics_fill_na_value | boolean | Fill any 'NA' values found in the columns specified with --var_qc_metrics to 'True' or 'False'. as False. |
--top_n_vars | integer multiple | Number of top vars to be used to calculate cumulative proportions. If not specified, proportions are not calculated. `--top_n_vars 20;50` finds cumulative proportion to the 20th and 50th most expressed vars. |
--output_obs_num_nonzero_vars | string | Name of column in .obs describing, for each observation, the number of stored values (including explicit zeroes). In other words, the name of the column that counts for each row the number of columns that contain data. |
--output_obs_total_counts_vars | string | Name of the column for .obs describing, for each observation (row), the sum of the stored values in the columns. |
Metrics added to .var
Name | Type & Properties | Description |
|---|---|---|
--output_var_num_nonzero_obs | string | Name of column describing, for each feature, the number of stored values (including explicit zeroes). In other words, the name of the column that counts for each column the number of rows that contain data. |
--output_var_total_counts_obs | string | Name of the column in .var describing, for each feature (column), the sum of the stored values in the rows. |
--output_var_obs_mean | string | Name of the column in .obs providing the mean of the values in each row. |
--output_var_pct_dropout | string | Name of the column in .obs providing for each feature the percentage of observations the feature does not appear on (i.e. is missing). Same as `--num_nonzero_obs` but percentage based. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file output | Output h5mu file. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output_obs_num_nonzero_vars: [ "num_nonzero_vars" ]
output_obs_total_counts_vars: [ "total_counts" ]
output_var_num_nonzero_obs: [ "num_nonzero_obs" ]
output_var_total_counts_obs: [ "total_counts" ]
output_var_obs_mean: [ "obs_mean" ]
output_var_pct_dropout: [ "pct_dropout" ]
output: "$id.$key.output.h5mu"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.3 \
-main-script target/nextflow/qc/calculate_qc_metrics/main.nf \
-params-file params.yaml Relationships
Used by
Current component
Uses
No component dependencies found.