workflows/multiomics/process_batches
Description
This workflow serves as an entrypoint into the 'full_pipeline' in order to
re-run the multisample processing and the integration setup. An input .h5mu file will
first be split in order to run the multisample processing per modality. Next, the modalities
are merged again and the integration setup pipeline is executed. Please note that this workflow
assumes that samples from multiple pipelines are already concatenated.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input -i | file required multiple | Path to the sample. |
--rna_layer | string | Input layer for the gene expression modality. If not specified, .X is used. |
--prot_layer | string | Input layer for the antibody capture modality. If not specified, .X is used. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
Highly variable features detection
Name | Type & Properties | Description |
|---|---|---|
--highly_variable_features_var_output --filter_with_hvg_var_output | string | In which .var slot to store a boolean array corresponding to the highly variable genes. |
--highly_variable_features_obs_batch_key --filter_with_hvg_obs_batch_key | string | If specified, highly-variable genes are selected within each batch separately and merged. This simple process avoids the selection of batch-specific genes and acts as a lightweight batch correction method. |
QC metrics calculation options
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. |
--top_n_vars | integer multiple | Number of top vars to be used to calculate cumulative proportions. If not specified, proportions are not calculated. `--top_n_vars 20,50` finds cumulative proportion to the 20th and 50th most expressed vars. |
PCA options
Name | Type & Properties | Description |
|---|---|---|
--pca_overwrite | boolean_true | Allow overwriting slots for PCA output. |
CLR options
Name | Type & Properties | Description |
|---|---|---|
--clr_axis | integer | Axis to perform the CLR transformation on. |
RNA Scaling options
Name | Type & Properties | Description |
|---|---|---|
--rna_enable_scaling | boolean_true | Enable scaling for the RNA modality. |
--rna_scaling_output_layer | string | Output layer where the scaled log-normalized data will be stored. |
--rna_scaling_pca_obsm_output | string | Name of the .obsm key where the PCA representation of the log-normalized and scaled data is stored. |
--rna_scaling_pca_loadings_varm_output | string | Name of the .varm key where the PCA loadings of the log-normalized and scaled data is stored. |
--rna_scaling_pca_variance_uns_output | string | Name of the .uns key where the variance and variance ratio will be stored as a map. The map will contain two keys: variance and variance_ratio respectively. |
--rna_scaling_umap_obsm_output | string | Name of the .obsm key where the UMAP representation of the log-normalized and scaled data is stored. |
--rna_scaling_max_value | double | Clip (truncate) data to this value after scaling. If not specified, do not clip. |
--rna_scaling_zero_center | boolean_false | If set, omit zero-centering variables, which allows to handle sparse input efficiently." |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
output: "$id.$key.output.h5mu"
highly_variable_features_var_output: [ "filter_with_hvg" ]
highly_variable_features_obs_batch_key: [ "sample_id" ]
var_qc_metrics: [ "filter_with_hvg" ]
top_n_vars: [ 50, 100, 200, 500 ]
clr_axis: [ 0 ]
rna_scaling_output_layer: [ "scaled" ]
rna_scaling_pca_obsm_output: [ "scaled_pca" ]
rna_scaling_pca_loadings_varm_output: [ "scaled_pca_loadings" ]
rna_scaling_pca_variance_uns_output: [ "scaled_pca_variance" ]
rna_scaling_umap_obsm_output: [ "scaled_umap" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/multiomics/process_batches/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
workflows/multiomics/process_batchesopenpipeline v4.0.0