workflows/rna/rna_multisample
Description
Processing unimodal multi-sample RNA transcriptomics data.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the concatenated file |
--input | file required | Path to the samples. |
--modality | string | Modality to process. |
--layer | string | Input layer to use. If not specified, .X is used. |
Output
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
Filtering highly variable features
Name | Type & Properties | Description |
|---|---|---|
--highly_variable_features_var_output --filter_with_hvg_var_output | string | In which .var slot to store a boolean array corresponding to the highly variable features. |
--highly_variable_features_obs_batch_key --filter_with_hvg_obs_batch_key | string | If specified, highly-variable features are selected within each batch separately and merged. This simple process avoids the selection of batch-specific features and acts as a lightweight batch correction method. For all flavors, featues are first sorted by how many batches they are highly variable. For dispersion-based flavors ties are broken by normalized dispersion. If flavor = 'seurat_v3', ties are broken by the median (across batches) rank based on within-batch normalized variance. |
--highly_variable_features_flavor --filter_with_hvg_flavor | string | Choose the flavor for identifying highly variable features. For the dispersion based methods in their default workflows, Seurat passes the cutoffs whereas Cell Ranger passes n_top_features. |
--highly_variable_features_n_top_features --filter_with_hvg_n_top_genes | integer | Number of highly-variable features to keep. Mandatory if filter_with_hvg_flavor is set to 'seurat_v3'. |
QC metrics calculation options
Name | Type & Properties | Description |
|---|---|---|
--var_qc_metrics | string multiple | Keys to select a boolean (containing only True or False) column from .var. For each cell, calculate the proportion of total values for genes which are labeled 'True', compared to the total sum of the values for all genes. |
--top_n_vars | integer multiple | Number of top vars to be used to calculate cumulative proportions. If not specified, proportions are not calculated. `--top_n_vars 20,50` finds cumulative proportion to the 20th and 50th most expressed vars. |
--output_obs_num_nonzero_vars | string | Name of column in .obs describing, for each observation, the number of stored values (including explicit zeroes). In other words, the name of the column that counts for each row the number of columns that contain data. |
--output_obs_total_counts_vars | string | Name of the column for .obs describing, for each observation (row), the sum of the stored values in the columns. |
--output_var_num_nonzero_obs | string | Name of column describing, for each feature, the number of stored values (including explicit zeroes). In other words, the name of the column that counts for each column the number of rows that contain data. |
--output_var_total_counts_obs | string | Name of the column in .var describing, for each feature (column), the sum of the stored values in the rows. |
--output_var_obs_mean | string | Name of the column in .obs providing the mean of the values in each row. |
--output_var_pct_dropout | string | Name of the column in .obs providing for each feature the percentage of observations the feature does not appear on (i.e. is missing). Same as `--num_nonzero_obs` but percentage based. |
RNA Scaling options
Name | Type & Properties | Description |
|---|---|---|
--enable_scaling | boolean_true | Enable scaling for the RNA modality. |
--scaling_output_layer | string | Output layer where the scaled log-normalized data will be stored. |
--scaling_max_value | double | Clip (truncate) data to this value after scaling. If not specified, do not clip. |
--scaling_zero_center | boolean_false | If set, omit zero-centering variables, which allows to handle sparse input efficiently." |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
highly_variable_features_var_output: [ "filter_with_hvg" ]
highly_variable_features_obs_batch_key: [ "sample_id" ]
highly_variable_features_flavor: [ "seurat" ]
var_qc_metrics: [ "filter_with_hvg" ]
top_n_vars: [ 50, 100, 200, 500 ]
output_obs_num_nonzero_vars: [ "num_nonzero_vars" ]
output_obs_total_counts_vars: [ "total_counts" ]
output_var_num_nonzero_obs: [ "num_nonzero_obs" ]
output_var_total_counts_obs: [ "total_counts" ]
output_var_obs_mean: [ "obs_mean" ]
output_var_pct_dropout: [ "pct_dropout" ]
scaling_output_layer: [ "scaled" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/workflows/rna/rna_multisample/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
workflows/rna/rna_multisampleopenpipeline v4.0.0