workflows/differential_expression/pseudobulk_deseq2
Description
Performs pseudobulk generation and subsequent differential expression analysis using DESeq2.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--input_layer | string | Input layer to use. If None, X is used. This layer must contain raw counts. |
--obs_cell_group | string required | .obs field containing the variable to group on. Typically this field contains cell groups such as annotated cell types or clusters. If provided, will generate DESeq2 analysis per cell group. |
--obs_groups | string multiple | .obs fields containing the experimental condition(s) to create pseudobulk samples for. |
--var_gene_names | string | Name of the .var field that contains gene symbols. If not provided, .var.index will be used. |
Pseudobulk Options
Name | Type & Properties | Description |
|---|---|---|
--aggregation_method | string | Method to aggregate the raw counts for pseudoreplicates. Either sum or mean. |
--random_state | integer | The random seed for sampling. |
Filtering options
Name | Type & Properties | Description |
|---|---|---|
--min_obs_per_sample | integer | Minimum number of observations per pseudobulk sample. |
--filter_genes_min_samples | integer | Minimum number of samples a gene must be expressed in to be included in the analysis. If None, no filtering is applied. |
--filter_gene_patterns | string multiple | List of regex patterns to filter out genes. Genes matching any of these patterns will be excluded from the analysis. |
DGEA options
Name | Type & Properties | Description |
|---|---|---|
--design_formula | string required | Design formula for DESeq2 analysis in R-style format. Specifies the statistical model to account for various factors. |
--contrast_column | string required | Column in the metadata to use for the contrast. This column should contain the conditions to compare. |
--contrast_values | string required multiple | Values to compare in the contrast column. First value is baseline, following values are different treatments. |
--p_adj_threshold | double | Adjusted p-value threshold for significance. Genes with adjusted p-values below this threshold will be considered significant. |
--log2fc_threshold | double | Log2 fold change threshold for significance. Genes with absolute log2 fold change above this threshold will be considered significant. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output directory for DESeq2 results, containing one CSV file per cell group (as defined by `--obs_cell_group`) |
--output_prefix | string | Prefix for output CSV files: files will be named "{prefix}_{cell_group}.csv" |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
aggregation_method: [ "sum" ]
random_state: [ 0 ]
min_obs_per_sample: [ 30 ]
p_adj_threshold: [ 0.05 ]
log2fc_threshold: [ 0 ]
output: "$id.$key.output"
output_prefix: [ "deseq2_analysis" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/workflows/differential_expression/pseudobulk_deseq2/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/differential_expression/pseudobulk_deseq2openpipeline v4.2.0