feature_annotation/align_query_reference
Description
Alignment of a query and reference dataset by:
Alignment of layers
Harmonization of .obs field names for batch and cell type labels
Harmonization of .var field name for gene names
Sanitation of gene names
Cross-checking of genes
Assignment of an id to the query and reference datasets
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. Note that the query and reference modalities should be the same. |
--input_layer | string | The layer in the input (query) data containing raw counts if .X is not to be used. |
--input_layer_lognormalized | string | The layer in the input (query) data containing log normalized counts if .X is not to be used. |
--input_var_gene_names | string | The name of the .var column in the input (query) data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--input_obs_batch | string required | The name of the .obs column in the input (query) data containing batch information. |
--input_obs_label | string | The name of the .obs column in the input (query) data containing cell type labels. If not provided, the --unkown_celltype_label will be assigned to all observations. |
--input_id | string | Meta id value to be assigned to the --output_obs_id .obs field of the aligned input (query) data. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train the CellTypist classifiers on. Only required if a pre-trained --model is not provided. |
--reference_layer | string | The layer in the reference data containing raw counts if .X is not to be used. Data are expected to be processed in the same way as the --input query dataset. |
--reference_layer_lognormalized | string | The layer in the reference data containing log normalized counts if .X is not to be used. Data are expected to be processed in the same way as the --input query dataset. |
--reference_var_gene_names | string | The name of the .var column in the reference data containing gene names; when no gene_name_layer is provided, the var index will be used. |
--reference_obs_batch | string required | The name of the .obs column in the reference data containing batch information. |
--reference_obs_label | string | The name of the .obs column in the reference data containing cell type labels. If not provided, the --unkown_celltype_label will be assigned to all observations. |
--reference_id | string | Meta id value to be assigned to the --output_obs_id .obs field of the aligned reference data. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output_query | file output | Aligned query data. |
--output_reference | file output | Aligned reference data. |
--output_layer | string | Name of the aligned layer containing raw counts in the output query and reference datasets. |
--output_layer_lognormalized | string | Name of the aligned layer containing log normalized counts in the output query and reference datasets. |
--output_var_gene_names | string | Name of the .var column in the output query and reference datasets containing the gene names. |
--output_obs_batch | string | Name of the .obs column in the output query and reference datasets containing the batch information. |
--output_obs_label | string | Name of the .obs column in the output query and reference datasets containing the cell type labels. |
--output_obs_id | string | Name of the .obs column in the output query and reference datasets containing the dataset id. |
--output_var_index | string | Name of the .var column to which the .var index of the --input and --reference datasets is stored. Only relevant if "--preserve_var_index" is False. |
--output_var_common_genes | string | Name of the .var column in the output query and reference datasets containing the boolean array indicating the common variables. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata H5 files. By default no compression is applied. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--input_reference_gene_overlap | integer | The minimum number of genes present in both the reference and query datasets. |
--align_layers_raw_counts | boolean | Whether to align the query and reference layers containing raw counts. |
--align_layers_lognormalized_counts | boolean_true | Whether to align the query and reference layers containing log normalized counts. |
--unkown_celltype_label | string | The label to assign to cells with an unknown cell type. |
--overwrite_existing_key | boolean_true | If set to true and the layer, obs or var key already exists in the query/reference file, the key will be overwritten. |
--preserve_var_index | boolean_true | If set to true, the .var index of the --input and --reference datasets will be preserved. If set to false (default behavior), the original .var index will be stored in the --output_var_index .var column and the .var index will be replaced with the sanitized & aligned gene names. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
input_id: [ "query" ]
reference_id: [ "reference" ]
output_query: "$id.$key.output_query.h5mu"
output_reference: "$id.$key.output_reference.h5mu"
output_layer: [ "_counts" ]
output_layer_lognormalized: [ "_log_normalized" ]
output_var_gene_names: [ "_gene_names" ]
output_obs_batch: [ "_sample_id" ]
output_obs_label: [ "_cell_type" ]
output_obs_id: [ "_dataset" ]
output_var_index: [ "_ori_var_index" ]
output_var_common_genes: [ "_common_vars" ]
input_reference_gene_overlap: [ 100 ]
align_layers_raw_counts: [ true ]
unkown_celltype_label: [ "Unknown" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/feature_annotation/align_query_reference/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
feature_annotation/align_query_referenceopenpipeline v4.2.0
Uses
0 relationships
No component dependencies found.