integrate/totalvi_scarches
Description
Performs totalVI integration by mapping the query dataset to a reference dataset or model.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file with query data to integrate with reference. |
--reference -r | file required | Input h5mu file with reference data to train the TOTALVI model. |
--force_retrain -f | boolean_true | If true, retrain the model and save it to reference_model_path |
--query_modality | string | |
--query_proteins_modality | string | Name of the modality in the input (query) h5mu file containing protein data |
--reference_modality | string | |
--reference_proteins_modality | string | Name of the modality containing proteins in the reference |
--input_layer | string | Input layer to use. If None, X is used |
--obs_batch | string | Column name discriminating between your batches. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--obsm_normalized_rna_output | string | In which .obsm slot to store the normalized RNA from TOTALVI. |
--obsm_normalized_protein_output | string | In which .obsm slot to store the normalized protein data from TOTALVI. |
--reference_model_path | file output | Directory with the reference model. If not exists, trained model will be saved there |
--query_model_path | file output | Directory, where the query model will be saved |
Learning parameters
Name | Type & Properties | Description |
|---|---|---|
--max_epochs | integer | Number of passes through the dataset |
--max_query_epochs | integer | Number of passes through the dataset, when fine-tuning model for query |
--weight_decay | double | Weight decay, when fine-tuning model for query |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
query_modality: [ "rna" ]
reference_modality: [ "rna" ]
reference_proteins_modality: [ "prot" ]
obs_batch: [ "sample_id" ]
output: "$id.$key.output"
obsm_output: [ "X_integrated_totalvi" ]
obsm_normalized_rna_output: [ "X_totalvi_normalized_rna" ]
obsm_normalized_protein_output: [ "X_totalvi_normalized_protein" ]
reference_model_path: "$id.$key.reference_model_path"
query_model_path: "$id.$key.query_model_path"
max_epochs: [ 400 ]
max_query_epochs: [ 200 ]
weight_decay: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/integrate/totalvi_scarches/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
integrate/totalvi_scarchesopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.