integrate/totalvi
Description
Performs TotalVI integration of CITE-seq and scRNA-seq data.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file to be integrated. |
--rna_modality | string required | |
--prot_modality | string required | Name of the modality in the input (query) h5mu file containing protein data |
--input_layer_rna | string | Input layer to use from the rna modality for gene counts. If None, X is used |
--input_layer_protein | string | Input layer to use from the protein modality for protein counts. If None, X is used |
--obs_batch | string | Column name discriminating between your batches. |
--obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. |
--var_gene_names | string | .var column in the rna modality containing gene names. By default, use mod['rna'].var_names. |
--var_protein_names | string | .var column in the protein modality containing protein names. By default, use mod['prot'].var_names. |
--var_input | string | .var column containing highly variable genes. By default, do not subset genes. |
Data validity checks
Name | Type & Properties | Description |
|---|---|---|
--n_obs_min_count | integer | Minimum number of cells threshold ensuring that every obs_batch category has sufficient observations (cells) for model training. |
--n_var_gene_min_count | integer | Minimum number of genes threshold ensuring that sufficient observations (genes) are present for model training. |
--n_var_protein_min_count | integer | Minimum number of proteins threshold ensuring that sufficient observations (genes) are present for model training. |
Model arguments
Name | Type & Properties | Description |
|---|---|---|
--n_dimensions_latent_space | integer | Dimensionality of the latent space. |
--gene_dispersion | string | Set the behavior for the dispersion for negative binomial distributions: - gene: dispersion parameter of negative binomial is constant per gene across cells - gene-batch: dispersion can differ between different batches - gene-label: dispersion can differ between different labels - gene-cell: dispersion can differ for every gene in every cell |
--protein_dispersion | string | Set the behavior for the dispersion for negative binomial distributions: - protein: dispersion parameter of negative binomial is constant per protein across cells - protein-batch: dispersion can differ between different batches - protein-label: dispersion can differ between different labels |
--gene_likelihood | string | Model used to generate the expression data from a count-based likelihood distribution. - nb: Negative binomial distribution - zinb: Zero-inflated negative binomial distribution |
--latent_distribution | string | Set the latent distribution to use: - normal: Normal distribution - ln: Log-normal distribution (Normal( 0, 1) transformed by softmax) |
--empirical_protein_background_prior | boolean | Set the initialization of protein background prior empirically. This option fits a GMM for each of 100 cells per batch and averages the distributions. Note that even with this option set to True, this only initializes a parameter that is learned during inference. If False, randomly initializes. By default, sets this to True if greater than 10 proteins are used. |
--override_missing_proteins | boolean_true | If True, will not treat proteins with all 0 expression in a particular batch as missing. |
Training arguments
Name | Type & Properties | Description |
|---|---|---|
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--output_model | file output | Directory with the reference model. If not exists, trained model will be saved there |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--obsm_normalized_rna_output | string | In which .obsm slot to store the normalized RNA data from TOTALVI. |
--obsm_normalized_protein_output | string | In which .obsm slot to store the normalized protein data from TOTALVI. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
rna_modality: [ "rna" ]
prot_modality: [ "prot" ]
obs_batch: [ "sample_id" ]
n_obs_min_count: [ 0 ]
n_var_gene_min_count: [ 0 ]
n_var_protein_min_count: [ 0 ]
n_dimensions_latent_space: [ 20 ]
gene_dispersion: [ "gene" ]
protein_dispersion: [ "protein" ]
gene_likelihood: [ "nb" ]
latent_distribution: [ "normal" ]
early_stopping: [ true ]
output: "$id.$key.output"
output_model: "$id.$key.output_model"
obsm_output: [ "X_integrated_totalvi" ]
obsm_normalized_rna_output: [ "X_totalvi_normalized_rna" ]
obsm_normalized_protein_output: [ "X_totalvi_normalized_protein" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/integrate/totalvi/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
integrate/totalviopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.