integrate/scarches
Description
Performs reference mapping with scArches
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file to use as a query |
--layer | string | Layer to be used for scArches, if .X is not to be used. |
--modality | string | Which modality from the input MuData file to process. |
--input_obs_batch | string | Name of the .obs column with batch information. |
--input_obs_label | string | Name of the .obs column with celltype information. |
--input_var_gene_names | string | Name of the .var column with gene names, if the var .index is not to be used. |
--input_obs_size_factor | string | Key in adata.obs for size factor information. Instead of using library size as a size factor, the provided size factor column will be used as offset in the mean of the likelihood. Assumed to be on linear scale. |
--input_obs_categorical_covariate | string multiple | Keys in adata.obs that correspond to categorical data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the categorical covariates matters and should match the order of the covariates in the trained reference model. |
--input_obs_continuous_covariate | string multiple | Keys in adata.obs that correspond to continuous data. These covariates can be added in addition to the batch covariate and are also treated as nuisance factors (i.e., the model tries to minimize their effects on the latent space). Thus, these should not be used for biologically-relevant factors that you do _not_ want to correct for. Important: the order of the continuous covariates matters and should match the order of the covariates in the trained reference model. |
--sanitize_ensembl_ids | boolean | Whether to sanitize ensembl ids by removing version numbers. |
Reference
Name | Type & Properties | Description |
|---|---|---|
--reference -r | file required | Path to the directory with reference model or a web link. |
--reference_class | string | For legacy models; the type of model (where the type of model was not saved with it; e.g. when they were generated with scvi-tools versions < 0.15). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--model_output | file output | Output directory for model |
--obsm_output | string | In which .obsm slot to store the resulting integrated embedding. |
--obs_output_predictions | string | In which .obs slot to store the resulting label predictions. Only relevant if a scANVI model was provided. |
--obs_output_probabilities | string | In which .obs slot to store the probabilities of the label predictions. Only relevant if a scANVI model was provided. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Early stopping arguments
Name | Type & Properties | Description |
|---|---|---|
--early_stopping | boolean | Whether to perform early stopping with respect to the validation set. |
--early_stopping_monitor | string | Metric logged during validation set epoch. |
--early_stopping_patience | integer | Number of validation epochs with no improvement after which training will be stopped. |
--early_stopping_min_delta | double | Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. |
Learning parameters
Name | Type & Properties | Description |
|---|---|---|
--max_epochs | integer | Number of passes through the dataset, defaults to (20000 / number of cells) * 400 or 400; whichever is smallest. |
--reduce_lr_on_plateau | boolean | Whether to monitor validation loss and reduce learning rate when validation set `lr_scheduler_metric` plateaus. |
--lr_factor | double | Factor to reduce learning rate. |
--lr_patience | double | Number of epochs with no improvement after which learning rate will be reduced. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
sanitize_ensembl_ids: [ true ]
output: "$id.$key.output"
model_output: "$id.$key.model_output"
obsm_output: [ "X_integrated_scanvi" ]
obs_output_predictions: [ "scanvi_pred" ]
obs_output_probabilities: [ "scanvi_proba" ]
early_stopping_monitor: [ "elbo_validation" ]
early_stopping_patience: [ 45 ]
early_stopping_min_delta: [ 0 ]
reduce_lr_on_plateau: [ true ]
lr_factor: [ 0.6 ]
lr_patience: [ 30 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/integrate/scarches/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
integrate/scarchesopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.