nichecompass/nichecompass
Description
Trains a NicheCompass model and generates the latent space for the provided input data.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input H5MU file |
--input_gp_mask | file required | JSON file containing a nested dictionary containing the gene programs, with keys being gene program names and values being dictionaries with keys `targets` and `sources`, where `targets` contains a list of the names of genes in the gene program for the reconstruction of the gene expression of the node itself (receiving node) and `sources` contains a list of the names of genes in the gene program for the reconstruction of the gene expression of the node's neighbors (transmitting nodes). |
--modality | string | Which modality from the input MuData file to process. |
--layer | string | Input layer containing raw counts to be use. If None, X is used |
--var_input | string | Key in adata.var that indicates which features to include as input for the NicheCompass model. If not provided, all features will be included. |
--input_obsp_spatial_connectivities | string | Key in adata.obsp where connectivities of the spatial neighborhood graph are stored |
--input_obs_covariates | string multiple | Keys of the adata.obs fields to use as covariates. |
Gene Program Mask
Name | Type & Properties | Description |
|---|---|---|
--min_genes_per_gp | integer | Minimum number of genes in a gene program including both target and source genes that need to be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--min_source_genes_per_gp | integer | Minimum number of source genes in a gene program that need to be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--min_target_genes_per_gp | integer | Minimum number of target genes in a gene program that need to be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--max_genes_per_gp | integer | Maximum number of genes in a gene program including both target and source genes that can be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--max_source_genes_per_gp | integer | Maximum number of source genes in a gene program that can be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--max_target_genes_per_gp | integer | Maximum number of target genes in a gene program that can be available in the adata (gene expression has been probed) for a gene program not to be discarded. |
--filter_genes_not_in_masks | boolean_true | Whether to remove the genes that are not in the gp masks from the adata object. |
NicheCompass Model Architecture
Name | Type & Properties | Description |
|---|---|---|
--covariate_edges | boolean multiple | List of booleans that indicate whether there can be edges between different categories of the categorical covariates. If this is `True` for a specific categorical covariate, this covariate will be excluded from the edge reconstruction loss. Needs to match the length and order of `--input_obs_covariates`. |
--covariate_embedding_injection_layers | string multiple | List of VGPGAE modules in which the categorical covariates embeddings are injected. |
--include_edge_recon_loss | boolean | Whether to include the edge reconstruction loss in the backpropagation. |
--include_gene_expr_recon_loss | boolean | Whether to include the gene expression reconstruction loss in the backpropagation. |
--include_cat_covariates_contrastive_loss | boolean | Whether to include the categorical covariates contrastive loss in the backpropagation. |
--gene_expr_recon_dist | string | The distribution used for gene expression reconstruction. If `nb`, uses a negative binomial distribution. If `zinb`, uses a zero-inflated negative binomial distribution. |
--log_variational | boolean | Whether to transform x by log(x+1) prior to encoding for numerical stability (not for normalization). |
--node_label_method | string | Node label method that will be used for omics reconstruction. If `one-hop-sum`, uses a concatenation of the node's input features with the sum of the input features of all nodes in the node's one-hop neighborhood. If `one-hop-norm`, uses a concatenation of the node's input features with the node's one-hop neighbors input features normalized as per Kipf, T. N. & Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. arXiv [cs.LG] (2016). If `one-hop-attention`, uses a concatenation of the node's input features with the node's one-hop neighbors input features weighted by an attention mechanism. |
--active_gp_thresh_ratio | double | Ratio that determines which gene programs are considered active and are used in the latent representation after model training. All inactive gene programs will be dropped during model training after a determined number of epochs. Aggregations of the absolute values of the gene weights of the gene expression decoder per gene program are calculated. The maximum value, i.e. the value of the gene program with the highest aggregated value will be used as a benchmark and all gene programs whose aggregated value is smaller than `--active_gp_thresh_ratio` times this maximum value will be set to inactive. If set to 0, all gene programs will be considered active. |
--active_gp_type | string | Type to determine active gene programs. Can be `mixed`, in which case active gene programs are determined across prior and add-on gene programs jointly, or `separate` in which case they are determined separately for prior and add-on gene programs. |
--n_fc_layers_encoder | integer | Number of fully connected layers in the encoder before message passing layers. |
--n_layers_encoder | integer | Number of message passing layers in the encoder. |
--n_hidden_encoder | integer | Number of nodes in the encoder hidden layers. If not provided, it will be determined automatically based on the number of input genes and gene programs. |
--conv_layer_encoder | string | Convolutional layer used as GNN in the encoder. |
--encoder_n_attention_heads | integer | Number of attention heads used in the GNN layers of the encoder. Only relevant if `--conv_layer_encoder gatv2conv`. |
--encoder_use_bn | boolean | Whether to use a batch normalization layer at the end of the encoder to normalize `mu`. |
--dropout_rate_encoder | double | Probability that nodes will be dropped in the encoder during training. |
--dropout_rate_graph_decoder | double | Probability that nodes will be dropped in the graph decoder during training. |
--n_addon_gp | integer | Number of addon gene programs (i.e. gene programs that are not included in masks but can be learned de novo). |
--cat_covariates_embeds_nums | integer multiple | Number of embedding nodes for all categorical covariates. Must be the same length as `--input_obs_covariates`. |
--random_state | integer | Random seed for reproducibility. |
NicheCompass Training Parameters
Name | Type & Properties | Description |
|---|---|---|
--n_epochs | integer | Number of training epochs |
--n_epochs_all_gps | integer | Number of epochs during which all gene programs are used for model training. After that only active gene programs are retained. |
--n_epochs_no_edge_recon | integer | Number of epochs during which the edge reconstruction loss is excluded from backpropagation for pretraining using the other loss components. |
--n_epochs_no_cat_covariates_contrastive | integer | Number of epochs during which the categorical covariates contrastive loss is excluded from backpropagation for pretraining using the other loss components. |
--lr | double | Learning rate |
--weight_decay | double | Weight decay (L2 penalty). |
--lambda_edge_recon | double | Lambda (weighting factor) for the edge reconstruction loss. If `>0`, this will enforce gene programs to be meaningful for edge reconstruction and, hence, to preserve spatial colocalization information. |
--lambda_gene_expr_recon | double | Lambda (weighting factor) for the gene expression reconstruction loss. If `>0`, this will enforce interpretable gene programs that can be combined in a linear way to reconstruct gene expression. |
--lambda_cat_covariates_contrastive | double | Lambda (weighting factor) for the categorical covariates contrastive loss. If `>0`, this will enforce observations with different categorical covariates categories with very similar latent representations to become more similar, and observations with different latent representations to become more different. |
--contrastive_logits_pos_ratio | double | Ratio for determining the logits threshold of positive contrastive examples of node pairs from different categorical covariates categories. The top (`contrastive_logits_pos_ratio` * 100)% logits of node pairs from different categorical covariates categories serve as positive labels for the contrastive loss. |
--contrastive_logits_neg_ratio | double | Ratio for determining the logits threshold of negative contrastive examples of node pairs from different categorical covariates categories. The bottom (`contrastive_logits_neg_ratio` * 100)% logits of node pairs from different categorical covariates categories serve as negative labels for the contrastive loss. |
--lambda_group_lasso | double | Lambda (weighting factor) for the group lasso regularization loss of gene programs. If `>0`, this will enforce sparsity of gene programs. |
--lambda_l1_masked | double | Lambda (weighting factor) for the L1 regularization loss of genes in masked gene programs. If `>0`, this will enforce sparsity of genes in masked gene programs. |
--l1_targets_categories | string multiple | Gene program mask targets categories for which l1 regularization loss will be applied. |
--l1_sources_categories | string multiple | Gene program mask sources categories for which l1 regularization loss will be applied. |
--lambda_l1_addon | double | Lambda (weighting factor) for the L1 regularization loss of genes in addon gene programs. If `>0`, this will enforce sparsity of genes in addon gene programs. |
--edge_val_ratio | double | Fraction of the data that is used as validation set on edge-level. The rest of the data will be used as training set on edge-level. |
--node_val_ratio | double | Fraction of the data that is used as validation set on node-level. The rest of the data will be used as training set on node-level. |
--edge_batch_size | integer | Batch size for the edge-level dataloaders. |
--node_batch_size | integer | Batch size for the node-level dataloaders. If not provided, is automatically determined based on `--edge_batch_size`. |
--n_sampled_neighbors | integer | Number of neighbors that are sampled during model training from the spatial neighborhood graph. If set to -1, all direct neighbors are included. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output H5MU file |
--output_model | file required output | Directory to save the trained NicheCompass model |
--output_obsm_embedding | string | Key of the obsm field where the latent / gene program representation of active gene programs will be stored after NicheCompass model training. |
--output_varm_gp_targets_mask | string | Key of the varm field where the binary gene program mask for target genes of a gene program will be stored (target genes are used for the reconstruction of the gene expression of the node itself (receiving node )). |
--output_varm_gp_sources_mask | string | Key of the varm field where the binary gene program mask for source genes of a gene program will be stored (source genes are used for the reconstruction of the gene expression of the node's neighbors (transmitting nodes)). |
--output_uns_gp_names | string | Key of the uns field where the gene program names will be stored. |
--output_uns_active_gp_names | string | Key of the uns field where the active gene program names will be stored. |
--output_uns_genes_index | string | Key of the uns field where the index of a concatenated vector of target and source genes that are in the gene program masks will be stored. |
--output_uns_target_genes_index | string | Key of the uns field where the index of the target genes that are in the gene program mask will be stored. |
--output_uns_source_genes_index | string | Key of the uns field where the index of the source genes that are in the gene program mask will be stored. |
--output_uns_covariate_embeddings | string multiple | Key of the uns fields where the covariate embeddings will be stored. Needs to match the length and order of `--input_obs_covariates`. |
--output_obsp_reconstructed_adj_edge_proba | string | Key of the obsp field where the reconstructed adjacency matrix edge probabilities will be stored. |
--output_obsp_agg_weights | string | Key of the obsp field where the aggregation weights of the node label aggregator will be stored. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
var_input: [ "filter_with_hvg" ]
input_obsp_spatial_connectivities: [ "spatial_connectivities" ]
min_genes_per_gp: [ 1 ]
min_source_genes_per_gp: [ 0 ]
min_target_genes_per_gp: [ 0 ]
covariate_embedding_injection_layers: [ "gene_expr_decoder", "chrom_access_decoder" ]
include_edge_recon_loss: [ true ]
include_gene_expr_recon_loss: [ true ]
include_cat_covariates_contrastive_loss: [ true ]
gene_expr_recon_dist: [ "nb" ]
log_variational: [ true ]
node_label_method: [ "one-hop-norm" ]
active_gp_thresh_ratio: [ 0.1 ]
active_gp_type: [ "separate" ]
n_fc_layers_encoder: [ 1 ]
n_layers_encoder: [ 1 ]
conv_layer_encoder: [ "gatv2conv" ]
encoder_n_attention_heads: [ 4 ]
encoder_use_bn: [ false ]
dropout_rate_encoder: [ 0 ]
dropout_rate_graph_decoder: [ 0 ]
n_addon_gp: [ 100 ]
random_state: [ 0 ]
n_epochs: [ 100 ]
n_epochs_all_gps: [ 25 ]
n_epochs_no_edge_recon: [ 0 ]
n_epochs_no_cat_covariates_contrastive: [ 5 ]
lr: [ 0.001 ]
weight_decay: [ 0.001 ]
lambda_edge_recon: [ 500000 ]
lambda_gene_expr_recon: [ 300 ]
lambda_cat_covariates_contrastive: [ 0 ]
contrastive_logits_pos_ratio: [ 0 ]
contrastive_logits_neg_ratio: [ 0 ]
lambda_group_lasso: [ 0 ]
lambda_l1_masked: [ 0 ]
lambda_l1_addon: [ 30 ]
edge_val_ratio: [ 0.1 ]
node_val_ratio: [ 0.1 ]
edge_batch_size: [ 256 ]
n_sampled_neighbors: [ -1 ]
output: "$id.$key.output"
output_model: "$id.$key.output_model"
output_obsm_embedding: [ "nichecompass_latent" ]
output_varm_gp_targets_mask: [ "nichecompass_gp_targets" ]
output_varm_gp_sources_mask: [ "nichecompass_gp_sources" ]
output_uns_gp_names: [ "nichecompass_gp_names" ]
output_uns_active_gp_names: [ "nichecompass_active_gp_names" ]
output_uns_genes_index: [ "nichecompass_genes_idx" ]
output_uns_target_genes_index: [ "nichecompass_target_genes_idx" ]
output_uns_source_genes_index: [ "nichecompass_source_genes_idx" ]
output_obsp_reconstructed_adj_edge_proba: [ "nichecompass_recon_connectivities" ]
output_obsp_agg_weights: [ "nichecompass_agg_weights" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_spatial.git \
-revision v0.7.0 \
-main-script target/nextflow/nichecompass/nichecompass/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
nichecompass/nichecompassopenpipeline_spatial v0.7.0
Uses
0 relationships
No component dependencies found.