workflows/niche/nichecompass_leiden
Description
A pipeline to compute the spatial neighborhood graph, perform nichecompass embedding followed by Leiden clustering.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input -i | file required | Path to the sample. |
--input_gp_mask | file required | JSON file containing a nested dictionary containing the gene programs, with keys being gene program names and values being dictionaries with keys `targets` and `sources`, where `targets` contains a list of the names of genes in the gene program for the reconstruction of the gene expression of the node itself (receiving node) and `sources` contains a list of the names of genes in the gene program for the reconstruction of the gene expression of the node's neighbors (transmitting nodes). |
--modality | string | Which modality to process. |
--layer | string | Use specified layer for calculation of qc metrics. If not specified, adata.X is used. |
--input_obs_batch_key | string | Key of the adata.obs field that contains the batch information. |
--input_obs_covariates | string multiple | Keys of the adata.obs fields to use as covariates. |
--input_obsm_spatial_coords | string | Key in adata.obsm where spatial coordinates are stored |
--var_input | string | Key in adata.var that indicates which features to include as input for the NicheCompass model. If not provided, all features will be included. |
Spatial Neighbors Calculation
Name | Type & Properties | Description |
|---|---|---|
--coord_type | string | Type of coordinate system. Valid options are: `grid` - grid coordinates. Only relevant for lattice-based layouts (e.g. Visium). Builds a graph assuming a regular grid topology, where neighbors are defined by grid adjacency rather than distance. `generic` - Generic coordinates. Recommended for imaging-based data. This option is appropriate for irregularly spaced points, such as cell centroids or spot coordinates from imaging-based assays (e.g. Xenium, CosMx). Setting this option avoids grid-assumption artifacts and ensures each node has spatial neighbors as defined by distance. |
--n_spatial_neighbors | integer | Depending on `--coord_type`: `grid` - number of neighboring tiles. `generic` - number of neighborhoods for non-grid data. Only used when `--delaunay False`. |
--delaunay | boolean | Whether to use Delaunay triangulation to determine spatial neighborhood graph. Only used when `--coord_type generic`. |
Gene Program Mask
Name | Type & Properties | Description |
|---|---|---|
--min_genes_per_gp | integer | Minimum number of genes in a gene program including both target and source genes that need to be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--min_source_genes_per_gp | integer | Minimum number of source genes in a gene program that need to be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--min_target_genes_per_gp | integer | Minimum number of target genes in a gene program that need to be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--max_genes_per_gp | integer | Maximum number of genes in a gene program inluding both target and source genes that can be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--max_source_genes_per_gp | integer | Maximum number of source genes in a gene program that can be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--max_target_genes_per_gp | integer | Maximum number of target genes in a gene program that can be available in the input data (gene expression has been probed) for a gene program not to be discarded. |
--filter_genes_not_in_masks | boolean_true | Whether to remove the genes that are not in the gp masks from the input data. |
NicheCompass Model Architecture
Name | Type & Properties | Description |
|---|---|---|
--covariate_edges | boolean multiple | List of booleans that indicate whether there can be edges between different categories of the categorical covariates. If this is `True` for a specific categorical covariate, this covariate will be excluded from the edge reconstruction loss. Needs to match the length and order of `--input_obs_covariates`. |
--gene_expr_recon_dist | string | The distribution used for gene expression reconstruction. If `nb`, uses a negative binomial distribution. If `zinb`, uses a zero-inflated negative binomial distribution. |
--log_variational | boolean | Whether to transform x by log(x+1) prior to encoding for numerical stability (not for normalization). |
--node_label_method | string | Node label method that will be used for omics reconstruction. If `one-hop-sum`, uses a concatenation of the node's input features with the sum of the input features of all nodes in the node's one-hop neighborhood. If `one-hop-norm`, uses a concatenation of the node's input features with the node's one-hop neighbors input features normalized as per Kipf, T. N. & Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. arXiv [cs.LG] (2016). If `one-hop-attention`, uses a concatenation of the node's input features with the node's one-hop neighbors input features weighted by an attention mechanism. |
--active_gp_thresh_ratio | double | Ratio that determines which gene programs are considered active and are used in the latent representation after model training. All inactive gene programs will be dropped during model training after a determined number of epochs. Aggregations of the absolute values of the gene weights of the gene expression decoder per gene program are calculated. The maximum value, i.e. the value of the gene program with the highest aggregated value will be used as a benchmark and all gene programs whose aggregated value is smaller than `--active_gp_thresh_ratio` times this maximum value will be set to inactive. If set to 0, all gene programs will be considered active. |
--active_gp_type | string | Type to determine active gene programs. Can be `mixed`, in which case active gene programs are determined across prior and add-on gene programs jointly, or `separate` in which case they are determined separately for prior and add-on gene programs. |
--n_addon_gp | integer | Number of addon gene programs (i.e. gene programs that are not included in masks but can be learned de novo). |
--cat_covariates_embeds_nums | integer multiple | Number of embedding nodes for all categorical covariates. Must be the same length as `--input_obs_covariates`. |
--random_state | integer | Random seed for reproducibility. |
NicheCompass Training Parameters
Name | Type & Properties | Description |
|---|---|---|
--n_epochs | integer | Number of training epochs |
--n_epochs_all_gps | integer | Number of epochs during which all gene programs are used for model training. After that only active gene programs are retained. |
--n_epochs_no_edge_recon | integer | Number of epochs during which the edge reconstruction loss is excluded from backpropagation for pretraining using the other loss components. |
--n_epochs_no_cat_covariates_contrastive | integer | Number of epochs during which the categorical covariates contrastive loss is excluded from backpropagation for pretraining using the other loss components. |
--lr | double | Learning rate |
--weight_decay | double | Weight decay (L2 penalty). |
--edge_val_ratio | double | Fraction of the data that is used as validation set on edge-level. The rest of the data will be used as training set on edge-level. |
--node_val_ratio | double | Fraction of the data that is used as validation set on node-level. The rest of the data will be used as training set on node-level. |
--edge_batch_size | integer | Batch size for the edge-level dataloaders. |
--node_batch_size | integer | Batch size for the node-level dataloaders. If not provided, is automatically determined based on `--edge_batch_size`. |
--n_sampled_neighbors | integer | Number of neighbors that are sampled during model training from the spatial neighborhood graph. If set to -1, all direct neighbors are included. |
Clustering options
Name | Type & Properties | Description |
|---|---|---|
--obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions resolutions specified in '--leiden_resolution'. |
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Umap options
Name | Type & Properties | Description |
|---|---|---|
--obsm_umap | string | In which .obsm slot to store the resulting UMAP embedding. |
Neighbour calculation
Name | Type & Properties | Description |
|---|---|---|
--uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
--output_model | file required output | Directory to save the trained NicheCompass model. |
--output_obsm_embedding | string | Key of the obsm field where the latent / gene program representation of active gene programs will be stored after NicheCompass model training. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
modality: [ "rna" ]
input_obs_batch_key: [ "sample_id" ]
input_obs_covariates: [ "sample_id" ]
input_obsm_spatial_coords: [ "spatial" ]
var_input: [ "filter_with_hvg" ]
coord_type: [ "generic" ]
n_spatial_neighbors: [ 6 ]
delaunay: [ false ]
min_genes_per_gp: [ 1 ]
min_source_genes_per_gp: [ 0 ]
min_target_genes_per_gp: [ 0 ]
gene_expr_recon_dist: [ "nb" ]
log_variational: [ true ]
node_label_method: [ "one-hop-norm" ]
active_gp_thresh_ratio: [ 0.1 ]
active_gp_type: [ "separate" ]
n_addon_gp: [ 100 ]
random_state: [ 0 ]
n_epochs: [ 100 ]
n_epochs_all_gps: [ 25 ]
n_epochs_no_edge_recon: [ 0 ]
n_epochs_no_cat_covariates_contrastive: [ 5 ]
lr: [ 0.001 ]
weight_decay: [ 0.001 ]
edge_val_ratio: [ 0.1 ]
node_val_ratio: [ 0.1 ]
edge_batch_size: [ 256 ]
n_sampled_neighbors: [ -1 ]
obs_cluster: [ "nichecompass_leiden" ]
leiden_resolution: [ 1 ]
obsm_umap: [ "X_leiden_nichecompass_umap" ]
uns_neighbors: [ "nichecompass_neighbors" ]
obsp_neighbor_distances: [ "nichecompass_distances" ]
obsp_neighbor_connectivities: [ "nichecompass_connectivities" ]
output: "$id.$key.output.h5mu"
output_model: "$id.$key.output_model"
output_obsm_embedding: [ "nichecompass_latent" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_spatial.git \
-revision v0.6.0 \
-main-script target/nextflow/workflows/niche/nichecompass_leiden/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/niche/nichecompass_leidenopenpipeline_spatial v0.6.0