workflows/integration/scanorama_leiden
Description
Run scanorama integration followed by neighbour calculations, leiden clustering and run umap on the result.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the sample. |
--input | file required | Path to the sample. |
--layer | string | use specified layer for expression values instead of the .X object from the modality. |
--modality | string | Which modality to process. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Destination path to the output. |
Neighbour calculation
Name | Type & Properties | Description |
|---|---|---|
--uns_neighbors | string | In which .uns slot to store various neighbor output objects. |
--obsp_neighbor_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_neighbor_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
Scanorama integration options
Name | Type & Properties | Description |
|---|---|---|
--obs_batch | string | Column name discriminating between your batches. |
--obsm_input | string | .obsm slot that points to embedding to run scanorama on. |
--obsm_output | string | The name of the field in adata.obsm where the integrated embeddings will be stored after running this function. Defaults to X_scanorama. |
--knn | integer | Number of nearest neighbors to use for matching. |
--batch_size | integer | The batch size used in the alignment vector computation. Useful when integrating very large (>100k samples) datasets. Set to large value that runs within available memory. |
--sigma | double | Correction smoothing parameter on Gaussian kernel. |
--approx | boolean | Use approximate nearest neighbors with Python annoy; greatly speeds up matching runtime. |
--alpha | double | Alignment score minimum cutoff |
Clustering options
Name | Type & Properties | Description |
|---|---|---|
--obs_cluster | string | Prefix for the .obs keys under which to add the cluster labels. Newly created columns in .obs will be created from the specified value for '--obs_cluster' suffixed with an underscore and one of the resolutions specified in '--leiden_resolution'. |
--leiden_resolution | double multiple | Control the coarseness of the clustering. Higher values lead to more clusters. |
Umap options
Name | Type & Properties | Description |
|---|---|---|
--obsm_umap | string | In which .obsm slot to store the resulting UMAP embedding. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
layer: [ "log_normalized" ]
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
uns_neighbors: [ "scanorama_integration_neighbors" ]
obsp_neighbor_distances: [ "scanorama_integration_distances" ]
obsp_neighbor_connectivities: [ "scanorama_integration_connectivities" ]
obs_batch: [ "sample_id" ]
obsm_input: [ "X_pca" ]
obsm_output: [ "X_scanorama" ]
knn: [ 20 ]
batch_size: [ 5000 ]
sigma: [ 15 ]
approx: [ true ]
alpha: [ 0.1 ]
obs_cluster: [ "scanorama_integration_leiden" ]
leiden_resolution: [ 1 ]
obsm_umap: [ "X_leiden_scanorama_umap" ]
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/workflows/integration/scanorama_leiden/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/integration/scanorama_leidenopenpipeline v4.2.0