dimred/umap
Description
UMAP (Uniform Manifold Approximation and Projection) is a manifold learning technique suitable for visualizing high-dimensional data. Besides tending to be faster than tSNE, it optimizes the embedding such that it best reflects the topology of the data, which we represent throughout Scanpy using a neighborhood graph. tSNE, by contrast, optimizes the distribution of nearest-neighbor distances in the embedding such that these best match the distribution of distances in the high-dimensional space. We use the implementation of umap-learn [McInnes18]. For a few comparisons of UMAP with tSNE, see this preprint.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string | Which modality from the input MuData file to process. |
--uns_neighbors | string | The `.uns` neighbors slot as output by the `find_neighbors` component. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--obsm_output | string | The pre/postfix under which to store the UMAP results. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--min_dist | double | The effective minimum distance between embedded points. Smaller values will result in a more clustered/clumped embedding where nearby points on the manifold are drawn closer together, while larger values will result on a more even dispersal of points. The value should be set relative to the spread value, which determines the scale at which embedded points will be spread out. |
--spread | double | The effective scale of embedded points. In combination with `min_dist` this determines how clustered/clumped the embedded points are. |
--num_components | integer | The number of dimensions of the embedding. |
--max_iter | integer | The number of iterations (epochs) of the optimization. Called `n_epochs` in the original UMAP. Default is set to 500 if neighbors['connectivities'].shape[0] <= 10000, else 200. |
--alpha | double | The initial learning rate for the embedding optimization. |
--gamma | double | Weighting applied to negative samples in low dimensional embedding optimization. Values higher than one will result in greater weight being given to negative samples. |
--negative_sample_rate | integer | The number of negative edge/1-simplex samples to use per positive edge/1-simplex sample in optimizing the low dimensional embedding. |
--init_pos | string | How to initialize the low dimensional embedding. Called `init` in the original UMAP. Options are: * Any key from `.obsm` * `'paga'`: positions from `paga()` * `'spectral'`: use a spectral embedding of the graph * `'random'`: assign initial embedding positions at random. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
uns_neighbors: [ "neighbors" ]
output: "$id.$key.output.h5mu"
obsm_output: [ "umap" ]
min_dist: [ 0.5 ]
spread: [ 1 ]
num_components: [ 2 ]
alpha: [ 1 ]
gamma: [ 1 ]
negative_sample_rate: [ 5 ]
init_pos: [ "spectral" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.0.0 \
-main-script target/nextflow/dimred/umap/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
dimred/umapopenpipeline v4.0.0
Uses
0 relationships
No component dependencies found.