wrappers/tools/umap
Description
UMAP (Uniform Manifold Approximation and Projection), running either the GPU
(rapids-singlecell) or the CPU (scanpy/squidpy) variant, selected with --device_type.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--uns_neighbors | string | The `.uns` neighbors slot as output by the `find_neighbors` component. |
Compute
Name | Type & Properties | Description |
|---|---|---|
--device_type | string | Which implementation to run: the GPU (rapids-singlecell) variant or the CPU (scanpy/squidpy) variant of the component. Selecting `gpu` requires a CUDA-capable GPU; the component errors out if none is available (there is no automatic fallback to CPU). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output h5mu file. |
--obsm_output | string | In which .obsm slot to store the resulting UMAP embedding. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--min_dist | double | The effective minimum distance between embedded points. Smaller values will result in a more clustered/clumped embedding where nearby points on the manifold are drawn closer together, while larger values will result on a more even dispersal of points. The value should be set relative to the spread value, which determines the scale at which embedded points will be spread out. |
--spread | double | The effective scale of embedded points. In combination with `min_dist` this determines how clustered/clumped the embedded points are. |
--num_components | integer | The number of dimensions of the embedding. |
--max_iter | integer | The number of iterations (epochs) of the optimization. Called `n_epochs` in the original UMAP. |
--alpha | double | The initial learning rate for the embedding optimization. |
--negative_sample_rate | integer | The number of negative edge/1-simplex samples to use per positive edge/1-simplex sample in optimizing the low dimensional embedding. |
--init_pos | string | How to initialize the low dimensional embedding. Called `init` in the original UMAP. Options are: * `'auto'`: chooses `'spectral'` for n_samples < 1000000, `'random'` otherwise. * `'spectral'`: use a spectral embedding of the graph. * `'random'`: assign initial embedding positions at random. * `'paga'`: use the paga() layout as initial embedding positions. * Any key from `.obsm`. |
rapids-singlecell options
Name | Type & Properties | Description |
|---|---|---|
--random_state | integer | Seed used by the random number generator. |
Scanpy options
Name | Type & Properties | Description |
|---|---|---|
--gamma | double | Weighting applied to negative samples in low dimensional embedding optimization. Values higher than one will result in greater weight being given to negative samples. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
uns_neighbors: [ "neighbors" ]
device_type: [ "gpu" ]
output: "$id.$key.output.h5mu"
obsm_output: [ "X_umap" ]
min_dist: [ 0.5 ]
spread: [ 1 ]
num_components: [ 2 ]
alpha: [ 1 ]
negative_sample_rate: [ 5 ]
init_pos: [ "spectral" ]
random_state: [ 0 ]
gamma: [ 1 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/wrappers/tools/umap/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
wrappers/tools/umapopenpipeline_rapids v0.1.3
Uses
2 relationships