dimred/tsne
Description
t-SNE (t-Distributed Stochastic Neighbor Embedding) is a dimensionality reduction technique used to visualize high-dimensional data in a low-dimensional space, revealing patterns and clusters by preserving local data similarities.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string required | Which modality from the input MuData file to process. |
--use_rep | string required | The `.obsm` slot to use as input for the tSNE computation. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file. |
--obsm_output | string | The .obsm key to use for storing the tSNE results. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata H5 files. By default no compression is applied. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--n_pcs | integer | The number of principal components to use for the tSNE computation. |
--perplexity | double | The perplexity is related to the number of nearest neighbors that is used in other manifold learning algorithms. Larger datasets usually require a larger perplexity. Consider selecting a value between 5 and 50. Different values can result in significantly different results. |
--min_dist | double | The effective minimum distance between embedded points. Smaller values will result in a more clustered/clumped embedding where nearby points on the manifold are drawn closer together, while larger values will result on a more even dispersal of points. The value should be set relative to the spread value, which determines the scale at which embedded points will be spread out. |
--metric | string | Distance metric to calculate neighbors on. |
--early_exaggeration | double | Controls how tight natural clusters in the original space are in the embedded space and how much space will be between them. For larger values, the space between natural clusters will be larger in the embedded space. Again, the choice of this parameter is not very critical. If the cost function increases during initial optimization, the early exaggeration factor or the learning rate might be too high. |
--learning_rate | double | The learning rate for t-SNE optimization. Typical values range between 10.0 and 1000.0. |
--random_state | integer | The random seed to use for the tSNE computation. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
obsm_output: [ "X_tsne" ]
n_pcs: [ 50 ]
perplexity: [ 30 ]
min_dist: [ 0.5 ]
metric: [ "euclidean" ]
early_exaggeration: [ 12 ]
learning_rate: [ 1000 ]
random_state: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.1.0 \
-main-script target/nextflow/dimred/tsne/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
dimred/tsneopenpipeline v4.1.0
Uses
0 relationships
No component dependencies found.