preprocessing/neighbors
Description
Compute a nearest-neighbor graph of observations on the GPU.
Wraps rapids-singlecell's rsc.pp.neighbors, whose neighbor search relies on
cuVS for fast (approximate) KNN search. The resulting graph is the basis for
downstream embedding and clustering (e.g. UMAP, Leiden). Connectivities can be
estimated with the UMAP fuzzy simplicial set ('umap'), an adaptive Gaussian
kernel ('gauss'), or the PhenoGraph Jaccard index ('jaccard').
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--obsm_input | string | Which .obsm slot to use as a starting embedding (use_rep). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file containing the found neighbors. |
--uns_output | string | In which .uns slot to store the neighbor graph metadata. |
--obsp_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--num_neighbors | integer | The size of the local neighborhood (in terms of number of neighboring data points) used for manifold approximation. Larger values result in more global views of the manifold, while smaller values result in more local data being preserved. In general values should be in the range 2 to 100. |
--n_pcs | integer | Number of dimensions of the `--obsm_input` representation to use. If not set, all available dimensions are used. |
--metric | string | The distance metric to use. |
--algorithm | string | The KNN query algorithm to use (provided by cuVS). See https://docs.rapids.ai/api/cuvs/stable/ for details. - brute: Brute-force exact search computing all pairwise distances. Exact but the most expensive; best for smaller datasets. - cagra: GPU graph-based approximate search optimized for high query throughput on large datasets. - ivfflat: Inverted-file index that clusters vectors into lists and only searches the nearest lists. Approximate, faster than brute. - ivfpq: Inverted-file index combined with product quantization of the vectors. Approximate with a low memory footprint; scales to very large datasets. - nn_descent: Graph-based approximate method that iteratively refines a KNN graph from a random initialization. |
--method | string | Method for computing connectivities. 'umap' uses the UMAP fuzzy simplicial set, 'gauss' uses an adaptive Gaussian kernel, 'jaccard' uses the PhenoGraph Jaccard index. |
--random_state | integer | A random seed. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
obsm_input: [ "X_pca" ]
output: "$id.$key.output"
uns_output: [ "neighbors" ]
obsp_distances: [ "distances" ]
obsp_connectivities: [ "connectivities" ]
num_neighbors: [ 15 ]
metric: [ "euclidean" ]
algorithm: [ "brute" ]
method: [ "umap" ]
random_state: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/preprocessing/neighbors/main.nf \
-params-file params.yaml Relationships
Used by
2 relationships
Current component
preprocessing/neighborsopenpipeline_rapids v0.1.3
Uses
0 relationships
No component dependencies found.