preprocessing/bbknn
Description
Compute a batch-balanced nearest-neighbor graph of observations on the GPU.
Wraps rapids-singlecell's rsc.pp.bbknn. For each cell it finds the nearest
neighbors within every batch separately and then merges them, correcting for
batch effects in the neighbor graph. The resulting graph is the basis for
downstream embedding and clustering (e.g. UMAP, Leiden) and is stored in the
same .uns/.obsp slots as a regular neighbor graph.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input -i | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--obsm_input | string | Which .obsm slot to use as a starting embedding (use_rep). |
--batch_key | string required | Which .obs column holds the batch assignment. Neighbors are searched within each batch separately and then merged. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output -o | file required output | Output h5mu file containing the found neighbors. |
--uns_output | string | In which .uns slot to store the neighbor graph metadata. |
--obsp_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
--overwrite | boolean | Allow overwriting the .uns/.obsp output slots if they already exist. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--neighbors_within_batch | integer | How many top neighbors to report for each batch. Total number of neighbors per cell is this value times the number of batches. |
--n_pcs | integer | Number of dimensions of the `--obsm_input` representation to use. If not set, all available dimensions are used. |
--metric | string | The distance metric to use. |
--algorithm | string | The KNN query algorithm to use (provided by cuVS). See https://docs.rapids.ai/api/cuvs/stable/ for details. - brute: Brute-force exact search computing all pairwise distances. Exact but the most expensive; best for smaller datasets. - cagra: GPU graph-based approximate search optimized for high query throughput on large datasets. - ivfflat: Inverted-file index that clusters vectors into lists and only searches the nearest lists. Approximate, faster than brute. - ivfpq: Inverted-file index combined with product quantization of the vectors. Approximate with a low memory footprint; scales to very large datasets. - mg_ivfflat: Multi-GPU variant of ivfflat. - mg_ivfpq: Multi-GPU variant of ivfpq. |
--trim | integer | Trim the neighbors of each cell to these many top connectivities. May help with population independence and improve the tidiness of clustering. If not set, defaults to ten times the total number of neighbors per cell. |
--random_state | integer | A random seed. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
obsm_input: [ "X_pca" ]
output: "$id.$key.output"
uns_output: [ "neighbors" ]
obsp_distances: [ "distances" ]
obsp_connectivities: [ "connectivities" ]
overwrite: [ false ]
neighbors_within_batch: [ 3 ]
metric: [ "euclidean" ]
algorithm: [ "brute" ]
random_state: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/preprocessing/bbknn/main.nf \
-params-file params.yaml Relationships
Used by
1 relationships
Current component
preprocessing/bbknnopenpipeline_rapids v0.1.3
Uses
0 relationships
No component dependencies found.