wrappers/preprocessing/bbknn
Description
Compute a batch-balanced nearest-neighbor graph of observations, running
either the GPU (rapids-singlecell) or the CPU (bbknn) variant of
bbknn, selected with --device_type.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file. |
--modality | string | Which modality from the input MuData file to process. |
--obsm_input | string | Which .obsm slot to use as a starting embedding (use_rep). |
--batch_key | string required | Which .obs column holds the batch assignment. Neighbors are searched within each batch separately and then merged. |
Compute
Name | Type & Properties | Description |
|---|---|---|
--device_type | string | Which implementation to run: the GPU (rapids-singlecell) variant or the CPU (scanpy/squidpy) variant of the component. Selecting `gpu` requires a CUDA-capable GPU; the component errors out if none is available (there is no automatic fallback to CPU). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output h5mu file containing the found neighbors. |
--uns_output | string | In which .uns slot to store the neighbor graph metadata. |
--obsp_distances | string | In which .obsp slot to store the distance matrix between the resulting neighbors. |
--obsp_connectivities | string | In which .obsp slot to store the connectivities matrix between the resulting neighbors. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Options
Name | Type & Properties | Description |
|---|---|---|
--neighbors_within_batch | integer | How many top neighbors to report for each batch. Total number of neighbors per cell is this value times the number of batches. |
--n_pcs | integer | Number of dimensions of the `--obsm_input` representation to use. If not set, all available dimensions are used. |
--trim | integer | Trim the neighbors of each cell to these many top connectivities. May help with population independence and improve the tidiness of clustering. If not set, defaults to ten times the total number of neighbors per cell. |
rapids-singlecell options
Name | Type & Properties | Description |
|---|---|---|
--overwrite | boolean | Allow overwriting the .uns/.obsp output slots if they already exist. |
--metric | string | The distance metric to use. |
--algorithm | string | The KNN query algorithm to use (provided by cuVS). See https://docs.rapids.ai/api/cuvs/stable/ for details. |
--random_state | integer | A random seed. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
obsm_input: [ "X_pca" ]
device_type: [ "gpu" ]
output: "$id.$key.output.h5mu"
uns_output: [ "neighbors" ]
obsp_distances: [ "distances" ]
obsp_connectivities: [ "connectivities" ]
neighbors_within_batch: [ 3 ]
overwrite: [ false ]
metric: [ "euclidean" ]
algorithm: [ "brute" ]
random_state: [ 0 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/wrappers/preprocessing/bbknn/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
wrappers/preprocessing/bbknnopenpipeline_rapids v0.1.3