wrappers/preprocessing/pca
Description
Computes PCA coordinates, loadings and variance decomposition, running either
the GPU (rapids-singlecell) or the CPU (scanpy) variant of pca, selected
with --device_type.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file. When running the GPU variant (`--device_type gpu`), genes with zero expression must be filtered out beforehand (e.g. with `filter_genes`), as rapids-singlecell PCA raises an error otherwise. The CPU variant tolerates zero-expression genes. |
--modality | string | Which modality from the input MuData file to process. |
--layer | string | Use specified layer for expression values instead of the .X object from the modality. |
--var_input | string | Column name in .var matrix that will be used to select which genes to run the PCA on. |
Compute
Name | Type & Properties | Description |
|---|---|---|
--device_type | string | Which implementation to run: the GPU (rapids-singlecell) variant or the CPU (scanpy/squidpy) variant of the component. Selecting `gpu` requires a CUDA-capable GPU; the component errors out if none is available (there is no automatic fallback to CPU). |
Options
Name | Type & Properties | Description |
|---|---|---|
--num_components | integer | Number of principal components to compute. Defaults to 50, or 1 - minimum dimension size of selected representation. |
--chunked | boolean_true | If True, perform an incremental PCA on segments of a predefined size. Setting this flag automatically implies zero centering. Must be specified together with --chunk_size. |
--chunk_size | integer | Number of observations to include in each chunk. Required if chunked=True was passed. |
--random_state | integer | Used to set the initial states for the optimization. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Output h5mu file. |
--obsm_output | string | In which .obsm slot to store the resulting embedding. |
--varm_output | string | In which .varm slot to store the resulting loadings matrix. |
--uns_output | string | In which .uns slot to store the resulting variance objects. |
--overwrite | boolean | Allow overwriting .obsm, .varm and .uns slots. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
device_type: [ "gpu" ]
output: "$id.$key.output.h5mu"
obsm_output: [ "X_pca" ]
varm_output: [ "pca_loadings" ]
uns_output: [ "pca_variance" ]
overwrite: [ false ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline_rapids.git \
-revision v0.1.3 \
-main-script target/nextflow/wrappers/preprocessing/pca/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
wrappers/preprocessing/pcaopenpipeline_rapids v0.1.3
Uses
2 relationships