labels_transfer/xgboost
Description
Performs label transfer from reference to query using XGBoost classifier
Input dataset (query) arguments
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The query data to transfer the labels to. Should be a .h5mu file. |
--modality | string | Which modality to use. |
--input_obsm_features | string | The `.obsm` key of the embedding to use for the classifier's inference. If not provided, the `.X` slot will be used instead. Make sure that embedding was obtained in the same way as the reference embedding (e.g. by the same model or preprocessing). |
Reference dataset arguments
Name | Type & Properties | Description |
|---|---|---|
--reference | file | The reference data to train classifiers on. |
--reference_obsm_features | string | The `.obsm` key of the embedding to use for the classifier's training. If not provided, the `.X` slot will be used instead. Make sure that embedding was obtained in the same way as the query embedding (e.g. by the same model or preprocessing). |
--reference_obs_targets | string multiple | The `.obs` key(s) of the target labels to tranfer. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | The query data in .h5mu format with predicted labels transfered from the reference. |
--output_obs_predictions | string multiple | In which `.obs` slots to store the predicted information. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_pred"` suffix. |
--output_obs_probability | string multiple | In which `.obs` slots to store the probability of the predictions. If provided, must have the same length as `--reference_obs_targets`. If empty, will default to the `reference_obs_targets` combined with the `"_probability"` suffix. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata objects. By default no compression is applied. |
Execution arguments
Name | Type & Properties | Description |
|---|---|---|
--force_retrain -f | boolean_true | Retrain models on the reference even if model_output directory already has trained classifiers. WARNING! It will rewrite existing classifiers for targets in the model_output directory! |
--use_gpu | boolean | Use GPU during models training and inference (recommended). |
--verbosity -v | integer | The verbosity level for evaluation of the classifier from the range [0,2] |
--model_output | file output | Output directory for model |
--output_uns_parameters | string | The key in `uns` slot of the output AnnData object to store the parameters of the XGBoost classifier. |
Learning parameters
Name | Type & Properties | Description |
|---|---|---|
--learning_rate --eta | double | Step size shrinkage used in update to prevents overfitting. Range: [0,1]. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--min_split_loss --gamma | double | Minimum loss reduction required to make a further partition on a leaf node of the tree. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--max_depth -d | integer | Maximum depth of a tree. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--min_child_weight | integer | Minimum sum of instance weight (hessian) needed in a child. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--max_delta_step | double | Maximum delta step we allow each leaf output to be. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--subsample | double | Subsample ratio of the training instances. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--sampling_method | string | The method to use to sample the training instances. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--colsample_bytree | double | Fraction of columns to be subsampled. Range (0, 1]. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--colsample_bylevel | double | Subsample ratio of columns for each level. Range (0, 1]. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--colsample_bynode | double | Subsample ratio of columns for each node (split). Range (0, 1]. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--reg_lambda --lambda | double | L2 regularization term on weights. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--reg_alpha --alpha | double | L1 regularization term on weights. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
--scale_pos_weight | double | Control the balance of positive and negative weights, useful for unbalanced classes. See https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster for the reference |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
reference_obs_targets:
[
"ann_level_1",
"ann_level_2",
"ann_level_3",
"ann_level_4",
"ann_level_5",
"ann_finest_level"
]
output: "$id.$key.output"
use_gpu: [ false ]
verbosity: [ 1 ]
model_output: "$id.$key.model_output"
output_uns_parameters: [ "xgboost_parameters" ]
learning_rate: [ 0.3 ]
min_split_loss: [ 0 ]
max_depth: [ 6 ]
min_child_weight: [ 1 ]
max_delta_step: [ 0 ]
subsample: [ 1 ]
sampling_method: [ "uniform" ]
colsample_bytree: [ 1 ]
colsample_bylevel: [ 1 ]
colsample_bynode: [ 1 ]
reg_lambda: [ 1 ]
reg_alpha: [ 0 ]
scale_pos_weight: [ 1 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v3.0.0 \
-main-script target/nextflow/labels_transfer/xgboost/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
labels_transfer/xgboostopenpipeline v3.0.0
Uses
0 relationships
No component dependencies found.