dataflow/split_h5mu_train_test
Description
Split mudata object into training and testing (and validation) datasets based on observations into separate mudata objects.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | The input (query) data to be labeled. Should be a .h5mu file. |
--modality | string | Which modality to process. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output_train | file required output | The output training data in mudata format. |
--output_test | file required output | The output testing data in mudata format. |
--output_val | file output | The output validation data in mudata format. |
--compression | string |
Split arguments
Name | Type & Properties | Description |
|---|---|---|
--test_size | double | The proportion of the dataset to include in the test split. |
--val_size | double | The proportion of the dataset to include in the validation split. |
--shuffle | boolean_true | Whether or not to shuffle the data before splitting. |
--random_state | integer | The seed used by the random number generator. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output_train: "$id.$key.output_train.h5mu"
output_test: "$id.$key.output_test.h5mu"
output_val: "$id.$key.output_val.h5mu"
test_size: [ 0.2 ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision 2.0.0 \
-main-script target/nextflow/dataflow/split_h5mu_train_test/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
dataflow/split_h5mu_train_testopenpipeline 2.0.0
Uses
0 relationships
No component dependencies found.