filter/filter_with_scrublet
Description
Doublet detection using the Scrublet method (Wolock, Lopez and Klein, 2019).
The method tests for potential doublets by using the expression profiles of
cells to generate synthetic potential doubles which are tested against cells.
The method returns a "doublet score" on which it calls for potential doublets.
For the source code please visit https://github.com/AllonKleinLab/scrublet.
For 10x we expect the doublet rates to be:
Multiplet Rate (%) - # of Cells Loaded - # of Cells Recovered
~0.4% ~800 ~500
~0.8% ~1,600 ~1,000
~1.6% ~3,200 ~2,000
~2.3% ~4,800 ~3,000
~3.1% ~6,400 ~4,000
~3.9% ~8,000 ~5,000
~4.6% ~9,600 ~6,000
~5.4% ~11,200 ~7,000
~6.1% ~12,800 ~8,000
~6.9% ~14,400 ~9,000
~7.6% ~16,000 ~10,000
Arguments
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string | |
--layer | string | Input layer to use as data for calculating doublets. .X is used not specified. |
--output | file output | Output h5mu file. |
--output_compression | string | The compression format to be used on the output h5mu object. |
--obs_name_filter | string | In which .obs slot to store a boolean array corresponding to which observations should be filtered out. |
--do_subset | boolean_true | Whether to subset before storing the output. |
--obs_name_doublet_score | string | Name of the doublet scores column in the obs slot of the returned object. |
--min_counts | integer | The number of minimal UMI counts per cell that have to be present for initial cell detection. |
--min_cells | integer | The number of cells in which UMIs for a gene were detected. |
--min_gene_variablity_percent | double | Used for gene filtering prior to PCA. Keep the most highly variable genes (in the top min_gene_variability_pctl percentile), as measured by the v-statistic [Klein et al., Cell 2015]. |
--num_pca_components | integer | Number of principal components to use during PCA dimensionality reduction. |
--distance_metric | string | The distance metric used for computing similarities. |
--allow_automatic_threshold_detection_fail | boolean_true | When scrublet fails to automatically determine the double score threshold, allow the component to continue and set the output columns to NA. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
obs_name_filter: [ "filter_with_scrublet" ]
obs_name_doublet_score: [ "scrublet_doublet_score" ]
min_counts: [ 2 ]
min_cells: [ 3 ]
min_gene_variablity_percent: [ 85 ]
num_pca_components: [ 30 ]
distance_metric: [ "euclidean" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision 1.0.2 \
-main-script target/nextflow/filter/filter_with_scrublet/main.nf \
-params-file params.yaml Relationships
Used by
Current component
Uses
No component dependencies found.