filter/filter_with_scrublet
Description
Doublet detection using the Scrublet method (Wolock, Lopez and Klein, 2019).
The method tests for potential doublets by using the expression profiles of
cells to generate synthetic potential doubles which are tested against cells.
The method returns a "doublet score" on which it calls for potential doublets.
For the source code please visit https://github.com/AllonKleinLab/scrublet.
For 10x we expect the doublet rates to be:
Multiplet Rate (%) - # of Cells Loaded - # of Cells Recovered
~0.4% ~800 ~500
~0.8% ~1,600 ~1,000
~1.6% ~3,200 ~2,000
~2.3% ~4,800 ~3,000
~3.1% ~6,400 ~4,000
~3.9% ~8,000 ~5,000
~4.6% ~9,600 ~6,000
~5.4% ~11,200 ~7,000
~6.1% ~12,800 ~8,000
~6.9% ~14,400 ~9,000
~7.6% ~16,000 ~10,000
Arguments
Name | Type & Properties | Description |
|---|---|---|
--input | file required | Input h5mu file |
--modality | string | Which modality from the input MuData file to process. |
--layer | string | Input layer to use as data for calculating doublets. .X is used not specified. |
--output | file output | Output h5mu file. |
--obs_name_filter | string | In which .obs slot to store a boolean array corresponding to which observations should be filtered out. |
--do_subset | boolean_true | Whether to subset before storing the output. |
--obs_name_doublet_score | string | Name of the doublet scores column in the obs slot of the returned object. |
--expected_doublet_rate | double | The estimated fraction of doublets as from the experimental setup. |
--stdev_doublet_rate | double | Uncertainty in the expected doublet rate. |
--n_neighbors | integer | Number of neighbors used to construct the KNN classifier of observed transcriptomes and simulated doublets. |
--sim_doublet_ratio | double | Number of doublets to simulate relative to the number of observed transcriptomes. |
--min_counts | integer | The number of minimal UMI counts per cell that have to be present for initial cell detection. |
--min_cells | integer | The number of cells in which UMIs for a gene were detected. |
--min_gene_variablity_percent | double | Used for gene filtering prior to PCA. Keep the most highly variable genes (in the top min_gene_variability_pctl percentile), as measured by the v-statistic [Klein et al., Cell 2015]. |
--num_pca_components | integer | Number of principal components to use during PCA dimensionality reduction. |
--distance_metric | string | The distance metric used for computing similarities. |
--scrublet_score_threshold | double | Manual doublet score threshold. Cells with a doublet score above this value are classified as doublets. If not provided, the threshold is determined automatically by Scrublet. |
--allow_automatic_threshold_detection_fail | boolean_true | When scrublet fails to automatically determine the double score threshold, allow the component to continue and set the output columns to NA. |
--output_compression | string | Compression format to use for the output AnnData and/or Mudata H5 files. By default no compression is applied. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
modality: [ "rna" ]
output: "$id.$key.output.h5mu"
obs_name_filter: [ "filter_with_scrublet" ]
obs_name_doublet_score: [ "scrublet_doublet_score" ]
min_counts: [ 2 ]
min_cells: [ 3 ]
min_gene_variablity_percent: [ 85 ]
num_pca_components: [ 30 ]
distance_metric: [ "euclidean" ]
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/filter/filter_with_scrublet/main.nf \
-params-file params.yaml Relationships
Used by
Current component
Uses
No component dependencies found.