workflows/ingestion/make_reference
Description
Build a transcriptomics reference into one of many formats.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--id | string required | ID of the reference. |
--genome_fasta | file required | Reference genome fasta. |
--transcriptome_gtf | file required | Reference transcriptome annotation. |
--ercc | file | ERCC sequence and annotation file. |
STAR Settings
Name | Type & Properties | Description |
|---|---|---|
--star_genome_sa_index_nbases | integer | Length (bases) of the SA pre-indexing string. Typically between 10 and 15. Longer strings will use much more memory, but allow faster searches. For small genomes, the parameter {genomeSAindexNbases must be scaled down to min(14, log2(GenomeLength)/2 - 1). |
BD Rhapsody Settings
Name | Type & Properties | Description |
|---|---|---|
--bdrhap_mitochondrial_contigs | string multiple | Names of the Mitochondrial contigs in the provided Reference Genome. Fragments originating from contigs other than these are identified as 'nuclear fragments' in the ATACseq analysis pipeline. |
--bdrhap_filtering_off | boolean_true | By default the input Transcript Annotation files are filtered based on the gene_type/gene_biotype attribute. Only features having the following attribute values are kept: - protein_coding - lncRNA - IG_LV_gene - IG_V_gene - IG_V_pseudogene - IG_D_gene - IG_J_gene - IG_J_pseudogene - IG_C_gene - IG_C_pseudogene - TR_V_gene - TR_V_pseudogene - TR_D_gene - TR_J_gene - TR_J_pseudogene - TR_C_gene If you have already pre-filtered the input Annotation files and/or wish to turn-off the filtering, please set this option to True. |
--bdrhap_wta_only_index | boolean_true | Build a WTA only index, otherwise builds a WTA + ATAC index. |
--bdrhap_extra_star_params | string | Additional parameters to pass to STAR when building the genome index. Specify exactly like how you would on the command line. |
Cellranger ARC options
Name | Type & Properties | Description |
|---|---|---|
--motifs_file | file | Path to file containing transcription factor motifs in JASPAR format. |
--non_nuclear_contigs | string multiple | Name(s) of contig(s) that do not have any chromatin structure, for example, mitochondria or plastids. These contigs are excluded from peak calling since the entire contig will be "open" due to a lack of chromatin structure. Leave empty if there are no such contigs. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--target | string multiple | Which reference indices to generate. |
--output_fasta | file output | Output genome sequence fasta. |
--output_gtf | file output | Output transcriptome annotation gtf. |
--output_cellranger | file output | Output index |
--output_cellranger_arc | file output | Output index |
--output_bd_rhapsody | file output | Output index |
--output_star | file output | Output index |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--subset_regex | string | Will subset the reference chromosomes using the given regex. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
id: "run"
star_genome_sa_index_nbases: [ 14 ]
bdrhap_mitochondrial_contigs: [ "chrM", "chrMT", "M", "MT" ]
target: [ "star" ]
output_fasta: "$id.$key.output_fasta.gz"
output_gtf: "$id.$key.output_gtf.gz"
output_cellranger: "$id.$key.output_cellranger.gz"
output_cellranger_arc: "$id.$key.output_cellranger_arc.gz"
output_bd_rhapsody: "$id.$key.output_bd_rhapsody.gz"
output_star: "$id.$key.output_star.gz"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/openpipeline.git \
-revision v4.2.0 \
-main-script target/nextflow/workflows/ingestion/make_reference/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
workflows/ingestion/make_referenceopenpipeline v4.2.0