star/star_genome_generate
genome
index
align
Description
Create index for STAR
Input
Name | Type & Properties | Description |
|---|---|---|
--genome_fasta_files | file required multiple | Path(s) to the fasta files with the genome sequences, separated by spaces. These files should be plain text FASTA files, they *cannot* be zipped. |
--sjdb_gtf_file | file | Path to the GTF file with annotations |
--sjdb_overhang | integer | Length of the donor/acceptor sequence on each side of the junctions, ideally = (mate_length - 1) |
--sjdb_gtf_chr_prefix | string | Prefix for chromosome names in a GTF file (e.g. 'chr' for using ENSMEBL annotations with UCSC genomes) |
--sjdb_gtf_feature_exon | string | Feature type in GTF file to be used as exons for building transcripts |
--sjdb_gtf_tag_exon_parent_transcript | string | GTF attribute name for parent transcript ID (default "transcript_id" works for GTF files) |
--sjdb_gtf_tag_exon_parent_gene | string | GTF attribute name for parent gene ID (default "gene_id" works for GTF files) |
--sjdb_gtf_tag_exon_parent_gene_name | string multiple | GTF attribute name for parent gene name |
--sjdb_gtf_tag_exon_parent_gene_type | string multiple | GTF attribute name for parent gene type |
--limit_genome_generate_ram | long | Maximum available RAM (bytes) for genome generation |
--genome_sa_index_nbases | integer | Length (bases) of the SA pre-indexing string. Typically between 10 and 15. Longer strings will use much more memory, but allow faster searches. For small genomes, this parameter must be scaled down to min(14, log2(GenomeLength)/2 - 1). |
--genome_chr_bin_nbits | integer | Defined as log2(chrBin), where chrBin is the size of the bins for genome storage. Each chromosome will occupy an integer number of bins. For a genome with large number of contigs, it is recommended to scale this parameter as min(18, log2[max(GenomeLength/NumberOfReferences,ReadLength)]). |
--genome_sa_sparse_d | integer | Suffux array sparsity, i.e. distance between indices. Use bigger numbers to decrease needed RAM at the cost of mapping speed reduction. |
--genome_suffix_length_max | integer | Maximum length of the suffixes, has to be longer than read length. Use -1 for infinite length. |
--genome_transform_type | string | Type of genome transformation None ... no transformation Haploid ... replace reference alleles with alternative alleles from VCF file (e.g. consensus allele) Diploid ... create two haplotypes for each chromosome listed in VCF file, for genotypes 1|2, assumes perfect phasing (e.g. personal genome) |
--genome_transform_vcf | file | path to VCF file for genome transformation |
Output
Name | Type & Properties | Description |
|---|---|---|
--index | file required output | STAR index directory. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
index: "$id.$key.index"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.4.2 \
-main-script target/nextflow/star/star_genome_generate/main.nf \
-params-file params.yaml Relationships
Used by
3 relationships, 2 components
Current component
star/star_genome_generatebiobox v0.4.2
Uses
0 relationships
No component dependencies found.