arriba
Gene fusion
RNA-Seq
Description
Detect gene fusions from RNA-Seq data
Inputs
Name | Type & Properties | Description |
|---|---|---|
--bam -x | file required | File in SAM/BAM/CRAM format with main alignments as generated by STAR (Aligned.out.sam). Arriba extracts candidate reads from this file. |
--genome -a | file required | FastA file with genome sequence (assembly). The file may be gzip-compressed. An index with the file extension .fai must exist only if CRAM files are processed. |
--gene_annotation -g | file required | GTF file with gene annotation. The file may be gzip-compressed. |
--known_fusions -k | file | File containing known/recurrent fusions. Some cancer entities are often characterized by fusions between the same pair of genes. In order to boost sensitivity, a list of known fusions can be supplied using this parameter. The list must contain two columns with the names of the fused genes, separated by tabs. |
--blacklist -b | file | File containing blacklisted events (recurrent artifacts and transcripts observed in healthy tissue). |
--structural_variants -d | file | Tab-separated file with coordinates of structural variants found using whole-genome sequencing data. These coordinates serve to increase sensitivity towards weakly expressed fusions and to eliminate fusions with low evidence. |
--tags -t | file | Tab-separated file containing fusions to annotate with tags in the 'tags' column. The first two columns specify the genes; the third column specifies the tag. The file may be gzip-compressed. |
--protein_domains -p | file | File in GFF3 format containing coordinates of the protein domains of genes. The protein domains retained in a fusion are listed in the column 'retained_protein_domains'. The file may be gzip-compressed. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--fusions -o | file required output | Output file with fusions that have passed all filters. |
--fusions_discarded -O | file output | Output file with fusions that were discarded due to filtering. |
Arguments
Name | Type & Properties | Description |
|---|---|---|
--max_genomic_breakpoint_distance -D | long | When a file with genomic breakpoints obtained via whole-genome sequencing is supplied via the --structural_variants parameter, this parameter determines how far a genomic breakpoint may be away from a transcriptomic breakpoint to consider it as a related event. For events inside genes, the distance is added to the end of the gene; for intergenic events, the distance threshold is applied as is. Default: 100000. |
--strandedness -s | string | Whether a strand-specific protocol was used for library preparation, and if so, the type of strandedness (auto/yes/no/reverse). When unstranded data is processed, the strand can sometimes be inferred from splice-patterns. But in unclear situations, stranded data helps resolve ambiguities. Default: auto |
--interesting_contigs -i | string multiple | List of interesting contigs. Fusions between genes on other contigs are ignored. Contigs can be specified with or without the prefix "chr". Asterisks (*) are treated as wild-cards. Default: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 X Y AC_* NC_* |
--viral_contigs -v | string multiple | List of viral contigs. Asterisks (*) are treated as wild-cards. Default: AC_* NC_* |
--disable_filters -f | string multiple | List of filters to disable. By default all filters are enabled. |
--max_e_value -E | double | Arriba estimates the number of fusions with a given number of supporting reads which one would expect to see by random chance. If the expected number of fusions (e-value) is higher than this threshold, the fusion is discarded by the 'relative_support' filter. Note: Increasing this threshold can dramatically increase the number of false positives and may increase the runtime of resource-intensive steps. Fractional values are possible. Default: 0.300000 |
--min_supporting_reads -S | integer | The 'min_support' filter discards all fusions with fewer than this many supporting reads (split reads and discordant mates combined). Default: 2 |
--max_mismappers -m | double | When more than this fraction of supporting reads turns out to be mismappers, the 'mismappers' filter discards the fusion. Default: 0.800000 |
--max_homolog_identity -L | double | Genes with more than the given fraction of sequence identity are considered homologs and removed by the 'homologs' filter. Default: 0.300000 |
--homopolymer_length -H | integer | The 'homopolymer' filter removes breakpoints adjacent to homopolymers of the given length or more. Default: 6 |
--read_through_distance -R | integer | The 'read_through' filter removes read-through fusions where the breakpoints are less than the given distance away from each other. Default: 10000 |
--min_anchor_length -A | integer | Alignment artifacts are often characterized by split reads coming from only one gene and no discordant mates. Moreover, the split reads only align to a short stretch in one of the genes. The 'short_anchor' filter removes these fusions. This parameter sets the threshold in bp for what the filter considers short. Default: 23 |
--many_spliced_events -M | integer | The 'many_spliced' filter recovers fusions between genes that have at least this many spliced breakpoints. Default: 4 |
--max_kmer_content -K | double | The 'low_entropy' filter removes reads with repetitive 3-mers. If the 3-mers make up more than the given fraction of the sequence, then the read is discarded. Default: 0.600000 |
--max_mismatch_pvalue -V | double | The 'mismatches' filter uses a binomial model to calculate a p-value for observing a given number of mismatches in a read. If the number of mismatches is too high, the read is discarded. Default: 0.010000 |
--fragment_length -F | integer | When paired-end data is given, the fragment length is estimated automatically and this parameter has no effect. But when single-end data is given, the mean fragment length should be specified to effectively filter fusions that arise from hairpin structures. Default: 200 |
--max_reads -U | integer | Subsample fusions with more than the given number of supporting reads. This improves performance without compromising sensitivity, as long as the threshold is high. Counting of supporting reads beyond the threshold is inaccurate, obviously. Default: 300 |
--quantile -Q | double | Highly expressed genes are prone to produce artifacts during library preparation. Genes with an expression above the given quantile are eligible for filtering by the 'in_vitro' filter. Default: 0.998000 |
--exonic_fraction -e | double | The breakpoints of false-positive predictions of intragenic events are often both in exons. True predictions are more likely to have at least one breakpoint in an intron, because introns are larger. If the fraction of exonic sequence between two breakpoints is smaller than the given fraction, the 'intragenic_exonic' filter discards the event. Default: 0.330000 |
--top_n -T | integer | Only report viral integration sites of the top N most highly expressed viral contigs. Default: 5 |
--covered_fraction -C | double | Ignore virally associated events if the virus is not fully expressed, i.e., less than the given fraction of the viral contig is transcribed. Default: 0.050000 |
--max_itd_length -l | integer | Maximum length of internal tandem duplications. Note: Increasing this value beyond the default can impair performance and lead to many false positives. Default: 100 |
--min_itd_allele_fraction -z | double | Required fraction of supporting reads to report an internal tandem duplication. Default: 0.070000 |
--min_itd_supporting_reads -Z | integer | Required absolute number of supporting reads to report an internal tandem duplication. Default: 10 |
--skip_duplicate_marking -u | boolean_true | Instead of performing duplicate marking itself, Arriba relies on duplicate marking by a preceding program using the BAM_FDUP flag. This makes sense when unique molecular identifiers (UMI) are used. |
--extra_information -X | boolean_true | To reduce the runtime and file size, by default, the columns 'fusion_transcript', 'peptide_sequence', and 'read_identifiers' are left empty in the file containing discarded fusion candidates (see parameter -O). When this flag is set, this extra information is reported in the discarded fusions file. |
--fill_gaps -I | boolean_true | If assembly of the fusion transcript sequence from the supporting reads is incomplete (denoted as '...'), fill the gaps using the assembly sequence wherever possible. |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
fusions: "$id.$key.fusions.tsv"
fusions_discarded: "$id.$key.fusions_discarded.tsv"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.1.0 \
-main-script target/nextflow/arriba/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
arribabiobox v0.1.0
Uses
0 relationships
No component dependencies found.