gffread
gff
conversion
validation
filtering
Description
Validate, filter, convert and perform various other operations on GFF files.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | A reference file in either the GFF3, GFF2 or GTF format. |
--chr_mapping -m | file | <chr_replace> is a name mapping table for converting reference sequence names, having this 2-column format: <original_ref_ID> <new_ref_ID>. |
--seq_info -s | file | <seq_info.fsize> is a tab-delimited file providing this info for each of the mapped sequences: <seq-name> <seq-length> <seq-description> (useful for --description option with mRNA/EST/protein mappings). |
--genome -g | file | Full path to a multi-fasta file with the genomic sequences for all input mappings, OR a directory with single-fasta files (one per genomic sequence, with file names matching sequence names). |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--outfile -o | file required output | Write the output records into <outfile>. |
--force_exons | boolean_true | Make sure that the lowest level GFF features are considered "exon" features. |
--gene2exon | boolean_true | For single-line genes not parenting any transcripts, add an exon feature spanning the entire gene (treat it as a transcript). |
--t_adopt | boolean_true | Try to find a parent gene overlapping/containing a transcript that does not have any explicit gene Parent. |
--decode -D | boolean_true | Decode url encoded characters within attributes. |
--merge_exons -Z | boolean_true | Merge very close exons into a single exon (when intron size<4). |
--junctions -j | boolean_true | Output the junctions and the corresponding transcripts. |
--spliced_exons -w | file output | Write a fasta file with spliced exons for each transcript. |
--w_add | integer | For the --spliced_exons option, extract additional <N> bases both upstream and downstream of the transcript boundaries. |
--w_nocds | boolean_true | For --spliced_exons, disable the output of CDS info in the FASTA file. |
--spliced_cds -x | file | Write a fasta file with spliced CDS for each GFF transcript. |
--tr_cds -y | file | Write a protein fasta file with the translation of CDS for each record. |
--w_coords -W | boolean_true | For --spliced_exons, --spliced_cds and -tr_cds options, write in the FASTA defline all the exon coordinates projected onto the spliced sequence. |
--stop_dot -S | boolean_true | For --tr_cds option, use '*' instead of '.' as stop codon translation. |
--id_version -L | boolean_true | Ensembl GTF to GFF3 conversion, adds version to IDs. |
--trackname -t | string | Use <trackname> in the 2nd column of each GFF/GTF output line. |
--gtf_output -T | boolean_true | Main output will be GTF instead of GFF3. |
--bed | boolean_true | Output records in BED format instead of default GFF3. |
--tlf | boolean_true | Output "transcript line format" which is like GFF but with exons and CDS related features stored as GFF attributes in the transcript feature line, like this: exoncount=N;exons=<exons>;CDSphase=<N>;CDS=<CDScoords> <exons> is a comma-delimited list of exon_start-exon_end coordinates; <CDScoords> is CDS_start:CDS_end coordinates or a list like <exons>. |
--table | string multiple | Output a simple tab delimited format instead of GFF, with columns having the values of GFF attributes given in <attrlist>; special pseudo-attributes (prefixed by @) are recognized: @id, @geneid, @chr, @start, @end, @strand, @numexons, @exons, @cds, @covlen, @cdslen If any of --spliced_exons/--tr_cds/--spliced_cds FASTA output files are enabled, the same fields (excluding @id) are appended to the definition line of corresponding FASTA records. |
--expose_dups -E -v | boolean_true | Expose (warn about) duplicate transcript IDs and other potential problems with the given GFF/GTF records. |
Options
Name | Type & Properties | Description |
|---|---|---|
--ids | file | Discard records/transcripts if their IDs are not listed in <IDs.lst>. |
--nids | file | Discard records/transcripts if their IDs are listed in <IDs.lst>. |
--maxintron -i | integer | Discard transcripts having an intron larger than <maxintron>. |
--minlen -l | integer | Discard transcripts shorter than <minlen> bases. |
--range -r | string | Only show transcripts overlapping coordinate range <start>..<end> (on chromosome/contig <chr>, strand <strand> if provided). |
--strict_range -R | boolean_true | For --range option, discard all transcripts that are not fully contained within the given range. |
--jmatch | string | Only output transcripts matching the given junction. |
--no_single_exon -U | boolean_true | Discard single-exon transcripts. |
--coding -C | boolean_true | Coding only: discard mRNAs that have no CDS features. |
--nc | boolean_true | Non-coding only: discard mRNAs that have CDS features. |
--ignore_locus | boolean_true | Discard locus features and attributes found in the input. |
--description -A | boolean_true | Use the description field from <seq_info.fsize> and add it as the value for a 'descr' attribute to the GFF record. |
Sorting
Name | Type & Properties | Description |
|---|---|---|
--sort_alpha | boolean_true | Chromosomes (reference sequences) are sorted alphabetically. |
--sort_by | file | Sort the reference sequences by the order in which their names are given in the <refseq.lst> file. |
Misc options
Name | Type & Properties | Description |
|---|---|---|
--keep_attrs -F | boolean_true | Keep all GFF attributes (for non-exon features). |
--keep_exon_attrs | boolean_true | For -F option, do not attempt to reduce redundant exon/CDS attributes. |
--no_exon_attrs -G | boolean_true | Do not keep exon attributes, move them to the transcript feature (for GFF3 output). |
--attrs | string | Only output the GTF/GFF attributes listed in <attr-list> which is a comma delimited list of attribute names to. |
--keep_genes | boolean_true | In transcript-only mode (default), also preserve gene records. |
--keep_comments | boolean_true | For GFF3 input/output, try to preserve comments. |
--process_other -O | boolean_true | process other non-transcript GFF records (by default non-transcript records are ignored). |
--rm_stop_codons -V | boolean_true | Discard any mRNAs with CDS having in-frame stop codons (requires --genome). |
--adj_cds_start -H | boolean_true | For --rm_stop_codons option, check and adjust the starting CDS phase if the original phase leads to a translation with an in-frame stop codon. |
--opposite_strand -B | boolean_true | For -V option, single-exon transcripts are also checked on the opposite strand (requires --genome). |
--coding_status -P | boolean_true | Add transcript level GFF attributes about the coding status of each transcript, including partialness or in-frame stop codons (requires --genome). |
--add_hasCDS | boolean_true | Add a "hasCDS" attribute with value "true" for transcripts that have CDS features. |
--adj_stop | boolean_true | Stop codon adjustment: enables --coding_status and performs automatic adjustment of the CDS stop coordinate if premature or downstream. |
--rm_noncanon -N | boolean_true | Discard multi-exon mRNAs that have any intron with a non-canonical splice site consensus (i.e. not GT-AG, GC-AG or AT-AC). |
--complete_cds -J | boolean_true | Discard any mRNAs that either lack initial START codon or the terminal STOP codon, or have an in-frame stop codon (i.e. only print mRNAs with a complete CDS). |
--no_pseudo | boolean_true | Filter out records matching the 'pseudo' keyword. |
--in_bed | boolean_true | Input should be parsed as BED format (automatic if the input filename ends with .bed*). |
--in_tlf | boolean_true | Input GFF-like one-line-per-transcript format without exon/CDS features (see --tlf option below); automatic if the input filename ends with .tlf). |
--stream | boolean_true | Fast processing of input GFF/BED transcripts as they are received (no sorting, exons must be grouped by transcript in the input data). |
Clustering
Name | Type & Properties | Description |
|---|---|---|
--merge -M | boolean_true | Cluster the input transcripts into loci, discarding "redundant" transcripts (those with the same exact introns and fully contained or equal boundaries). |
--dupinfo -d | file | For --merge option, write duplication info to file <dupinfo>. |
--cluster_only | boolean_true | Same as --merge but without discarding any of the "duplicate" transcripts, only create "locus" features. |
--rm_redundant -K | boolean_true | For --merge option: also discard as redundant the shorter, fully contained transcripts (intron chains matching a part of the container). |
--no_boundary -Q | boolean_true | For --merge option, no longer require boundary containment when assessing redundancy (can be combined with --rm_redundant); only introns have to match for multi-exon transcripts, and >=80% overlap for single-exon transcripts. |
--no_overlap -Y | boolean_true | For --merge option, enforce --no_boundary but also discard overlapping single-exon transcripts, even on the opposite strand (can be combined with --rm_redudant). |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
outfile: "$id.$key.outfile.gff"
spliced_exons: "$id.$key.spliced_exons.fa"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/biobox.git \
-revision v0.3.1 \
-main-script target/nextflow/gffread/main.nf \
-params-file params.yaml Relationships
Used by
0 relationships
No components use this component.
Current component
gffreadbiobox v0.3.1
Uses
0 relationships
No component dependencies found.