utils/gtffilter
Description
Clean up a GTF file before it is used to build reference indices.
The following records are removed:
Records on a sequence (contig) that is not present in the genome FASTA.
RSEM refuses transcripts with exons on such a sequence.Records without a
transcript_idattribute.Duplicated records: records that share their sequence, feature type, start, end,
strand andtranscript_idwith an earlier record. Only the first one is kept.
STAR adds up duplicated exons, which makes the transcript lengths in its
transcriptome BAM disagree with the transcript FASTA used by Salmon.
Comment lines (starting with #) are kept. The tool fails when a record does not
have 9 tab-separated columns, or when no records are left after filtering.
Inputs
Name | Type & Properties | Description |
|---|---|---|
--input | file required | GTF file to filter (plain or gzipped). |
--genome_fasta | file required | Genome FASTA file (plain or gzipped). The sequence names (the first word of each header) determine which GTF records are kept. |
Outputs
Name | Type & Properties | Description |
|---|---|---|
--output | file required output | Filtered GTF file (plain). |
Run this component
Run the following command to execute this component with Nextflow:
cat > params.yaml <<'EOM'
output: "$id.$key.output.gtf"
id: "run"
publish_dir: "output/"
EOM
nextflow run https://packages.viash-hub.com/vsh/bulk-rnaseq.git \
-revision v0.3.2 \
-main-script target/nextflow/utils/gtffilter/main.nf \
-params-file params.yaml Relationships
Used by
Current component
Uses
No component dependencies found.