Evaluation of gene structure prediction programs

被引：523

作者：

Burset, M

Guigo, R

机构：

[1] INST MUNICIPAL INVEST MED, DEPT MED INFORMAT, E-08003 BARCELONA, SPAIN

[2] UNIV BARCELONA, FAC BIOL, DEPT ESTAD, BARCELONA, SPAIN

来源：

GENOMICS | 1996年 / 34卷 / 03期

关键词：

D O I：

10.1006/geno.1996.0298

中图分类号：

Q81 [生物工程学（生物技术）]; Q93 [微生物学];

学科分类号：

071005 ; 0836 ; 090102 ; 100705 ;

摘要：

We evaluate a number of computer programs designed to predict the structure of protein coding genes in genomic DNA sequences. Computational gene identification is set to play an increasingly important role in the development of the genome projects, as emphasis turns from mapping to large-scale sequencing. The evaluation presented here serves both to assess the current status of the problem and to identify the most promising approaches to ensure further progress. The programs analyzed were uniformly tested on a large set of vertebrate sequences with simple gene structure, and several measures of predictive accuracy were computed at the nucleotide, exon, and protein product levels. The results indicated that the predictive accuracy of the programs analyzed was lower than originally found. The accuracy was even lower when considering only those sequences that had recently been entered and that did not show any similarity to previously entered sequences. This indicates that the programs are overly dependent on the particularities of the examples they learn from. For most of the programs, accuracy in this test set ranged from 0.60 to 0.70 as measured by the Correlation Coefficient (where 1.0 corresponds to a perfect prediction and 0.0 is the value expected for a random prediction), and the average percentage of exons exactly identified was less than 50%. Only those programs including protein sequence database searches showed substantially greater accuracy. The accuracy of the programs was severely affected by relatively high rates of sequence errors. Since the set on which the programs were tested included only relatively short sequences with simple gene structure, the accuracy of the programs is likely to be even lower when used for large uncharacterized genomic sequences with complex structure. While in such cases, programs currently available may still be of great use in pinpointing the regions Likely to contain exons, they are far from being powerful enough to elucidate its genomic structure completely. (C) 1996 Academic Press, Inc.

引用

页码：353 / 367

页数：15

共 47 条

[1]

ALTSCHUL SF, 1990, J MOL BIOL, V215, P403, DOI 10.1006/jmbi.1990.9999

[2]

ANDERBERG M, 1973, CLUSTER ANAL APPLICA

[3]

[Anonymous], METHOD ENZYMOL

[4]

[Anonymous], P 2 INT C BIOINF SUP