Evaluation of gene structure prediction programs

被引:523
作者
Burset, M
Guigo, R
机构
[1] INST MUNICIPAL INVEST MED, DEPT MED INFORMAT, E-08003 BARCELONA, SPAIN
[2] UNIV BARCELONA, FAC BIOL, DEPT ESTAD, BARCELONA, SPAIN
关键词
D O I
10.1006/geno.1996.0298
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
We evaluate a number of computer programs designed to predict the structure of protein coding genes in genomic DNA sequences. Computational gene identification is set to play an increasingly important role in the development of the genome projects, as emphasis turns from mapping to large-scale sequencing. The evaluation presented here serves both to assess the current status of the problem and to identify the most promising approaches to ensure further progress. The programs analyzed were uniformly tested on a large set of vertebrate sequences with simple gene structure, and several measures of predictive accuracy were computed at the nucleotide, exon, and protein product levels. The results indicated that the predictive accuracy of the programs analyzed was lower than originally found. The accuracy was even lower when considering only those sequences that had recently been entered and that did not show any similarity to previously entered sequences. This indicates that the programs are overly dependent on the particularities of the examples they learn from. For most of the programs, accuracy in this test set ranged from 0.60 to 0.70 as measured by the Correlation Coefficient (where 1.0 corresponds to a perfect prediction and 0.0 is the value expected for a random prediction), and the average percentage of exons exactly identified was less than 50%. Only those programs including protein sequence database searches showed substantially greater accuracy. The accuracy of the programs was severely affected by relatively high rates of sequence errors. Since the set on which the programs were tested included only relatively short sequences with simple gene structure, the accuracy of the programs is likely to be even lower when used for large uncharacterized genomic sequences with complex structure. While in such cases, programs currently available may still be of great use in pinpointing the regions Likely to contain exons, they are far from being powerful enough to elucidate its genomic structure completely. (C) 1996 Academic Press, Inc.
引用
收藏
页码:353 / 367
页数:15
相关论文
共 47 条
[1]  
ALTSCHUL SF, 1990, J MOL BIOL, V215, P403, DOI 10.1006/jmbi.1990.9999
[2]  
ANDERBERG M, 1973, CLUSTER ANAL APPLICA
[3]  
[Anonymous], METHOD ENZYMOL
[4]  
[Anonymous], P 2 INT C BIOINF SUP
[5]   INTRINSIC AND EXTRINSIC APPROACHES FOR DETECTING GENES IN A BACTERIAL GENOME [J].
BORODOVSKY, M ;
RUDD, KE ;
KOONIN, EV .
NUCLEIC ACIDS RESEARCH, 1994, 22 (22) :4756-4767
[6]   GENMARK - PARALLEL GENE RECOGNITION FOR BOTH DNA STRANDS [J].
BORODOVSKY, M ;
MCININCH, J .
COMPUTERS & CHEMISTRY, 1993, 17 (02) :123-133
[7]   IDENTIFICATION AND ANALYSIS OF MULTIGENE FAMILIES BY COMPARISON OF EXON FINGERPRINTS [J].
BROWN, NP ;
WHITTAKER, AJ ;
NEWELL, WR ;
RAWLINGS, CJ ;
BECK, S .
JOURNAL OF MOLECULAR BIOLOGY, 1995, 249 (02) :342-359
[8]   PREDICTION OF HUMAN MESSENGER-RNA DONOR AND ACCEPTOR SITES FROM THE DNA-SEQUENCE [J].
BRUNAK, S ;
ENGELBRECHT, J ;
KNUDSEN, S .
JOURNAL OF MOLECULAR BIOLOGY, 1991, 220 (01) :49-65
[9]  
CLAVERIE JM, 1990, METHOD ENZYMOL, V183, P237
[10]  
DAYHOFF MO, 1978, ATLAS PROTEIN SEQ S5, V3, P1