Published July 1999 | Version public
Journal Article

A Dictionary-Based Approach for Gene Annotation

Abstract

This paper describes a fast and fully automated dictionary-based approach to gene annotation and exon prediction. Two dictionaries are constructed, one from the nonredundant protein OWL database and the other from the dbEST database. These dictionaries are used to obtain O(1) time lookups of tuples in the dictionaries (4 tuples for the OWL database and 11 tuples for the dbEST database). These tuples can be used to rapidly find the longest matches at every position in an input sequence to the database sequences. Such matches provide very useful information pertaining to locating common segments between exons, alternative splice sites, and frequency data of long tuples for statistical purposes. These dictionaries also provide the basis for both homology determination, and statistical approaches to exon prediction.

Additional Information

© 1999 Mary Ann Liebert, Inc.

Additional details

Identifiers

Eprint ID
74981
Resolver ID
CaltechAUTHORS:20170309-113000311

Dates

Created
2017-03-10
Created from EPrint's datestamp field
Updated
2021-11-15
Created from EPrint's last_modified field