Weerayuth Kittichotirat. Development of Secondary Structure Analysis Tool for Identifying Starch Synthase Isoforms. Master's Degree(Bioinformatics). King Mongkut's University of Technology Thonburi. Library. : King Mongkut's University of Technology Thonburi, 2004.
Development of Secondary Structure Analysis Tool for Identifying Starch Synthase Isoforms
Organization :
King Mongkut's University of Technology Thonburi. School of Bioresources and Technology and School of Information Technology. Bioinformatics
Abstract:
Starch is an indispensable raw material in many industries. The utilization of starch mostly depends on its physical and chemical properties, which differ according to plant origins and their varieties. While native forms of starch have many uses, sometimes the manufacturers find them unsuitable for their industrial requirement. These problems are often handled by exposing the native form of starch to physical or chemical modification processes in order to tailor them appropriately according to the industrial needs. However, both physical and chemical modifications are cost intensive and generate high chemical waste. Moreover, large portion of starches are likely to be destroyed during the process. Thus, a major challenge would be to predict the effect of genetic changes on the functional properties of starch to produce designer starches that suit specific uses. The enzymes in starch biosynthesis pathway generally have a variety of isoforms that have specific roles, which are believed to affect starch properties. However, their precise roles have not been identified. Understanding of relationships between enzyme isoforms in starch biosynthesis process and properties of starch granule is therefore one of the main goals of plant molecular biologist. Because of the difficulties in obtaining experimental evidences for each enzyme isoform as they are usually instable, hard to purify to homogeneity and exist in low abundance, classifying them into families using computational method based on the presence of shared features is one of the common strategies for functional analysis. Since the mechanisms by which different isofoms catalyze the same reaction and yet generate polymer variation are not fully understood together with the evidences that structures are more conserved than sequences in evolution as well as secondary structure agreement between predicted secondary structures can be used to identify distantly related protein sequences, we present a novel method for classifying and identifying starch enzyme isoforms based on predicted secondary structures. We have developed a method that can automatically perform unsupervised cluster discovery from the secondary structure of starch biosynthesis enzyme isoforms data input based on their similarity. During each iteration, a Smith-Waterman dynamic programming alignment algorithm together with our novel scoring model is used to calculate all pairwise similarity of secondary structure dataset. The highest score pair is then used to construct a representing pattern where it will be further considered in the successive rounds of the algorithm, instead of using its parental secondary structures. Finally, the clustering of the isoforms result is displayed in the form of a tree in which isoforms with similar secondary structure occupy neighboring leaves. In addition, we The method was applied to analyze a set of 201 isoforms of all key starch biosynthesis enzymes from seven selected plants and showed that our novel similarity measures can be used to cluster structurally related enzyme isoforms together. I describe how this method can be used to identify enzyme isoforms that might have been misclassified. The visualization method can also be used to identify the structural features that might constitute the minimum catalytic unit of each group of structurally related enzyme , isoforms. On the secondary structure comparison analysis of 61 starch synthase isoforms, it was found that the rice starch synthase isoform V contain structural regions that exhibit similar features to both isoform IV group and isoform V group. From this result, it might be suggested that rice starch synthase isoform V enzyme is the link between starch synthases that were classified to be in isoform IV group and those to the isoform V group as the secondary structure analysis show that these two isoform groups are structurally related. The result also suggests that the region where the rice starch synthase isoform V show high secondary structural residue identities with the rest of the isoform V enzymes might constitute to the minimum catalytic unit of the starch synthase isoform V family.
Abstract:
The completions of many genome sequencing projects of many organisms have brought about the need for automated, large-scale and accurate function assignment of gene products. Many methods and databases have been developed and implemented to annotate the large amount of genomic sequence data. However, these methods usually use different algorithm and thus each has its own strength and weaknesses and all of them do not provide precise reliable estimates or confidence measures for their predictions. Jason el al., 2005 have develop the Bioverse database and computational biology framework which is an integrated, knowledge-base resource to facilitate the understanding of molecular and organismal biology. Automated functional annotation by various domain, site and family tools were performed on protein sequences in the Bioverse. These function predictions were used as the primary functional annotation of each protein molecule with the tool(s) that produce that specific prediction regarded as the prediction evidence of that functional assignment. I have taken a step further by developing a prediction confidence or correctness measure which makes use of the comparison between prediction and experimentally derived annotation. In my approach, Homo sapiens and Drosophila melanogaster were used as the model organism. Artificial neural network was used to learn the integration of quality measures from various tools to predict the new integrated correctness score and it can be used in the assignment of correctness score for hnctional predictions that do not have experimental evidences.