Paper Details: Downloads: 490
Serial Number: P1120820001
Title: NTCA: A Novel Text Clustering AlgorithmBuild on Cellular automata Based local search and K-Means Algorithm for Identifying the Protein Coding Regions in Genomic DNA
Authors: P.KIRAN SREE, I.RAMESH BABU, USHA DEVI,
Abstract: Genes carry the instructions for making proteins that are found in a cell as a specific sequence of nucleotides that are found in DNA molecules. But, the regions of these genes that code for proteins may occupy only a small region of the sequence. Identifying the coding regions play a vital role in understanding these genes. This paper presents a new text clustering algorithm based on Cellular Automata Based Local Search and K-Means (LSKM) for identifying these protein coding regions and explains the characteristic of this algorithm theoretically. Experimental results confirm the scalability of the proposed LSKM based classifier to any datasets irrespective of the number of classes, tuples and attributes. We note an increase in accuracy of more than 5.2%, over any existing standard algorithms for addressing this problem. This was the first algorithm to identify protein coding regions in mixed and non overlapping exon-inton boundary DNA sequences also
Keywords: Cellular Automata, DNA, Local search, K-Means Algorithm
Journal/Conference: International Journal of Artificial Intelligence and Machine Learning
Volume:
Issue:
Submission Date: 5/13/2008 12:00:00 AM
Review Date: 6/9/2008 12:00:00 AM
Publishing Date: 6/28/2008 12:00:00 AM
Article Downloads: 490
Download:

Facebook