Paper Details: Downloads: 460
Serial Number: P1120812006
Title: A Statistical Approach for Qur'an Vowel Restoration
Authors: A.A. EL-Harby and M.A. EL-Shehawey and R.S. El-Barogy
Abstract: This paper presents an automatic system that has the ability to restore diacritics (vowels) for non-diacritic Qur’an words, using a unigram base-line model and a bigram Hidden Markov Model (HMM). The proposed system was very robust and reliable without using morphological analysis methods for diacritics restoration. It was found that the HMMs are useful tools for the task of diacritics restoration in Arabic language. The used technique is simple to apply and does not require any language specific knowledge to be embedded in the model. Qur’an was used as corpora; our system was implemented and also tested on many parts of Qur’an as training set. For instance, the proposed system was implemented on 1366 words starting from the beginning of the Qur'an, and the best performance was 94.3% word accuracy for a unigram model and 95.2% word accuracy for a bigram HMM model.
Keywords: Diacritics in Arabic language, Vowel Restoration, Statistical Model, HMM.
Journal/Conference: International Journal of Artificial Intelligence and Machine Learning
Volume: 8
Issue: 3
Submission Date: 4/1/2008 12:00:00 AM
Review Date: 6/1/2008 12:00:00 AM
Publishing Date: 12/1/2008 12:00:00 AM
Article Downloads: 460
Download:

Facebook