| Abstract: | The objective of this paper is the segmentation of handwritten Arabic script into connected components using a combination between Hough
Transform (HT) and Mathematical Morphology (MM). We start by a presentation of these two methods, their applications and their limits. According to the complexity of the segmented document, different kinds of combinations between these two methods are done in order to extract blocks, words, sub-words and some times characters. The main contribution of MM is binarisation, elimination of background and noisy, construction and extraction of straight lines and connected components. The HT is used essentially for line detection. Different kinds of lines can be extracted from a document and can be useful for its segmentation: the script baseline, pre-printed lines and horizontal and vertical lines used for the identification of boxes. Pre-printed lines and boxes are generally used for user enter handwriting script. For a document, three levels of segmentations can be done: segmentation of document in blocks, blocks in connected components and connected components in characters. We are interested in this paper by the two first level of segmentation. Extracted connected components from an Arabic document can in many cases be isolated characters. The application of the two proposed methods and their combination is done on several kinds of documents. Handwritten multilingual bank check, newspapers and printed formulas. The extraction rate, on CENPARMI Arabic check database and on the IFN/ENIT database attempts 99.63% for block documents and 85% for sub-words
|