| Abstract: | The digitization of large collections of documents aims at more than just producing digital versions of scan originals. Digital media allows for new reader experiences, some of them are traditional like searching for data; and some others are new, e.g., linking content to online real-time news. These tasks, involving advanced image analysis techniques, may require as well intensive human focus, process, care and resources if the final output of a digitalization system is to lead to a successful and engaging reading experience. We define and use a system that allows these tasks to happen in an automatic manner.
In this paper we present techniques for enhancing the quality of digital collections of documents reconstructed from their basic scanned originals by providing a pleasant experience to the final reader. We achieve this goal by minimizing the consequences of occasionally poor initial scans on a document and then by adding information layers with online content.
We used this process to recapture 80 years of weekly magazines published by Time [1]. The historical collection is scanned, automatically processed by advanced document analysis components to extract articles, manually verified for accuracy, and converted in a form suitable for web access. We used the proposed process to enhance the visual quality the PDF files generated from the conversion system.
|