Extracting semi-structured data through examples
Proceedings of the eighth international conference on Information and knowledge management
DEByE - Date extraction by example
Data & Knowledge Engineering
Hi-index | 0.00 |
In this paper, we propose an innovative approach to extracting semi-structured data from Web sources. The idea is to collect a couple of example objects from the user and to use this information to extract new objects from new pages or texts. We propose a top-down strategy that extracts complex objects decomposing them in objects less complex, until atomic objects have been extracted. Through experimentation, we demonstrate that with a small number of given examples our strategy is able to extract most of the objects present in a Web source given as input.