Adaptive web-page content identification
Proceedings of the 9th annual ACM international workshop on Web information and data management
RENS --- Enabling a Robot to Identify a Person
ICIRA '09 Proceedings of the 2nd International Conference on Intelligent Robotics and Applications
Facilitating wrapper generation with page analysis
ISI'09 Proceedings of the 2009 IEEE international conference on Intelligence and security informatics
Peddling or creating? investigating the role of twitter in news reporting
ECIR'11 Proceedings of the 33rd European conference on Advances in information retrieval
An automatic web news article contents extraction system based on RSS feeds
Journal of Web Engineering
Story graphs: Tracking document set evolution using dynamic graphs
Intelligent Data Analysis - Dynamic Networks and Knowledge Discovery
Hi-index | 0.00 |
We developed and tested a heuristic technique for extracting the main article from news site Web pages. We construct the DOM tree of the page and score every node based on the amount of text, the number of links it contains and additional heuristics. The method is site-independent and does not use any language-based features. We tested our algorithm on a set of 1120 news article pages from 27 domains. Our algorithm achieved over 97% precision and 98% recall, and an average processing speed of under 15ms per page.