Coreex: content extraction from online news articles

Authors:
Jyotika Prasad;Andreas Paepcke
Affiliations:
Stanford University, Stanford, CA, USA;Stanford University, Stanford, CA, USA
Venue:
Proceedings of the 17th ACM conference on Information and knowledge management
Year:
2008

Citing 1
Cited 5

Adaptive web-page content identification

Proceedings of the 9th annual ACM international workshop on Web information and data management

RENS --- Enabling a Robot to Identify a Person

ICIRA '09 Proceedings of the 2nd International Conference on Intelligent Robotics and Applications
Facilitating wrapper generation with page analysis

ISI'09 Proceedings of the 2009 IEEE international conference on Intelligence and security informatics
Peddling or creating? investigating the role of twitter in news reporting

ECIR'11 Proceedings of the 33rd European conference on Advances in information retrieval
An automatic web news article contents extraction system based on RSS feeds

Journal of Web Engineering
Story graphs: Tracking document set evolution using dynamic graphs

Intelligent Data Analysis - Dynamic Networks and Knowledge Discovery

Quantified Score

Hi-index	0.00

Visualization

Abstract

We developed and tested a heuristic technique for extracting the main article from news site Web pages. We construct the DOM tree of the page and score every node based on the amount of text, the number of links it contains and additional heuristics. The method is site-independent and does not use any language-based features. We tested our algorithm on a set of 1120 news article pages from 27 domains. Our algorithm achieved over 97% precision and 98% recall, and an average processing speed of under 15ms per page.