Automated Metadata and Instance Extraction from News Web Sites

Authors:
Srinivas Vadrevu;Saravanakumar Nagarajan;Fatih Gelgi;Hasan Davulcu
Affiliations:
Arizona State University;Arizona State University;Arizona State University;Arizona State University
Venue:
WI '05 Proceedings of the 2005 IEEE/WIC/ACM International Conference on Web Intelligence
Year:
2005

Citing 0
Cited 3

Personal News RSS Feeds Generation Using Existing News Feeds

ICWE '9 Proceedings of the 9th International Conference on Web Engineering
Semantic search in the World News domain using automatically extracted metadata files

Knowledge-Based Systems
Improving web data annotations with spreading activation

WISE'05 Proceedings of the 6th international conference on Web Information Systems Engineering

Quantified Score

Hi-index	0.00

Visualization

Abstract

Over the past few years World Wide Web has established as a vital resource for news. With the continuous growth in the number of available news Web sites and the diversity in their presentation of content, there is an increasing need to organize the news related information on the Web and keep track of it. In this paper, we present automated techniques for extracting metadata instance information by organizing and mining a set of news Web sites. We develop algorithms that detect and utilize HTML regularities in the Web documents to turn them into hierarchical semantic structures encoded as XML. The tree-mining algorithms that we present identify key domain concepts and their taxonomical relationships. We also extract semi-structured concept instances annotated with their labels whenever they are available. We report experimental evaluation for the news domain to demonstrate the efficacy of our algorithms.