Automatic extraction rules generation based on XPath pattern learning

  • Authors:
  • Jingwei Zhang;Can Zhang;Weining Qian;Aoying Zhou

  • Affiliations:
  • Institute of Massive Computing, East China Normal University, Shanghai, China;Institute of Massive Computing, East China Normal University, Shanghai, China;Institute of Massive Computing, East China Normal University, Shanghai, China;Institute of Massive Computing, East China Normal University, Shanghai, China

  • Venue:
  • WISS'10 Proceedings of the 2010 international conference on Web information systems engineering
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

Web forums have become important information sources on the Web due to their rich content contributed by millions of Internet users every day. Data extraction from Web pages is a key but cumbersome step for data analysis because of significant human intervention. Web forums have fairly regular structures which allow us to generate extraction rules automatically according to their paths. In this paper, we introduce formal expressions for XPath patterns and pattern mapping rules, and advise machine learning methods to generate extraction rules for automatic data extraction from Web forums. The experimental results on real-life Web forums show good feasibility and accuracy for forum data.