On the construction of a large scale Chinese web test collection

  • Authors:
  • Hongfei Yan;Chong Chen;Bo Peng;Xiaoming Li

  • Affiliations:
  • School of Electronics Engineering and Computer Science, Peking University, Beijing, P.R. China;School of Electronics Engineering and Computer Science, Peking University, Beijing, P.R. China;School of Electronics Engineering and Computer Science, Peking University, Beijing, P.R. China;School of Electronics Engineering and Computer Science, Peking University, Beijing, P.R. China

  • Venue:
  • AIRS'08 Proceedings of the 4th Asia information retrieval conference on Information retrieval technology
  • Year:
  • 2008

Quantified Score

Hi-index 0.00

Visualization

Abstract

The lack of a large scale Chinese test collection is an obstacle to the Chinese information retrieval development. In order to address this issue, we built such a collection composed of millions of Chinese web pages, known as the Chinese Web Test collection with 100 gigabyte (CWT100g) in data volume, which is the largest Chinese web test collection as of this writing, and has been used by several dozen research groups besides being adopted in the evaluation of the SEWM-2004 Chinese Web Track[1] and the HTRDPE-2004[2]. We present the total solution for constructing a large scale test collection like the CWT100g. Further, we found that: 1) the distribution of the number of pages within sites obeys a Zipf-like law instead of a power law proposed by Adamic and Huberman [3, 4]; 2) and an appropriate filtering method on host alias will economize resources for about 25% while crawling pages. The Zipf-like law and the method of filtering host alias proposed in the paper will facilitate both to model the Web and to perfect a search engine. Finally, we report on the results of the SEWM-2004 Chinese Web Track.