Host-IP clustering technique for deep web characterization

Authors:
Denis Shestakov;Tapio Salakoski
Affiliations:
Helsinki University of Technology, Finland;University of Turku, Finland
Venue:
Proceedings of the 2010 ACM Symposium on Applied Computing
Year:
2010

Citing 4
Cited 0

Characterization of national Web domains

ACM Transactions on Internet Technology (TOIT)
Accessing the deep web

Communications of the ACM - ACM at sixty: a look back in time
On building a search interface discovery system

RED'09 Proceedings of the 2nd international conference on Resource discovery
On estimating the scale of national deep web

DEXA'07 Proceedings of the 18th international conference on Database and Expert Systems Applications

Quantified Score

Hi-index	0.01

Visualization

Abstract

The part of the Web, known as the deep Web, is to date relatively unexplored and even major characteristics such as number of searchable databases on the Web is somewhat disputable. In this paper, we are aimed at more accurate estimation of main parameters of the deep Web by sampling one national web domain. We propose the Host-IP clustering sampling technique that addresses drawbacks of existing approaches to characterize the deep Web and report our findings based on the survey of Russian Web conducted in September 2006. Obtained estimates together with a proposed sampling method could be useful for further studies to handle data in the deep Web.