THESUS: Organizing Web document collections based on link semantics

Authors:
Maria Halkidi;Benjamin Nguyen;Iraklis Varlamis;Michalis Vazirgiannis
Affiliations:
76 Patision Street, Athens University of Economics and Business, Athens, Greece;Domaine de Voluceau, INRIA, 78153, Le Chesnay, France;76 Patision Street, Athens University of Economics and Business, Athens, Greece;76 Patision Street, Athens University of Economics and Business, Athens, Greece
Venue:
The VLDB Journal — The International Journal on Very Large Data Bases
Year:
2003

Citing 24
Cited 25

The R*-tree: an efficient and robust access method for points and rectangles

SIGMOD '90 Proceedings of the 1990 ACM SIGMOD international conference on Management of data
Concept based query expansion

SIGIR '93 Proceedings of the 16th annual international ACM SIGIR conference on Research and development in information retrieval
A rough set model of information retrieval

Fundamenta Informaticae - Special issue: to the memory of Prof. Helena Rasiowa
Web document clustering: a feasibility demonstration

Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval
Automatic resource compilation by analyzing hyperlink structure and associated text

WWW7 Proceedings of the seventh international conference on World Wide Web 7
The anatomy of a large-scale hypertextual Web search engine

WWW7 Proceedings of the seventh international conference on World Wide Web 7
Fast and effective text mining using linear-time document clustering

KDD '99 Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data mining
On the merits of building categorization systems by supervised clustering

KDD '99 Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data mining
Authoritative sources in a hyperlinked environment

Journal of the ACM (JACM)
Hierarchical classification of Web content

SIGIR '00 Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval
Efficient and tumble similar set retrieval

SIGMOD '01 Proceedings of the 2001 ACM SIGMOD international conference on Management of data
Evaluating strategies for similarity search on the web

Proceedings of the 11th international conference on World Wide Web
Using web structure for classifying and describing web pages

Proceedings of the 11th international conference on World Wide Web
Introduction to Modern Information Retrieval

Introduction to Modern Information Retrieval
On Clustering Validation Techniques

Journal of Intelligent Information Systems
Mining the Web's Link Structure

Computer
Knowledge Acquisition Via Incremental Conceptual Clustering

Machine Learning
An Information-Theoretic Definition of Similarity

ICML '98 Proceedings of the Fifteenth International Conference on Machine Learning
Incremental Clustering for Mining in a Data Warehousing Environment

VLDB '98 Proceedings of the 24rd International Conference on Very Large Data Bases
Web Document Searching Using Enhanced Hyperlink Semantics Based on XML

IDEAS '01 Proceedings of the International Database Engineering & Applications Symposium
Robust Hyperlinks Cost Just Five Words Each

Robust Hyperlinks Cost Just Five Words Each
Verbs semantics and lexical selection

ACL '94 Proceedings of the 32nd annual meeting on Association for Computational Linguistics
Pattern Recognition, Third Edition

Pattern Recognition, Third Edition
Using information content to evaluate semantic similarity in a taxonomy

IJCAI'95 Proceedings of the 14th international joint conference on Artificial intelligence - Volume 1

SEWeP: using site semantics and a taxonomy to enhance the Web personalization process

Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining
Web personalization integrating content semantics and navigational patterns

Proceedings of the 6th annual ACM international workshop on Web information and data management
Automatically Generating an E-textbook on the Web

World Wide Web
A software infrastructure for RSS deployment and linking on the web

WebMedia '05 Proceedings of the 11th Brazilian Symposium on Multimedia and the web
Category ranking for personalized search

Data & Knowledge Engineering
Integrating recommendation models for improved web page prediction accuracy

ACSC '08 Proceedings of the thirty-first Australasian conference on Computer science - Volume 74
A comparative evaluation of different link types on enhancing document clustering

Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval
Semantically driven snippet selection for supporting focused web searches

Data & Knowledge Engineering
French EuroWordNet Lexical Database Improvements

CICLing '07 Proceedings of the 8th International Conference on Computational Linguistics and Intelligent Text Processing
State of the Art in Semantic Focused Crawlers

ICCSA '09 Proceedings of the International Conference on Computational Science and Its Applications: Part II
Comparison of similarity measures for clustering Turkish documents

Intelligent Data Analysis
Navigating among search results: an information content approach

WISE'07 Proceedings of the 8th international conference on Web information systems engineering
An integrated model for next page access prediction

International Journal of Knowledge and Web Intelligence
A framework for discovering and classifying ubiquitous services in digital health ecosystems

Journal of Computer and System Sciences
Web personalization by assimilating usage data and semantics expressed in ontology terms

Proceedings of the International Conference & Workshop on Emerging Trends in Technology
An integrated technique for web site usage semantic analysis: the organ system

Journal of Web Engineering
Discriminating biased web manipulations in terms of link oriented measures

ISCIS'05 Proceedings of the 20th international conference on Computer and Information Sciences
Enriching short text representation in microblog for clustering

Frontiers of Computer Science in China
Web directory construction using lexical chains

NLDB'05 Proceedings of the 10th international conference on Natural Language Processing and Information Systems
IKUM: an integrated web personalization platform based on content structures and user behavior

ITWP'03 Proceedings of the 2003 international conference on Intelligent Techniques for Web Personalization
Factors affecting web page similarity

ECIR'05 Proceedings of the 27th European conference on Advances in Information Retrieval Research
Semi-automatic creation and maintenance of web resources with webtopic

EWMF'05/KDO'05 Proceedings of the 2005 joint international conference on Semantics, Web and Mining
Introducing semantics in web personalization: the role of ontologies

EWMF'05/KDO'05 Proceedings of the 2005 joint international conference on Semantics, Web and Mining
Category labelling for automatic classification scheme generation

FDIA'07 Proceedings of the 1st BCS IRSG conference on Future Directions in Information Access
Measuring web page similarity based on textual and visual properties

ICAISC'12 Proceedings of the 11th international conference on Artificial Intelligence and Soft Computing - Volume Part II

Quantified Score

Hi-index	0.00

Visualization

Abstract

The requirements for effective search and management of the WWW are stronger than ever. Currently Web documents are classified based on their content not taking into account the fact that these documents are connected to each other by links. We claim that a page’s classification is enriched by the detection of its incoming links’ semantics. This would enable effective browsing and enhance the validity of search results in the WWW context. Another aspect that is underaddressed and strictly related to the tasks of browsing and searching is the similarity of documents at the semantic level. The above observations lead us to the adoption of a hierarchy of concepts (ontology) and a thesaurus to exploit links and provide a better characterization of Web documents. The enhancement of document characterization makes operations such as clustering and labeling very interesting. To this end, we devised a system called THESUS. The system deals with an initial sets of Web documents, extracts keywords from all pages’ incoming links, and converts them to semantics by mapping them to a domain’s ontology. Then a clustering algorithm is applied to discover groups of Web documents. The effectiveness of the clustering process is based on the use of a novel similarity measure between documents characterized by sets of terms. Web documents are organized into thematic subsets based on their semantics. The subsets are then labeled, thereby enabling easier management (browsing, searching, querying) of the Web. In this article, we detail the process of this system and give an experimental analysis of its results.