Superimposed Code-Based Indexing Method for Extracting MCTs from XML Documents

Authors:
Wenxin Liang;Takeshi Miki;Haruo Yokota
Affiliations:
CREST, Japan Science and Technology Agency (JST), and Global Scientific Information and Computing Center, Tokyo Institute of Technology,;Nomura Research Institute,;Department of Computer Science, Tokyo Institute of Technology, and Global Scientific Information and Computing Center, Tokyo Institute of Technology,
Venue:
DEXA '08 Proceedings of the 19th international conference on Database and Expert Systems Applications
Year:
2008

Citing 17
Cited 1

Fast algorithms for finding nearest common ancestors

SIAM Journal on Computing
On finding lowest common ancestors: simplification and parallelization

SIAM Journal on Computing
Finding lowest common ancestors in arbitrarily directed trees

Information Processing Letters
Applications of Path Compression on Balanced Trees

Journal of the ACM (JACM)
Selective families, superimposed codes, and broadcasting on unknown radio networks

SODA '01 Proceedings of the twelfth annual ACM-SIAM symposium on Discrete algorithms
DBXplorer: enabling keyword search over relational databases

Proceedings of the 2002 ACM SIGMOD international conference on Management of data
Querying XML Documents Made Easy: Nearest Concept Queries

Proceedings of the 17th International Conference on Data Engineering
On finding lowest common ancestors in trees

STOC '73 Proceedings of the fifth annual ACM symposium on Theory of computing
XRANK: ranked keyword search over XML documents

Proceedings of the 2003 ACM SIGMOD international conference on Management of data
DBXplorer: A System for Keyword-Based Search over Relational Databases

ICDE '02 Proceedings of the 18th International Conference on Data Engineering
Keyword Searching and Browsing in Databases using BANKS

ICDE '02 Proceedings of the 18th International Conference on Data Engineering
Efficient keyword search for smallest LCAs in XML databases

Proceedings of the 2005 ACM SIGMOD international conference on Management of data
Kalchas: a dynamic XML search engine

Proceedings of the 14th ACM international conference on Information and knowledge management
Keyword Proximity Search in XML Trees

IEEE Transactions on Knowledge and Data Engineering
Discover: keyword search in relational databases

VLDB '02 Proceedings of the 28th international conference on Very Large Data Bases
XSEarch: a semantic search engine for XML

VLDB '03 Proceedings of the 29th international conference on Very large data bases - Volume 29
Schema-free XQuery

VLDB '04 Proceedings of the Thirtieth international conference on Very large data bases - Volume 30

A Low-Storage-Consumption XML Labeling Method for Efficient Structural Information Extraction

DEXA '09 Proceedings of the 20th International Conference on Database and Expert Systems Applications

Quantified Score

Hi-index	0.00

Visualization

Abstract

With the exponential increase in the amount of XML data on the Internet, information retrieval techniques on tree-structured XML documents such as keyword search become important. The search results for this retrieval technique are often represented by minimum connecting trees (MCTs) rooted at the lowest common ancestors (LCAs) of the nodes containing all the search keywords. Recently, effective methods such as the stack-based algorithm for generating the lowest grouped distance MCTs (GDMCTs), which derive a more compact representation of the query results, have been proposed. However, when the XML documents and the number of search keywords become large, these methods are still expensive. To achieve more efficient algorithms for extracting MCTs, especially lowest GDMCTs, we first consider two straightforward LCA detection methods: keyword B+trees with Dewey-order labels and superimposed code-based indexing methods. Then, we propose a method for efficiently detecting the LCAs, which combines the two straightforward indexing methods for LCA detection. We also present an effective solution for the false drop problem caused by the superimposed code. Finally, the proposed LCA detection methods are applied to generate the lowest GDMCTs. We conduct detailed experiments to evaluate the benefits of our proposed algorithms and show that the proposed combined method can completely solve the false drop problem and outperforms the stack-based algorithm in extracting the lowest GDMCTs.