Phrase processing methods for Japanese text retrieval

Authors:
Noriko Kando;Kyo Kageura;Masaharu Yoshioka;Keizo Oyama
Affiliations:
Research and Development Department, Tokyo, Japan;Research and Development Department, Tokyo, Japan;Research and Development Department, Tokyo, Japan;Research and Development Department, Tokyo, Japan
Venue:
ACM SIGIR Forum
Year:
1998

Citing 0
Cited 3

A Corpus-Based Learning Method of Compound Noun Indexing Rules for Korean

Information Retrieval
Corpus-based learning of compound noun indexing

RANLPIR '00 Proceedings of the ACL-2000 workshop on Recent advances in natural language processing and information retrieval: held in conjunction with the 38th Annual Meeting of the Association for Computational Linguistics - Volume 11
Query structuring and expansion with two-stage term dependence for Japanese web retrieval

Information Retrieval

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper examines the effectiveness of different phrase identification and weighting methods for Japanese text retrieval in an operational information retrieval (IR) system, called NACSIS-IR. Based on our previous experiments, we used character-based indexing with positional information and word-or phrase-based query processing, which allowed us to implement sophisticated linguistic analysis on large-scale databases while maintaining adequate efficiency. The results of retrieval experiments on a large-scale Japanese test collection showed that the combination of enhanced phrase identification using patterns defined over part-of-speech tags and our algorithms Phrase2 and Phrase5 made a significant positive contribution to retrieval effectiveness. The paper also discusses indexing and phrase processing of Japanese or East Asian languages.