A co-learning framework for learning user search intents from rule-generated training data

Authors:
Jun Yan;Zeyu Zheng;Li Jiang;Yan Li;Shuicheng Yan;Zheng Chen
Affiliations:
Microsoft Research Asia, Beijing, China;Peking University, Beijing, China;Microsoft Corporation, Redmond, WA, USA;Microsoft Corporation, Redmond, WA, USA;National University of Singapore, Singapore, Singapore;Microsoft Research Asia, Beijing, China
Venue:
Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval
Year:
2010

Citing 1
Cited 0

Task Behaviors During Web Search: The Difficulty of Assigning Labels

HICSS '09 Proceedings of the 42nd Hawaii International Conference on System Sciences

Quantified Score

Hi-index	0.00

Visualization

Abstract

Learning to understand user search intents from their online behaviors is crucial for both Web search and online advertising. However, it is a challenging task to collect and label a sufficient amount of high quality training data for various user intents such as "compare products", "plan a travel", etc. Motivated by this bottleneck, we start with some user common sense, i.e. a set of rules, to generate training data for learning to predict user intents. The rule-generated training data are however hard to be used since these data are generally imperfect due to the serious data bias and possible data noises. In this paper, we introduce a Co-learning Framework (CLF) to tackle the problem of learning from biased and noisy rule-generated training data. CLF firstly generates multiple sets of possibly biased and noisy training data using different rules, and then trains the individual user search intent classifiers over different training datasets independently. The intermediate classifiers are then used to categorize the training data themselves as well as the unlabeled data. The confidently classified data by one classifier are added to other training datasets and the incorrectly classified ones are instead filtered out from the training datasets. The algorithmic performance of this iterative learning procedure is theoretically guaranteed.