Mining Surprising Periodic Patterns

  • Authors:
  • Jiong Yang;Wei Wang;Philip S. Yu

  • Affiliations:
  • Computer Science Department, University of Illinois at Urbana Champaign, 201 N. Goodwin Ave., Urbana, IL 61801, USA. jioyang@cs.uiuc.edu;Computer Science Department, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA. weiwang@cs.unc.edu;IBM T. J. Watson Research Center, 19 Skyline Dr., Hawthorne, NY 10532, USA. psyu@us.ibm.com

  • Venue:
  • Data Mining and Knowledge Discovery
  • Year:
  • 2004

Quantified Score

Hi-index 0.00

Visualization

Abstract

In this paper, we focus on mining surprising periodic patterns in a sequence of events. In many applications, e.g., computational biology, an infrequent pattern is still considered very significant if its actual occurrence frequency exceeds the prior expectation by a large margin. The traditional metric, such as support, is not necessarily the ideal model to measure this kind of surprising patterns because it treats all patterns equally in the sense that every occurrence carries the same weight towards the assessment of the significance of a pattern regardless of the probability of occurrence. A more suitable measurement, information, is introduced to naturally value the degree of surprise of each occurrence of a pattern as a continuous and monotonically decreasing function of its probability of occurrence. This would allow patterns with vastly different occurrence probabilities to be handled seamlessly. As the accumulated degree of surprise of all repetitions of a pattern, the concept of information gain is proposed to measure the overall degree of surprise of the pattern within a data sequence. The bounded information gain property is identified to tackle the predicament caused by the violation of the downward closure property by the information gain measure and in turn provides an efficient solution to this problem. Furthermore, the user has a choice between specifying a minimum information gain threshold and choosing the number of surprising patterns wanted. Empirical tests demonstrate the efficiency and the usefulness of the proposed model.