Load Shedding for Window Joins on Multiple Data Streams

Authors:
Yan-Nei Law;Carlo Zaniolo
Affiliations:
Bioinformatics Institute, 30 Biopolis Street, Singapore 138671. lawyn@bii.a-star.edu.sg;Computer Science Dept., UCLA, Los Angeles, CA 90095, USA. zaniolo@cs.ucla.edu
Venue:
ICDEW '07 Proceedings of the 2007 IEEE 23rd International Conference on Data Engineering Workshop
Year:
2007

Citing 0
Cited 2

A load shedding framework for XML stream joins

DEXA'10 Proceedings of the 21st international conference on Database and expert systems applications: Part I
Load shedding for multi-way stream joins based on arrival order patterns

Journal of Intelligent Information Systems

Quantified Score

Hi-index	0.00

Visualization

Abstract

We consider the problem of semantic load shedding for continuous queries containing window joins on multiple data streams and propose a robust approach that is effective with the different semantic accuracy criteria that are required in different applications. In fact, our approach can be used to (i) maximize the number of output tuples produced by joins, and (ii) optimize the accuracy of complex aggregates estimates under uniform random sampling. We first consider the problem of computing maximal subsets of approximate window joins over multiple data streams. Previously proposed approaches are based on multiple pair-wise joins and, in their load-shedding decisions, disregard the content of streams outside the joined pairs. To overcome these limitations, we optimize our load-shedding policy using various predictors of the productivity of each tuple in the window. To minimize processing costs, we use a fast-and-light sketching technique to estimate the productivity of the tuples. We then show that our method can be generalized to produce statistically accurate samples, as needed in, e.g., the computation of averages, quantiles, and stream mining queries. Tests performed on both synthetic and real-life data demonstrate that our method outperforms previous approaches, while requiring comparable amounts of time and space.