Towards Identify Anonymization in Large Survey Rating Data

Authors:
Xiaoxun Sun;Hua Wang
Affiliations:
-;-
Venue:
NSS '10 Proceedings of the 2010 Fourth International Conference on Network and System Security
Year:
2010

Citing 0
Cited 1

Positive influence dominating set in e-learning social networks

ICWL'11 Proceedings of the 10th international conference on Advances in Web-Based Learning

Quantified Score

Hi-index	0.00

Visualization

Abstract

We study the challenge of identity protection in the large public survey rating data. Even though the survey participants do not reveal any of their ratings, their survey records are potentially identifiable by using information from other public sources. None of the existing anonymisation principles (e.g., $k$-anonymity, $l$-diversity, etc.) can effectively prevent such breaches in large survey rating data sets. In this paper, we tackle the problem by defining the $ (k, \epsilon)$-anonymity principle. The principle requires for each transaction $t$ in the given survey rating data $T$, at least $ (k-1)$ other transactions in $T$ must have ratings similar with $t$, where the similarity is controlled by $\epsilon$. We propose a greedy approach to anonymize survey rating data and apply the method to two real-life data sets to demonstrate their efficiency and practical utility.