Massively Categorical Variables: Revealing the Information in Zip Codes

Authors:
Thomas J. Steenburgh;Andrew Ainslie;Peder Hans Engebretson
Affiliations:
-;-;-
Venue:
Marketing Science
Year:
2003

Citing 0
Cited 8

Editorial: The Impact of Advancing Technology on Marketing and Academic Research

Marketing Science
Modeling Movie Life Cycles and Market Share

Marketing Science
Understanding Geographical Markets of Online Firms Using Spatial Models of Customer Choice

Marketing Science
Real-Time Evaluation of E-mail Campaign Performance

Marketing Science
Incorporating neighborhood effects in customer relationship management models

ISMIS'11 Proceedings of the 19th international conference on Foundations of intelligent systems
Traditional and IS-Enabled Customer Acquisition on the Internet

Management Science
Including spatial interdependence in customer acquisition models: A cross-category comparison

Expert Systems with Applications: An International Journal
Customer event history for churn prediction: How long is long enough?

Expert Systems with Applications: An International Journal

Quantified Score

Hi-index	0.00

Visualization

Abstract

We introduce the idea of a massively categorical variable, a variable such as zip code that takes on too many values to treat in the standard manner. We show how to use a massively categorical variable directly as an explanatory variable.As an application of this concept, we explore several of the issues that analysts confront when trying to develop a direct marketing campaign. We begin by pointing out that the data contained in many of the common sources are masked through aggregation in order to protect consumer privacy. This creates some difficulty when trying to construct models of individual level behavior.We show how to take full advantage of such data through a hierarchical Bayesian variance components (HBVC) model. The flexibility of our approach allows us to combine several sources of information, some of which may not be aggregated, in a coherent manner. We show that the conventional modeling practice understates the uncertainty with regard to its parameter values.We explore an array of financial considerations, including ones in which the marginal benefit is non-linear, to make robust model comparisons. To implement the decision rules that determine the optimal number of prospects to contact, we develop an algorithm based on the Monte Carlo Markov chain output from parameter estimation. We conclude the analysis by demonstrating how to determine an organization's willingness to pay for additional data.