On the Selection of an Optimal Set of Indexes

Authors:
M. Y. L. Ip;L. V. Saxton;V. V. Raghavan
Affiliations:
Datatron Processing & Systems Ltd.;-;-
Venue:
IEEE Transactions on Software Engineering
Year:
1983

Citing 0
Cited 6

Physical Database Design: the database professional's guide to exploiting indexes, views, storage, and more

Physical Database Design: the database professional's guide to exploiting indexes, views, storage, and more
Compressing Very Large Database Workloads for Continuous Online Index Selection

DEXA '08 Proceedings of the 19th international conference on Database and Expert Systems Applications
An index selection method without repeated optimizer estimations

Information Sciences: an International Journal
Data mining-based materialized view and index selection in data warehouses

Journal of Intelligent Information Systems
A genetic algorithm for the index selection problem

EvoWorkshops'03 Proceedings of the 2003 international conference on Applications of evolutionary computing
Intrusion recovery for database-backed web applications

SOSP '11 Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles

Quantified Score

Hi-index	0.00

Visualization

Abstract

A problem of considerable interest in the design of a database is the selection of indexes. In this paper, we present a probabilistic model of transactions (queries, updates, insertions, and deletions) to a file. An evaluation function, which is based on the cost saving (in terms of the number of page accesses) attributable to the use of an index set, is then developed. The maximization of this function would yield an optimal set of indexes. Unfortunately, algorithms known to solve this maximization problem require an order of time exponential in the total number of attributes in the file. Consequently, we develop the theoretical basis which leads to an algorithm that obtains a near optimal solution to the index selection problem in polynomial time. The theoretical result consists of showing that the index selection problem can be solved by solving a properly chosen instance of the knapsack problem. A theoretical bound for the amount by which the solution obtained by this algorithm deviates from the true optimum is provided. This result is then interpreted in the light of evidence gathered through experiments.