Accelerating the dynamic programming for the optimal polygon triangulation on the GPU

Authors:
Kazufumi Nishida;Koji Nakano;Yasuaki Ito
Affiliations:
Department of Information Engineering, Hiroshima University, Higashi Hiroshima, Japan;Department of Information Engineering, Hiroshima University, Higashi Hiroshima, Japan;Department of Information Engineering, Hiroshima University, Higashi Hiroshima, Japan
Venue:
ICA3PP'12 Proceedings of the 12th international conference on Algorithms and Architectures for Parallel Processing - Volume Part I
Year:
2012

Citing 8
Cited 0

Introduction to algorithms

Introduction to algorithms
A Survey of Longest Common Subsequence Algorithms

SPIRE '00 Proceedings of the Seventh International Symposium on String Processing Information Retrieval (SPIRE'00)
Efficient Canny Edge Detection Using a GPU

ICNC '10 Proceedings of the 2010 First International Conference on Networking and Computing
GPU Computing Gems Emerald Edition

GPU Computing Gems Emerald Edition
A GPU Implementation of Computing Euclidean Distance Map with Efficient Memory Access

ICNC '11 Proceedings of the 2011 Second International Conference on Networking and Computing
Fast and Accurate Template Matching Using Pixel Rearrangement on the GPU

ICNC '11 Proceedings of the 2011 Second International Conference on Networking and Computing
Fast Ellipse Detection Algorithm Using Hough Transform on the GPU

ICNC '11 Proceedings of the 2011 Second International Conference on Networking and Computing
Accelerating the Dynamic Programming for the Matrix Chain Product on the GPU

ICNC '11 Proceedings of the 2011 Second International Conference on Networking and Computing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Modern GPUs (Graphics Processing Units) can be used for general purpose parallel computation. Users can develop parallel programs running on GPUs using programming architecture called CUDA (Compute Unified Device Architecture). The optimal polygon triangulation problem for a convex polygon is an optimization problem to find a triangulation with minimum total weight. It is known that this problem can be solved using the dynamic programming technique in O(n3) time using a work space of size O(n2). The main contribution of this paper is to present an efficient parallel implementation of this O(n3)-time algorithm on the GPU. In our implementation, we have used two new ideas to accelerate the dynamic programming. The first idea (granularity adjustment) is to partition the dynamic programming algorithm into many sequential kernel calls of CUDA, and to select the best size and number of blocks and threads for each kernel call. The second idea (sliding and mirroring arrangements) is to arrange the temporary data for coalesced access of the global memory in the GPU to minimize the memory access overhead. Our implementation using these two ideas solves the optimal polygon triangulation problem for a convex 16384-gon in 69.1 seconds on the NVIDIA GeForce GTX 580, while a conventional CPU implementation runs in 17105.5 seconds. Thus, our GPU implementation attains a speedup factor of 247.5.