In-place transposition of rectangular matrices on accelerators

Authors:
I-Jui Sung;Juan Gómez-Luna;José María González-Linares;Nicolás Guil;Wen-Mei W. Hwu
Affiliations:
MulticoreWare, Inc, Champaign, IL, USA;University of Córdoba, Córdoba, Spain;University of Málaga, Málaga, Spain;University of Málaga, Málaga, Spain;University of Illinois at Urbana-Champaign, Urbana, IL, USA
Venue:
Proceedings of the 19th ACM SIGPLAN symposium on Principles and practice of parallel programming
Year:
2014

Citing 15
Cited 1

Efficient transposition algorithms for large matrices

Proceedings of the 1993 ACM/IEEE conference on Supercomputing
A Method for Transposing a Matrix

Journal of the ACM (JACM)
Algorithm 513: Analysis of In-Situ Transposition [F1]

ACM Transactions on Mathematical Software (TOMS)
Anatomy of high-performance matrix multiplication

ACM Transactions on Mathematical Software (TOMS)
Benchmarking GPUs to tune dense linear algebra

Proceedings of the 2008 ACM/IEEE conference on Supercomputing
Auto-tuning of fast fourier transform on graphics processors

Proceedings of the 16th ACM symposium on Principles and practice of parallel programming
Dymaxion: optimizing memory access patterns for heterogeneous systems

Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis
Parallel and Cache-Efficient In-Place Matrix Storage Format Conversion

ACM Transactions on Mathematical Software (TOMS)
Performance models for asynchronous data transfers on consumer Graphics Processing Units

Journal of Parallel and Distributed Computing
GPU-vote: a framework for accelerating voting algorithms on GPU

Euro-Par'12 Proceedings of the 18th international conference on Parallel Processing
K-Means for Parallel Architectures Using All-Prefix-Sum Sorting and Updating Steps

IEEE Transactions on Parallel and Distributed Systems
An optimized approach to histogram computation on GPU

Machine Vision and Applications
Improving GPU Performance Prediction with Data Transfer Modeling

IPDPSW '13 Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing Workshops and PhD Forum
Performance Modeling of Atomic Additions on GPU Scratchpad Memory

IEEE Transactions on Parallel and Distributed Systems
A decomposition for in-place matrix transposition

Proceedings of the 19th ACM SIGPLAN symposium on Principles and practice of parallel programming

A decomposition for in-place matrix transposition

Proceedings of the 19th ACM SIGPLAN symposium on Principles and practice of parallel programming

Quantified Score

Hi-index	0.00

Visualization

Abstract

Matrix transposition is an important algorithmic building block for many numeric algorithms such as FFT. It has also been used to convert the storage layout of arrays. With more and more algebra libraries offloaded to GPUs, a high performance in-place transposition becomes necessary. Intuitively, in-place transposition should be a good fit for GPU architectures due to limited available on-board memory capacity and high throughput. However, direct application of CPU in-place transposition algorithms lacks the amount of parallelism and locality required by GPUs to achieve good performance. In this paper we present the first known in-place matrix transposition approach for the GPUs. Our implementation is based on a novel 3-stage transposition algorithm where each stage is performed using an elementary tiled-wise transposition. Additionally, when transposition is done as part of the memory transfer between GPU and host, our staged approach allows hiding transposition overhead by overlap with PCIe transfer. We show that the 3-stage algorithm allows larger tiles and achieves 3X speedup over a traditional 4-stage algorithm, with both algorithms based on our high-performance elementary transpositions on the GPU. We also show our proposed low-level optimizations improve the sustained throughput to more than 20 GB/s. Finally, we propose an asynchronous execution scheme that allows CPU threads to delegate in-place matrix transposition to GPU, achieving a throughput of more than 3.4 GB/s (including data transfers costs), and improving current multithreaded implementations of in-place transposition on CPU.