Vectorizing Unstructured Mesh Computations for Many-core Architectures

Authors:
I. Z. Reguly;E. László;G. R. Mudalige;M. B. Giles
Affiliations:
Oxford e-Research Centre, University of Oxford and Faculty of Information Technology and Bionics, Pázmány Péter Catholic University;Oxford e-Research Centre, University of Oxford and Faculty of Information Technology and Bionics, Pázmány Péter Catholic University;Oxford e-Research Centre, University of Oxford;Oxford e-Research Centre, University of Oxford
Venue:
Proceedings of Programming Models and Applications on Multicores and Manycores
Year:
2014

Citing 9
Cited 0

Multicolor ICCG methods for vector computers

SIAM Journal on Numerical Analysis
Page placement policies for NUMA multiprocessors

Journal of Parallel and Distributed Computing
Beyond Traditional Microprocessors for Geoscience High-Performance Computing Applications

IEEE Micro
Performance Analysis and Optimization of the OP2 Framework on Many-Core Architectures

The Computer Journal
Analysis and Optimization of Financial Analytics Benchmark on Modern Multi- and Many-core IA-Based Architectures

SCC '12 Proceedings of the 2012 SC Companion: High Performance Computing, Networking Storage and Analysis
Design and Implementation of the Linpack Benchmark for Single and Multi-node Systems Based on Intel® Xeon Phi Coprocessor

IPDPS '13 Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing
Exploring SIMD for Molecular Dynamics, Using Intel® Xeon® Processors and Intel® Xeon Phi Coprocessors

IPDPS '13 Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing
Designing OP2 for GPU architectures

Journal of Parallel and Distributed Computing
Design and initial performance of a high-level unstructured mesh framework on heterogeneous parallel systems

Parallel Computing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Achieving optimal performance on the latest multi-core and many-core architectures depends more and more on making efficient use of the hardware's vector processing capabilities. While auto-vectorizing compilers do not require the use of vector processing constructs, they are only effective on a few classes of applications with regular memory access and computational patterns. Irregular application classes require the explicit use of parallel programming models; CUDA and OpenCL are well established for programming GPUs, but it is not obvious what model to use to exploit vector units on architectures such as CPUs or the Xeon Phi. Therefore it is of growing interest what programming models are available, such as Single Instruction Multiple Threads (SIMT) or Single Instruction Multiple Data (SIMD), and how they map to vector units. This paper presents results on achieving high performance through vectorization on CPUs and the Xeon Phi on a key class of applications: unstructured mesh computations. By exploring the SIMT and SIMD execution and parallel programming models, we show how abstract unstructured grid computations map to OpenCL or vector intrinsics through the use of code generation techniques, and how these in turn utilize the hardware. We benchmark a number of systems, including Intel Xeon CPUs and the Intel Xeon Phi, using an industrially representative CFD application and compare the results against previous work on CPUs and NVIDIA GPUs to provide a contrasting comparison of what could be achieved on current many-core systems. By carrying out a performance analysis study, we identify key performance bottlenecks due to computational, control and bandwidth limitations. We show that the OpenCL SIMT model does not map efficiently to CPU vector units due to auto-vectorization issues and threading overheads. We demonstrate that while the use of SIMD vector intrinsics imposes some restrictions, and requires more involved programming techniques, it does result in efficient code and near-optimal performance, that is up to 2 times faster than the non-vectorized code. We observe that the Xeon Phi does not provide good performance for this class of applications, but is still on par with a pair of high-end Xeon chips. CPUs and GPUs do saturate the available resources, giving performance very near to the optimum.