End-to-end Performance Modeling of Distributed GPU Applications

International Conference on Supercomputing (ICS) 2020
Pulication Type: Paper
Download: pdf ps

Abstract

With the growing number of GPU-based supercomputing platforms and GPU-enabled applications, the ability to accurately model the performance of such applications is becoming increasingly important. Most current performance models for GPU-enabled applications are limited to single node performance. In this work, we propose a methodology for end-to-end performance modeling of distributed GPU applications. Our work strives to create performance models that are both accurate and easily applicable to any distributed GPU application. We combine trace-driven simulation of MPI communication based on the TraceR-CODES framework with a profiling-based roofline model for GPU kernels. We make substantial modifications to these models to capture the complex effects of both on-node and off-node networks in today's multi-GPU supercomputers. We validate our model against empirical data from GPU platforms and also vary tunable parameters of our model to observe how they affect application performance.

Text Ref


						

BibTex

@inproceedings{10.1145/3392717.3392737,
author = {Choi, Jaemin and Richards, David F. and Kale, Laxmikant V. and Bhatele, Abhinav},
title = {End-to-End Performance Modeling of Distributed GPU Applications},
year = {2020},
isbn = {9781450379830},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3392717.3392737},
doi = {10.1145/3392717.3392737},
booktitle = {Proceedings of the 34th ACM International Conference on Supercomputing},
articleno = {30},
numpages = {12},
keywords = {GPU computing, performance modeling, communication, trace-driven simulation},
location = {Barcelona, Spain},
series = {ICS ’20}
}