Communication Optimizations
The communication cost of a parallel application can greatly affect its scalability. Although communication bandwidth increases have kept pace with increases in processor speed over the past decade, the communication latencies (including the software overhead) for each message have not decreased proportionately.
The framework currently has three major motivations
- Optimize collective communication operations like AlltoAll personalized communication, AlltoAll multicast, AllReduce etc. Collective communication operations often involve most processors in a system. They are also time consuming and can involve massive data movement. These operations can be optimized by using message combining for small messages and smart message sequencing for large messages. Message combining is achieved by imposing a virtual topology on the processors and routing messages along that topology. Messages destined to a group of processors are agglomerated into a single message. The combined message is then sent to a representative processor which forwards the message to the correct destination. For example if the virtual topology is Hypercube, dimensional exchange can be used to combine messages. There will be log(p) stages and in stage i each processor will exchange messages with its ith dimension neighbor. We have also implemented two other virtual topologies 2D Mesh and 3D Grid. For large messages smart message sequencing like prefix send can be used to reduce network contention.
- Optimize implementations of the Charm++ machine layers to exploit the special features provided by the lower lever API's.
- Develop a learning framework which will learn the communication patterns of an application and use known strategies to optimize those patterns.
Papers / Talks
-
22-102022
Phd Thesis -
22-082022
PaperImproving Communication Asynchrony and Concurrency for Adaptive MPI Endpoints
- Sam White
- Laxmikant Vasudeo Kale
-
22-052022
PaperImproving Scalability with GPU-Aware Asynchronous Tasks
- Jaemin Choi
- David F. Richards
- Laxmikant Vasudeo Kale
-
22-032022
PaperOptimizing Non-Commutative Allreduce over Virtualized, Migratable MPI Ranks
- Sam White
- Laxmikant Vasudeo Kale
-
18-012018
Talk -
17-102017
PaperOptimizing Point-to-Point Communication between Adaptive MPI Endpoints in Shared Memory
- Sam White
- Laxmikant Vasudeo Kale
-
15-042015
Phd Thesis -
14-182014
PaperTRAM: Optimizing Fine-grained Communication with Topological Routing and Aggregation of Messages
- Lukasz Wesolowski
- Ramprasad Venkataraman
- Abhishek Gupta
- Jae-Seung Yeom
- Keith Bisset
- Yanhua Sun
- Pritish Jetley
- Thomas Quinn
- Laxmikant Vasudeo Kale
-
14-122014
PaperPICS: A Performance-Analysis-Based Introspective Control System to Steer Parallel Applications
- Yanhua Sun
- Jonathan Lifflander
- Laxmikant Vasudeo Kale
-
11-042011
PaperEvaluation of Simple Causal Message Logging for Large-Scale Fault Tolerant HPC Systems
- Esteban Meneses
- Greg Bronevetsky
- Laxmikant Vasudeo Kale
-
08-112008
PaperCkDirect: Unsynchronized One-Sided Communication in a Message-Driven Paradigm
- Eric Bohm
- Sayantan Chakravorty
- Pritish Jetley
- Abhinav Bhatele
- Laxmikant Vasudeo Kale
-
06-102006
MS Thesis -
05-022005
PaperArchitecture for supporting Hardware Collectives in Output-Queued High-Radix Routers
- Sameer Kumar
- Laxmikant Vasudeo Kale
- Craig Stunkel
-
03-152003
PaperOpportunities and Challenges of Modern Communication Architectures:Case Study with QsNet
- Sameer Kumar
- Laxmikant Vasudeo Kale
-
03-112003
PaperScaling Collective Multicast on Fat-tree Networks
- Sameer Kumar
- Laxmikant Vasudeo Kale
-
03-042003
PaperScaling Collective Multicast on High Performance Clusters
- Laxmikant Vasudeo Kale
- Sameer Kumar
-
02-102002
PaperA Framework for Collective Personalized Communication
- Laxmikant Vasudeo Kale
- Sameer Kumar
- Krishnan Varadarajan
-
94-071994
MS Thesis -
92-101992
PaperDynamic Adaptive Scheduling in an Implementation of a Data Parallel Language
- Ed Kornkven
- Laxmikant Vasudeo Kale