Power and energy efficiency are important challenges for the High Performance Computing (HPC) community. Excessive power consumption is a main limitation for further scaling of HPC systems, and researchers believe that current technology trends will not provide Exascale performance within a reasonable power budget in near future. Hardware innovations such as the proposed Exascale architectures and Near Threshold Computing are expected to improve power efficiency significantly, but more innovations are required in this domain to make Exascale possible.
To help shrink the power efficiency gap, we argue that adaptive runtime systems can be exploited. The runtime system (RTS) can save significant power, since it is aware of both the hardware properties and the application behavior.
Adaptive Runtime Systems. We use application-centric analysis of different architectures to design automatic adaptive RTS techniques that save significant power in different system components, only with minor hardware support. In a nutshell, we analyze different modern architectures and common applications and illustrate that some system components such as caches and network links consume extensive power disproportionately for common HPC applications. We demonstrate how a large fraction of power consumed in caches and networks can be saved using our approach automatically. In these cases, the hardware support the RTS needs is the ability to turn off ways of set-associative caches and network links.
Optimizing Energy Consumption. Furthermore, cooling energy needs to be considered for large-scale systems. As of today, most of the research has focused on saving machine energy consumption leaving behind energy spent on cooling which takes about 40% of the total energy consumption for a datacenter. Our focus is to extend energy optimization work beyond machine energy saving so that we reduce cooling energy. Most datacenters do excessive cooling in order to avoid hotspots (areas in the machine room which are at a much higher temperature than other parts of the room). We are working on a runtime system which uses Dynamic Voltage and Frequency Scaling (DVFS) in order to minimize the occurrence of hotspots by keeping core temperatures in check. While doing so, one of our schemes reduces the timing penalty associated with using just DVFS by doing chare migration in order to load balance the application. Our results show that we can save considerable cooling energy using this temperature aware load balancing. Part of our recent research is exploring the possibility of load balancing chares in a way that we place 'less-frequency-sensitive' chares on hotter cores so that we can further reduce DVFS induced slowdown.
Performance Optimization Under Power Budget. Recent advances in processor and memory hardware designs have made it possible for the user to control the power consumption of the CPU and memory through software, e.g., the power consumption of Intel’s Sandy Bridge family of processors can be user-controlled through the Running Average Power Limit (RAPL) library. It has been shown that increase in the power allowed to the processor (and/or memory) does not yield a proportional increase in the application’s performance. As a result, for a given power budget, it can be better to run an application on larger number of nodes with each node capped at lower power than fewer nodes each running at its TDP. This is also called as overprovisioning. The optimal resource configuration for an application can be determined by profiling an application’s performance for varying number of nodes, CPU power and memory power and then selecting the best performing configuration for the given power budget. In our recent work, we propose a performance modeling scheme that estimates the essential power characteristics of a job at any scale. Our online resource manager uses these performance characteristics for making scheduling and resource allocation decisions that maximize the job throughput of the supercomputer under a given power budget. With a power budget of 4.75 MW, we can obtain up to 5.2X improvement in job throughput when compared with the SLURM scheduling policy that is power-unaware. With real experiments on a relatively small scale cluster, we obtained 1.7X improvement. An adaptive runtime system allows further improvement by allowing already running jobs to shrink and expand for optimal resource allocation.
Several of our new online softwares and methods such as Power Aware Resource Manager [14-15] and Variation Aware Scheduler [15-01] use linear/integer programming to come up with superior solutions as compared to solutions obtained from suboptimal heuristics.
Papers / Talks
-
19-052019
PaperFine-Grained Energy Efficiency Using Per-CoreDVFS with an Adaptive Runtime System
- Bilge Acun
- Kavitha Chandrasekar
- Laxmikant Vasudeo Kale
-
16-132016
PaperNeural Network-Based Task Scheduling with Preemptive Fan Control
- Bilge Acun
- Eun Kyung Lee
- Yoonho Park
- Laxmikant Vasudeo Kale
-
16-122016
PaperPower, Reliability, Performance: One System to Rule Them All
- Bilge Acun
- Akhil Langer
- Esteban Meneses
- Harshitha Menon
- Osman Sarood
- Ehsan Totoni
- Laxmikant Vasudeo Kale
-
16-102016
PaperEnergy-optimal Configuration Selection for Manycore Chips with Variation
- Akhil Langer
- Ehsan Totoni
- Udatta S Palekar
- Laxmikant Vasudeo Kale
-
16-082016
PaperVariation Among Processors Under Turbo Boost in HPC Systems
- Bilge Acun
- Phil Miller
- Laxmikant Vasudeo Kale
-
16-032016
PaperMitigating Processor Variation through Dynamic Load Balancing
- Bilge Acun
- Laxmikant Vasudeo Kale
-
15-112015
PaperAnalyzing Energy-Time Tradeoff in Power Overprovisioned HPC Data Centers
- Akhil Langer
- Harshit Dokania
- Laxmikant Vasudeo Kale
- Udatta S Palekar
-
15-052015
Phd Thesis -
15-012015
PaperEnergy-efficient Computing for HPC Workloads on Heterogeneous Manycore Chips
- Akhil Langer
- Ehsan Totoni
- Udatta S Palekar
- Laxmikant Vasudeo Kale
-
14-352014
PaperScheduling for HPC Systems with Process Variation Heterogeneity
- Ehsan Totoni
- Akhil Langer
- Josep Torrellas
- Laxmikant Vasudeo Kale
-
14-272014
PaperPower Management of Extreme-scale Networks with On/Off Links in Runtime Systems
- Ehsan Totoni
- Nikhil Jain
- Laxmikant Vasudeo Kale
-
14-232014
PaperUsing an Adaptive HPC Runtime System to Reconfigure the Cache Hierarchy
- Ehsan Totoni
- Josep Torrellas
- Laxmikant Vasudeo Kale
-
14-192014
Paper- Laxmikant Vasudeo Kale
- Akhil Langer
- Osman Sarood
-
14-152014
PaperMaximizing Throughput of Overprovisioned HPC Data Centers Under a Strict Power Budget
- Osman Sarood
- Akhil Langer
- Abhishek Gupta
- Laxmikant Vasudeo Kale
-
14-022014
PaperEnergy Profile of Rollback-Recovery Strategies in High Performance Computing
- Esteban Meneses
- Osman Sarood
- Laxmikant Vasudeo Kale
-
13-562013
PaperEasy, Fast and Energy Efficient Object Detection on Heterogeneous On-Chip Architectures
- Ehsan Totoni
- Mert Dikmen
- Maria Garzaran
-
13-502013
TalkA ‘Cool’ Way of Improving the Reliability of HPC Machines
- Osman Sarood
- Esteban Meneses
- Laxmikant Vasudeo Kale
-
13-332013
PaperThermal Aware Automated Load Balancing for HPC Applications
- Harshitha Menon
- Bilge Acun
- Simon Garcia De Gonzalo
- Osman Sarood
- Laxmikant Vasudeo Kale
-
13-252013
PaperA ‘Cool’ Way of Improving the Reliability of HPC Machines
- Osman Sarood
- Esteban Meneses
- Laxmikant Vasudeo Kale
-
13-202013
PaperOptimizing Power Allocation to CPU and Memory Subsystems in Overprovisioned HPC Systems
- Osman Sarood
- Akhil Langer
- Laxmikant Vasudeo Kale
- Barry Rountree
- Bronis de Supinski
-
13-102013
Talk -
13-092013
PaperToward Runtime Power Management of Exascale Networks by On/Off Control of Links
- Ehsan Totoni
- Nikhil Jain
- Laxmikant Vasudeo Kale
-
12-432012
TalkAssessing Energy Efficiency of Fault Tolerance Protocols for HPC Systems
- Esteban Meneses
- Osman Sarood
- Laxmikant Vasudeo Kale
-
12-372012
PaperAssessing Energy Efficiency of Fault Tolerance Protocols for HPC Systems
- Esteban Meneses
- Osman Sarood
- Laxmikant Vasudeo Kale
-
12-272012
PaperCloud Friendly Load Balancing for HPC Applications: Preliminary Work
- Osman Sarood
- Abhishek Gupta
- Laxmikant Vasudeo Kale
-
12-202012
Paper‘Cool’ Load Balancing for High Performance Computing Data Centers
- Osman Sarood
- Phil Miller
- Ehsan Totoni
- Laxmikant Vasudeo Kale
-
12-102012
Talk -
11-102011
PaperTemperature Aware Load Balancing for Parallel Applications: Preliminary Work
- Osman Sarood
- Abhishek Gupta
- Laxmikant Vasudeo Kale
-
06-152006
PaperParallel Adaptive Simulations of Dynamic Fracture Events
- Sandhya Mangala
- Terry Wilmarth
- Sayantan Chakravorty
- Nilesh Choudhury
- Laxmikant Vasudeo Kale
- Philippe Geubelle
-
05-082005
PaperAn Integration Framework for Simulations of Solid Rocket Motors
- Xiangmin Jiao
- Gengbin Zheng
- Orion Lawlor
- Phil Alexander
- Mike Campbell
- Michael Heath
- Robert Fiedler
-
96-101996
PaperStructured Dagger: A Coordination Language for Message-Driven Programming
- Laxmikant Vasudeo Kale
- Milind Bhandarkar
-
93-131993
PaperA Load Balancing Strategy For Prioritized Execution of Tasks
- Amitabh Sinha
- Laxmikant Vasudeo Kale