VLDB 2026 Research / reviewers in the wild / expert
Di Zhu 0002
dblp:21/2144-2
· DBLP profile ↗
18ranked-venue papers
9as first author
0since 2021 · last 2019
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 9 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Interconnection networks and networks-on-chip · 32% Energy-efficient computing · 30% Processor architecture and microarchitecture · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Energy systems and smart grids · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy systems and smart grids
energy storage |
0.2 | 1 | 2016 | Toward a Profitable Grid-Connected Hybrid Electrical Energy Storage System for Residential Use · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Energy systems and smart grids › energy storage
hybrid energy storage system |
0.2 | 1 | 2016 | Toward a Profitable Grid-Connected Hybrid Electrical Energy Storage System for Residential Use · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.2 | 1 | 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip Multiprocessors · IEEE Trans. Computers 2016 |
Processor architecture and microarchitecture
many-core architecture |
0.2 | 1 | 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip Multiprocessors · IEEE Trans. Computers 2016 |
Interconnection networks and networks-on-chip › router architecture
network-on-chip router |
0.2 | 1 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 |
Energy-efficient computing
power gating |
0.2 | 1 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 |
Energy-efficient computing
power management |
0.2 | 1 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 |
Energy systems and smart grids
demand-side management |
0.1 | 1 | 2016 | Toward a Profitable Grid-Connected Hybrid Electrical Energy Storage System for Residential Use · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip Multiprocessors · IEEE Trans. Computers 2016 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2sensitivity analysis · 0.2heuristic algorithm · 0.2full-system simulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Express Link Placement for NoC-Based Many-Core PlatformsabstractWith the integration of up to hundreds of cores in recent general-purpose processors that can be used in parallel processing systems, it is critical to design scalable and low-latency networks-on-chip (NoCs) to support various on-chip communications. An effective way to reduce on-chip latency and improve network scalability is to add express links between pairs of non-adjacent routers. However, increasing the number of express links may result in smaller bandwidth per link due to the limited total bisection bandwidth on chip, thus leading to higher serialization latency of packets in the network. Unlike previous works on application-specific designs or on fixed placement of express links, this paper aims at finding effective placement of express links for general-purpose processors considering all the possible placement options. We formulate the problem mathematically and propose an efficient algorithm that utilizes an initial solution generation heuristic and enhanced candidate generator in simulated annealing. Evaluation on 4x4, 8x8 and 16x16 networks using multi-threaded PARSEC benchmarks and various synthetic traffic patterns shows significant reduction of average packet latency over previous works. Yunfan Li 0002, Di Zhu 0002, Lizhong Chen |
ICPP | 2 |
| 2019 | On Trade-off Between Static and Dynamic Power Consumption in NoC Power GatingabstractRecent research has proposed to minimize network-on-chip (NoC) static power by proactively power-gating selected routers when not all the cores are active. However, as more routers are powered off, on-chip packets are forced to take detours more frequently, resulting in a higher hop count and increased dynamic power that may potentially offset the static power savings. This paper investigates such a trade-off between static and dynamic power in detail, and explores how overall NoC power consumption can be minimized through proactive power-gating. Three efficient and effective algorithms are proposed to reduce NoC static power, dynamic power, and overall power consumption, respectively. Evaluation results based on PARSEC benchmarks demonstrate the importance of the trade-off and show a substantial improvement in total NoC power savings of the proposed algorithms, compared with previous work that did not give full consideration to both static and dynamic power. Di Zhu 0002, Yunfan Li 0002, Lizhong Chen |
ISLPED | 1 |
| 2017 | CALM: Contention-Aware Latency-Minimal Application Mapping for Flattened Butterfly On-Chip NetworksabstractWith the emergence of many-core multiprocessor system-on-chips (MPSoCs), on-chip networks are facing serious challenges in providing fast communication among various tasks and cores. One promising on-chip network design approach shown in recent studies is to add express channels to traditional mesh network as shortcuts to bypass intermediate routers, thereby reducing packet latency. This approach not only changes the packet latency models, but also greatly affects network traffic behaviors, both of which have not been fully exploited in existing mapping algorithms. In this article, we explore the opportunities in optimizing application mapping for flattened butterfly, a popular express channel-based on-chip network. Specifically, we identify the unique characteristics of flattened butterfly, analyze the opportunities and new challenges, and propose an efficient heuristic mapping algorithm. The proposed algorithm Contention-Aware Latency Minimal (CALM) is able to reduce unnecessary turns that would otherwise impose additional router pipeline latency to packets, as well as adjust forwarding traffic to reduce network contention latency. Simulation results show that the proposed algorithm can achieve, on average, 3.4X reduction in the number of turns, 24.8% reduction in contention latency, and 14.12% reduction in the overall packet latency. Di Zhu 0002, Siyu Yue, Massoud Pedram, Lizhong Chen |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | Simulation of NoC power-gating: Requirements, optimizations, and the Agate simulator
Lizhong Chen, Di Zhu 0002, Massoud Pedram, Timothy M. Pinkston |
J. Parallel Distributed Comput. | 2 |
| 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip MultiprocessorsabstractThis paper addresses the problem of balancing the on-chip packet latencies in a chip multi-processor (CMP), which is simultaneously executing multiple applications. Specifically, this paper presents a balanced application-to-core mapping algorithm that aims to minimize the maximum on-chip packet latency of all running applications. The paper starts by formulating the balanced mapping problem for CMPs and proving its NP-completeness. Next it presents an efficient heuristic algorithm for solving the aforesaid problem, which utilizes the characteristics of on-chip cache and memory accesses in CMPs and takes into account the workload variations among applications. Simulation results on PARSEC benchmark suite show that the proposed algorithm lowers the maximum average packet latency of all applications by 11 percent while cutting the standard deviation of on-chip packet latencies by 99 percent. This is achieved by very little overhead in terms of the overall packet latency and power consumption averaged over all packets. Di Zhu 0002, Lizhong Chen, Siyu Yue, Timothy M. Pinkston, Massoud Pedram |
IEEE Trans. Computers | 1 |
| 2016 | Toward a Profitable Grid-Connected Hybrid Electrical Energy Storage System for Residential UseabstractHybrid electrical energy storage (HEES) systems have the potential to result in considerable cost savings by reducing the electric bills of home users. This paper first presents grid-connected dual-bank HEES system design and management to maximize the electric bill savings for residential users, and subsequently provides a comprehensive sensitivity analysis of the economic feasibility of residential HEES systems. Specifically, the paper describes a daily management policy based on energy buffering strategy with one bank as the main storage bank and the other as the energy buffering bank, and then derive the global design of HEES specifications based on the daily management results. Simulation results prove the effectiveness of energy buffering strategy and show the proposed HEES system is capable of bringing in profits under current input variables. Finally, a detailed analysis is conducted to show how each input variable affects the final design of the proposed residential HEES system and the maximum annual profits it achieves. Together with the design and control mechanism, the proposed analysis provides potential customers with the comprehensive knowledge of how HEES systems can be deployed to achieve savings in their electric bills. Di Zhu 0002, Siyu Yue, Naehyuck Chang, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | TAPP: temperature-aware application mapping for NoC-based many-core processors
Di Zhu 0002, Lizhong Chen, Timothy M. Pinkston, Massoud Pedram |
DATE | 1 |
| 2015 | Power punch: Towards non-blocking power-gating of NoC routersabstractAs chip designs penetrate further into the dark silicon era, innovative techniques are much needed to power off idle or under-utilized system components while having minimal impact on performance. On-chip network routers are potentially good targets for power-gating, but packets in the network can be significantly delayed as their paths may be blocked by powered-off routers. In this paper, we propose Power Punch, a novel performance-aware, power reduction scheme that aims to achieve non-blocking power-gating of on-chip network routers. Two mechanisms are proposed that not only allow power control signals to utilize existing slack at source nodes to wake up powered-off routers along the first few hops before packets are injected, but also allow these signals to utilize hop count slack by staying ahead of packets to "punch through " any blocked routers along the imminent path of packets, preventing packets from having to suffer router wakeup latency or packet detour latency. Full system evaluation on PARSEC benchmarks shows Power Punch saves more than 83% of router static energy while having an execution time penalty of less than 0.4%, effectively achieving near non-blocking power-gating of on-chip network routers. Lizhong Chen, Di Zhu 0002, Massoud Pedram, Timothy M. Pinkston |
HPCA | 2 |
| 2014 | Application mapping for express channel-based networks-on-chipabstractWith the emergence of many-core multiprocessor system-on-chips (MPSoCs), the on-chip networks are facing serious challenges in providing fast communication for various tasks and cores. One promising solution shown in recent studies is to add express channels to the network as shortcuts to bypass intermediate routers, thereby reducing packet latency. However, this approach also greatly changes the packet delay estimation and traffic behaviors of the network, both of which have not yet been exploited in existing mapping algorithms. In this paper, we explore the opportunities in optimizing application mapping for express channel-based on-chip networks. Specifically, we derive a new delay model for this type of networks, identify their unique characteristics, and propose an efficient heuristic mapping algorithm that increases the bypassing opportunities by reducing unnecessary turns that would otherwise impose the entire router pipeline delay to packets. Simulation results show that the proposed algorithm can achieve a 2∼4X reduction in the number of turns and 10∼26% reduction in the average packet delay. Di Zhu 0002, Lizhong Chen, Siyu Yue, Massoud Pedram |
DATE | 1 |
| 2014 | Optimal design and management of a smart residential PV and energy storage systemabstractSolar photovoltaic (PV) technology has been widely deployed in large power plants operated by utility companies. However, the home owners are not yet convinced of the saving cost benefits of this technology, and consequently, in spite of government subsidies, they have been reluctant to install PV systems in their homes. The main reason for this is the absence of a complete and truthful analysis which could explain to home owners under what conditions spending money on a PV system can actually save them money over a long-term, but known, time horizon. This paper thus presents a design and management mechanism for a smart residential energy system comprising PV modules, electrical energy storage banks, and conversion circuits connected to the power grid. First, we figure out how much savings can be achieved by a system with given PV modules and EES bank capacities by optimally solving the daily energy flow control problem of such a system. Based on the daily optimization results, we come up with the optimal system specifications with a fixed budget. Experiments are conducted for various electricity prices and different profiles of PV output power and load demand. Results show that the designed system breaks even in 6 years and in the system lifetime achieves up to 8% annual profit besides paying back the budget. Di Zhu 0002, Yanzhi Wang 0001, Naehyuck Chang, Massoud Pedram |
DATE | 1 |
| 2014 | Model-free learning-based online management of hybrid electrical energy storage systems in electric vehiclesabstractTo improve the cycle efficiency and peak output power density of energy storage systems in electric vehicles (EVs), supercapacitors have been proposed as auxiliary energy storage elements to complement the mainstream Lithium-ion (Li-ion) batteries. The performance of such a hybrid electrical energy storage (HEES) system is highly dependent on the implemented management policy. This paper presents a model-free reinforcement learning-based approach to dynamically manage the current flows from and into the battery and supercapacitor banks under various scenarios (combinations of EV specs and driving patterns). Experimental results demonstrate that the proposed approach achieves up to 25% higher efficiency compared to a Li-ion battery only storage system and outperforms other online HEES system control policies in all test cases. Siyu Yue, Yanzhi Wang 0001, Qing Xie 0001, Di Zhu 0002, Massoud Pedram, Naehyuck Chang |
IECON | 4 |
| 2014 | Balancing On-Chip Network Latency in Multi-application Mapping for Chip-MultiprocessorsabstractAs the number of cores continues to grow in chip multiprocessors (CMPs), application-to-core mapping algorithms that leverage the non-uniform on-chip resource access time have been receiving increasing attention. However, existing mapping methods for reducing overall packet latency cannot meet the requirement of balanced on-chip latency when multiple applications are present. In this paper, we address the looming issue of balancing minimized on-chip packet latency with performance-awareness in the multi-application mapping of CMPs. Specifically, the proposed mapping problem is formulated, its NP-completeness is proven, and an efficient heuristic-based algorithm for solving the problem is presented. Simulation results show that the proposed algorithm is able to reduce the maximum average packet latency by 10.42% and the standard deviation of packet latency by 99.65% among concurrently running applications and, at the same time, incur little degradation in the overall performance. Di Zhu 0002, Lizhong Chen, Siyu Yue, Timothy M. Pinkston, Massoud Pedram |
IPDPS | 1 |
| 2014 | Smart butterfly: reducing static power dissipation of network-on-chip with core-state-awarenessabstractWhile power gating is a promising technique to reduce the static power consumption of network-on-chip (NoC), its effectiveness is often hindered by the requirement of maintaining network connec-tivity and the limited knowledge of traffic behaviors. In this paper, we present Smart Butterfly, a core-state-aware NoC power-gating scheme based on flattened butterfly that utilizes the active/sleep state information of processing cores to improve power-gating effectiveness. Smart Butterfly exploits the rich connectivity of the flattened butterfly topology to allow more on-chip routers to be power-gated when their attached cores are asleep. We present two heuristic algorithms to determine the set of routers to be turned on to maintain connectivity and allow tradeoff between power consumption and average packet latency. Simulation results show an average of 42.85% and 60.48% power reduction of Smart Butterfly over prior art on 4x4 and 8x8 networks, respectively. Siyu Yue, Lizhong Chen, Di Zhu 0002, Timothy M. Pinkston, Massoud Pedram |
ISLPED | 3 |
| 2013 | An efficient scheduling algorithm for multiple charge migration tasks in hybrid electrical energy storage systemsabstractHybrid electrical energy storage (HEES) systems are comprised of multiple banks of heterogeneous electrical energy storage (EES) elements with distinct properties. This paper defines and solves the problem of scheduling multiple charge migration tasks in HEES systems with the objective of minimizing the total energy drawn from the source banks. The solution approach consists of two steps: (i) Finding the best charging current profile and voltage level setting for the Charge Transfer Interconnect (CTI) bus for each charge migration task, and (ii) Merging and scheduling the charge migration tasks. Experimental results demonstrate improvements of up to 32.2% in the charge migration efficiency compared to baseline setups in an example HEES system. Qing Xie 0001, Di Zhu 0002, Yanzhi Wang 0001, Massoud Pedram, Younghyun Kim 0001, Naehyuck Chang |
ASP-DAC | 2 |
| 2013 | Maximizing return on investment of a grid-connected hybrid electrical energy storage systemabstractThis paper is the first to present a comprehensive analysis of the profitability of the hybrid electrical energy storage (HEES) systems while further providing a HEES design and control optimization framework to maximize the total return on investment (ROI). The solution consists of two steps: (i) Derivation of an optimal HEES management policy to maximize the daily energy cost saving and (ii) Optimal design of the HEES system to maximize the amortized annual profit under budget and system volume constraints. We consider a HEES system comprised of lead-acid and Li-ion batteries for a case study. The optimal HEES system achieves an annual ROI of up to 60% higher than a lead-acid battery-only system (Li-ion battery-only) system. Di Zhu 0002, Yanzhi Wang 0001, Siyu Yue, Qing Xie 0001, Massoud Pedram, Naehyuck Chang |
ASP-DAC | 1 |
| 2013 | SIMES: A simulator for hybrid electrical energy storage systemsabstractState-of-the-art electrical energy storage (EES) systems are mainly homogeneous, i.e., they consist of a single type of EES elements. None of the existing EES elements is capable of simultaneously fulfilling all the desired features of an ideal EES system, e.g., high charge/discharge efficiency, high energy density, low cost per unit capacity, long cycle life. A novel technology, i.e., a hybrid EES system that employs heterogeneous EES elements organized in a hierarchy of storage banks and linked by appropriate charge transfer interconnects, has shown great promise in overcoming the aforesaid limitations of conventional EES systems. However, the widespread adoption/deployment of hybrid EES systems is hampered by lack of a hybrid EES system simulator. This paper thus presents SIMES, a powerful and scalable simulator for hybrid EES systems, which provides fast and accurate system simulations, while accounting for key characteristics of various EES elements, power converters, charge transfer interconnect schemes, etc. Experimental results on two different applications (one targeting load shifting for households, the other related to battery rate capacity effect minimization in portable electronic devices) demonstrate the value and usefulness of SIMES for designing energy-aware facilities and products. Siyu Yue, Di Zhu 0002, Yanzhi Wang 0001, Massoud Pedram, Younghyun Kim 0001, Naehyuck Chang |
ISLPED | 2 |
| 2012 | Online fault detection and tolerance for photovoltaic energy harvesting systemsabstractPhotovoltaic energy harvesting systems (PV systems) are subject to PV cell faults, which decrease the efficiency of PV systems and even shorten the PV system lifespan. Manual PV cell fault detection and elimination are expensive and nearly impossible for remote PV systems, e.g., PV systems on satellites. Therefore, online fault detection techniques and fault tolerance solutions are needed that can detect and tolerate PV cell faults without manual intervention. In this work, we present an online fault detection and tolerance technique for remote PV systems, which is capable of dynamically locating faulty PV cells and tolerating PV cell faults. More precisely, we present a modified PV panel structure and an efficient algorithm for our online fault detection and tolerance. Our fault detection and tolerance technique reduces output power degradation due to PV cell faults in a PV system by up to 81.31%. Xue Lin 0001, Yanzhi Wang 0001, Di Zhu 0002, Naehyuck Chang, Massoud Pedram |
ICCAD | 3 |
| 2012 | Reinforcement learning based dynamic power management with a hybrid power supplyabstractDynamic power management (DPM) in battery-powered mobile systems attempts to achieve higher energy efficiency by selectively setting idle components to a sleep state. However, re-activating these components at a later time consumes a large amount of energy, which means that it will create a significant power draw from the battery supply in the system. This is known as the energy overhead of the “wakeup” operation. We start from the observation that, due to the rate capacity effect in Li-ion batteries which are commonly used to power mobile systems, the actual energy overhead is in fact larger than previously thought. Next we present a model-free reinforcement learning (RL) approach for an adaptive DPM framework in systems with bursty workloads, using a hybrid power supply comprised of Li-ion batteries and supercapacitors. Simulation results show that our technique enhances power efficiency by up to 9% compared to a battery-only power supply. Our RL-based DPM approach also achieves a much lower energy-delay product compared to a previously reported expert-based learning approach. Siyu Yue, Di Zhu 0002, Yanzhi Wang 0001, Massoud Pedram |
ICCD | 2 |