Hui Zhao 0013

dblp:39/6153-13 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3683-4077ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 LiteNoC: Developing Low-Cost Network-on-Chips for Deep Neural Networks
Khoa Ho, Siamak Biglari, Justin Garrigus, Hui Zhao 0013, Saraju P. Mohanty
ACM Great Lakes Symposium on VLSI4
2025 GenomeDPU: A Cost-Effective In-Memory Data Processing Unit for GPU-based Genome Analysis
abstract
In recent years, High-Throughput Sequencing (HTS) technologies have revolutionized the research field of genomics analysis. With the advances in HTS, the amount of genomics data has grown exponentially, creating a major challenge in efficient genomic analysis. Graphics Processing Units (GPU) have been widely used as a major platform to accelerate genomics analysis, but their massive parallel processing power cannot be fully exploited due to the memory bottleneck. Our investigation shows that memory accesses account for more than 60% of GPU pipeline stalls for Genomics applications.In this work, we propose a lightweight process-in-memory (PIM) architecture that can significantly speed up GPU-based genome analysis by minimizing off-chip data movement. We first identify several commonly used memory-bound functions in sequence alignments. Then, a data processing unit is developed to integrate processing capabilities directly within the memory system. The proposed PIM technique can significantly reduce the memory transfer between GPU processors and memory, thereby exploiting the GPU computational power more effectively. Compared with other Process-in-Memory techniques that involve complicated design and memory-embedded processing cores, GenomeDPU only adds simple logics to the DRAM interface and has a much smaller overhead. Our evaluation results show that GenomeDPU can improve performance by as much as 8.37x and reduce energy consumption by 55%.
Shouzhe Zhang, Zhuren Liu, Ruixiao Huang, Hui Zhao 0013
ICCAD4
2024 Designing Reconfigurable Interconnection Network of Heterogeneous Chiplets Using Kalman Filter
abstract
Heterogeneous chiplets have been proposed for accelerating high-performance computing tasks. Integrated inside one package, CPU and GPU chiplets can share a common interconnection network that can be implemented through the interposer. However, CPU and GPU applications have very different traffic patterns in general. Without effective management of the network resource, some chiplets can suffer significant performance degradation because the network bandwidth is taken away by communication-intensive applications. Therefore, techniques need to be developed to effectively manage the shared network resources. In a chiplet-based system, resource management needs to not only react in real-time but also be cost-efficient. In this work, we propose a reconfigurable network architecture, leveraging Kalman Filter to make accurate predictions on network resources needed by the applications and then adaptively change the resource allocation. Using our design, the network bandwidth can be fairly allocated to avoid starvation or performance degradation. Our evaluation results show that the proposed reconfigurable interconnection network can dynamically react to the changes in traffic demand of the chiplets and improve the system performance with low cost and design complexity.
Siamak Biglari, Ruixiao Huang, Hui Zhao 0013, Saraju P. Mohanty
ACM Great Lakes Symposium on VLSI3
2023 Genomics-GPU: A Benchmark Suite for GPU-accelerated Genome Analysis
abstract
Genomic analysis is the study of genes which includes the identification, measurement, or comparison of genomic features. Genomics research is of great importance to our society because it can be used to detect diseases, create vaccines, and develop drugs and treatments. As a type of general-purpose accelerators with massive parallel processing capability, GPUs have been recently used for genomics analysis. Developing GPUbased hardware and software frameworks for genome analysis is becoming a promising research area. To support this type of research, benchmarks are needed that can feature representative, concurrent, and diverse applications running on GPUs. In this work, we created a benchmark suite called Genomics-GPU, which contains 10 widely-used genomic analysis applications. It covers genome comparison, matching, and clustering for DNAs and RNAs. We also adapted these applications to exploit the CUDA Dynamic Parallelism (CDP), a recent advanced feature supporting dynamic GPU programming, to further improve the performance. Our benchmark suite can serve as a basis for algorithm optimization and also facilitate GPU architecture development for genomics analysis.
Zhuren Liu, Shouzhe Zhang, Justin Garrigus, Hui Zhao 0013
ISPASS4
2021 APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGA
abstract
In this paper, we introduce APCNN, which explores algorithm-hardware co-design and provides a CNN acceleration framework with multi-layer cooperative optimization and customized design on FPGA. In terms of the algorithm design, the pooling layer is moved before the non-linear activation function and normalization in APCNN, which we prove causes negligible accuracy loss; the pooling layer is then co-optimized with the convolutional layer by means of redundant multiplication elimination, local addition reuse, and global addition reuse. We further design a dedicated accelerator to take full advantage of convolutional-pooling cross-layer optimization to not only accelerate computation but also reduce on-off chip data communication on FPGA. We demonstrate that our novel APCNN can achieve 75% multiplication and 75% addition reduction in the best case. For on-off chip data communication, a max{Row,Col} /(Row x Col) percent of memory footprint can be eliminated, where Row and Col are the number of rows and columns in the activation feature map respectively. We have implemented a prototype of APCNN and evaluated its performance on LeNet-5 and VGG16 using both an accelerator-level cycle and energy model and an RTL implementation. Our experimental results show that APCNN achieves a 2.5× speedup and 4.7× energy efficiency compared with the dense CNN. (This research was supported in part by NSF grants CCF-1563750, OAC-2017564, and CNS-2037982.)
Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Hui Zhao 0013, Song Fu
FPGA6
2021 Distance-in-time versus distance-in-space
abstract
Cache behavior is one of the major factors that influence the performance of applications. Most of the existing compiler techniques that target cache memories focus exclusively on reducing data reuse distances in time (DIT). However, current manycore systems employ distributed on-chip caches that are connected using an on-chip network. As a result, a reused data element/block needs to travel over this on-chip network, and the distance to be traveled -- reuse distance in space (DIS) -- can be as influential in dictating application performance as reuse DIT. This paper represents the first attempt at defining a compiler framework that accommodates both DIT and DIS. Specifically, it first classifies data reuses into four groups: G1: (low DIT, low DIS), G2: (high DIT, low DIS), G3: (low DIT, high DIS), and G4: (high DIT, high DIS). Then, observing that reuses in G1 represent the ideal case and there is nothing much to be done in computations in G4, it proposes a "reuse transfer" strategy that transfers select reuses between G2 and G3, eventually, transforming each reuse to either G1 or G4. Finally, it evaluates the proposed strategy using a set of 10 multithreaded applications. The collected results reveal that the proposed strategy reduces parallel execution times of the tested applications between 19.3% and 33.3%.
Mahmut T. Kandemir, Xulong Tang, Hui Zhao 0013, Jihyun Ryoo, Mustafa Karaköy
PLDI3
2020 Collective Affinity Aware Computation Mapping
abstract
This work defines the concept of collective affinity. It is claimed that collective affinity has more potential than single core-centric affinity, for data locality optimization in manycores. The reason is that collective affinity captures the potential benefits of transferring computations originally assigned to one core to other cores. Next, building upon the collective affinity concept and a cache content estimation strategy, it presents a computation-to-core mapping strategy, specifically tuned for exploiting near data computing by reducing distance-to-data.
Mahmut T. Kandemir, Jihyun Ryoo, Hui Zhao 0013, Myoungsoo Jung, Mustafa Karaköy
PACT3
2020 AMOEBA: a coarse grained reconfigurable architecture for dynamic GPU scaling
abstract
Different GPU applications exhibit varying scalability patterns with network-on-chip (NoC), coalescing, memory and control divergence, and L1 cache behavior. A GPU consists of several Streaming Multi-processors (SMs) that collectively determine how shared resources are partitioned and accessed. Recent years have seen divergent paths in SM scaling towards scale-up (fewer, larger SMs) vs. scale-out (more, smaller SMs). However, neither scaling up nor scaling out can meet the scalability requirement of all applications running on a given GPU system, which inevitably results in performance degradation and resource under-utilization for some applications. In this work, we investigate major design parameters that influence GPU scaling. We then propose AMOEBA, a solution to GPU scaling through reconfigurable SM cores. AMOEBA monitors and predicts application scalability at run-time and adjusts the SM configuration to meet program requirements. AMOEBA also enables dynamic creation of heterogeneous SMs through independent fusing or splitting. AMOEBA is a microarchitecture-based solution and requires no additional programming effort or custom compiler support. Our experimental evaluations with application programs from various benchmark suites indicate that AMOEBA is able to achieve a maximum performance gain of 4.3x, and generates an average performance improvement of 47% when considering all benchmarks tested.
Xianwei Cheng, Hui Zhao 0013, Mahmut T. Kandemir, Beilei Jiang, Gayatri Mehta
ICS2
2019 A Low-Cost and Energy-Efficient NoC Architecture for GPGPUs
abstract
GPGPU accelerated systems demand high throughput in data communication in order to fully exploit thread-level parallelism. Most of current GPGPU Network-on-Chips (NoCs) employ topology adapted from CPUs, such as mesh and crossbar. However, the trade-off between performance and cost for such networks is sub-optimal, due to the unique traffic pattern of GPUs. In this work, we propose a novel NoC architecture called fused fat tree which modifies the fat tree to match GPU traffic pattern. By separately connecting memory controllers and computing cores to tree roots and leaves, protocol deadlocks can be avoided using just one physical network. However, this modification removes the advantage of path diversity in the original fat tree topology and makes the network vulnerable to hotspot-caused congestion. To solve this problem, we propose to fuse routers with side links to create multiple paths. A load-balancing routing algorithm is also proposed in order to increase network throughput. We also propose a novel preemptive bandwidth allocation scheme to improve resource utilization by taking advantage of request message slacks. Our evaluation results show that our design can improve performance by 46% while achieving 27 % and 25 % area and energy savings on the average.
Xianwei Cheng, Yang Zhao 0013, Mohammadreza Robaei, Beilei Jiang, Hui Zhao 0013, Juan Fang 0004
ANCS5
2018 Packet pump: overcoming network bottleneck in on-chip interconnects for GPGPUs
abstract
In order to fully exploit GPGPU's parallel processing power, on-chip interconnects need to provide bandwidth efficient data communication. GPGPUs exhibit a many-to-few-to-many traffic pattern which makes the memory controller connected routers the network bottleneck. Inefficient design of conventional routers causes long queues of packets blocked at memory controllers and thus greatly constrained the network bandwidth. In this work, we employ heterogeneous design techniques and propose a novel decoupled architecture for routers connected with memory controllers. To further improve performance, we propose techniques called Injection Virtual Circuit and Memory-aware Adaptive Routing. We show that our scheme can effectively eliminate NoC bottleneck and improve performance by 78% on average.
Xianwei Cheng, Yang Zhao 0013, Hui Zhao 0013, Yuan Xie 0001
DAC3
2017 DEMM: A Dynamic Energy-Saving Mechanism for Multicore Memories
abstract
Since main memory system contributes to a large and increasing fraction of server/datacenter energy consumption, there have been several efforts to reduce its power and energy consumption. DVFS schemes have been used to reduce the memory power, but they come with a performance penalty. In this work, we propose DEMM, an OS-based, high performance DVFS mechanism that reduces memory power by dynamically scaling individual memory channel frequencies/voltages. Our strategy also involves clustering the running applications based on their sensitivities to memory latency, and assigning memory channels to the application clusters. We introduce a new metric called Discrete Misses per Kilo Cycle (DMPKC) to capture the performance sensitivities of the applications to memory frequency modulation. DEMM allows us to save power in the memory system with negligible impact on performance. We demonstrate around 25% savings in the memory system energy and 10% savings in the total system energy, with only a 4% loss in workload performance.
Akbar Sharifi, Wei Ding 0008, Diana R. Guttman, Hui Zhao 0013, Xulong Tang, Mahmut T. Kandemir, Chita R. Das
MASCOTS4
2015 Phase Detection with Hidden Markov Models for DVFS on Many-Core Processors
abstract
The energy concerns of many-core processors are increasing with the number of cores. We provide a new method that reduces energy consumption of an application on many-core processors by identifying unique segments to apply dynamic voltage and frequency scaling (DVFS). Our method, phase-based voltage and frequency scaling (PVFS), hinges on the identification of phases, i.e., Segments of code with unique performance and power attributes, using hidden Markov Models. In particular, we demonstrate the use of this method to target hardware components on many-core processors such as Network-on-Chip (NoC). PVFS uses these phases to construct a static power schedule that uses DVFS to reduce energy with minimal performance penalty. This general scheme can be used with a variety of performance and power metrics to match the needs of the system and application. More importantly, the flexibility in the general scheme allows for targeting of the unique hardware components of future many-core processors. We provide an in-depth analysis of PVFS applied to five threaded benchmark applications, and demonstrate the advantage of using PVFS for 4 to 32 cores in a single socket. Empirical results of PVFS show a reduction of up to 10.1% of total energy while only impacting total time by at most 2.7% across all core counts. Furthermore, PVFS outperforms standard coarse-grain time-driven DVFS, while scaling better in terms of energy savings with increasing core counts.
Joshua Dennis Booth, Jagadish Kotra, Hui Zhao 0013, Mahmut T. Kandemir, Padma Raghavan
ICDCS3
2015 Memory Row Reuse Distance and its Role in Optimizing Application Performance
abstract
Continuously increasing dataset sizes of large-scale applications overwhelm on-chip cache capacities and make the performance of last-level caches (LLC) increasingly important. That is, in addition to maximizing LLC hit rates, it is becoming equally important to reduce LLC miss latencies. One of the critical factors that influence LLC miss latencies is row-buffer locality (i.e., the fraction of LLC misses that hit in the large buffer attached to a memory bank). While there has been a plethora of recent works on optimizing row-buffer performance, to our knowledge, there is no study that quantifies the full potential of row-buffer locality and impact of maximizing it on application performance.
Mahmut T. Kandemir, Hui Zhao 0013, Xulong Tang, Mustafa Karaköy
SIGMETRICS2
2012 A hybrid NoC design for cache coherence optimization for chip multiprocessors
abstract
On chip many-core systems, evolving from prior multi-processor systems, are considered as a promising solution to the performance scalability and power consumption problems. The long communication distance between the traditional multi-processors makes directory-based cache coherence protocols better solutions compared to bus-based snooping protocols even with the overheads from indirections. However, much smaller distances between the CMP cores enhance the reachability of buses, revitalizing the applicability of snooping protocols for cache-to-cache transfers. In this work, we propose a hybrid NoC design to provide optimized support for cache coherency. In our design, on-chip links can be dynamically configured as either point-to-point links between NoC nodes or short buses to facilitate localized snooping. By taking advantage of the best of both worlds, bus-based snooping coherency and NoC-based directory coherency, our approach brings both power and performance benefits.
Hui Zhao 0013, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir, Mary Jane Irwin
DAC1
2011 Exploring heterogeneous NoC design space
abstract
The Network-on-Chip (NoC) plays a crucial role in designing low cost chip multiprocessors (CMPs) as the number of cores on a chip keeps increasing. However, buffers in NoC routers increase the cost of CMPs in terms of both area and power. Recently, bufferless routers have been proposed to reduce such costs by removing buffers from the routers. However, bufferless routers can provide competitive performance only when network utilization is moderate. In this paper, we propose a novel heterogeneous design that employs both buffered and bufferless routers in the same NoC to achieve high performance at low cost. We evaluate a variety of plans to place buffered and bufferless routers in an NoC based CMP according to performance requirements and power allowances. In order to take full advantage of these heterogeneous NoCs, we also propose novel strategies for buffered-router-aware application thread mapping and a routing algorithm (once the router placement is fixed). Our evaluations show that, by utilizing the techniques we proposed, a heterogeneous NoC does not only achieve performance comparable to that of the NoCs with buffered routers but also reduces buffer costs and energy consumption.
Hui Zhao 0013, Mahmut T. Kandemir, Wei Ding 0008, Mary Jane Irwin
ICCAD1
2011 Feedback control based cache reliability enhancement for emerging multicores
abstract
Focusing on data reliability, we propose a control theory centric approach designed to improve transient error resilience in shared caches of emerging multicores while satisfying performance goals. The proposed scheme takes, as input, two quality of service (QoS) specifications: performance QoS and reliability QoS. The first of these indicates the minimum workload-wide cache (L2) hit rate value acceptable, whereas the second one captures the reliability bound on an application basis, with the help of a metric called the Reads-with-Replica (RwR). We present an extensive experimental evaluation of the proposed scheme on various workloads formed using the applications from the SPEC2006 benchmark suite. The proposed scheme is able to satisfy, in most of the tested cases, both performance and reliability QoS targets, by successfully modulating the total size of the data replication area and partitioning of this area among the co-runner applications. The collected results also show that our scheme achieves consistent improvements under different values of the major simulation parameters.
Hui Zhao 0013, Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir
ICCAD1
2010 Feedback control for providing QoS in NoC based multicores
abstract
In this paper, we employ formal feedback control theory to achieve desired communication throughput across a network-on-chip (NoC) based multicore. When the output of the system needs to follow a certain reference input over time, our controller regulates the system to obtain the desired effect on the output. In this work, targeting a multicore that executes multiple applications simultaneously, we demonstrate how to design and employ a PID (Proportional Integral Derivative) controller to obtain the desired throughput for communications by tuning the weights of the virtual channels of the routers in the NoC. We also propose a global controller architecture that implements policies to handle situations in which the network cannot provide the overlapping communications with sufficient resources or the throughputs of the communications can be enhanced (beyond their specified values) due to the availability of excess resources. Finally, we discuss how our novel control architecture works under different scenarios by presenting experimental results obtained using four embedded applications. These results show how the global controller adjusts the virtual channels weights to achieve the desired throughputs of different communications across the NoC, and as a result, the system output successfully tracks the specified input.
Akbar Sharifi, Hui Zhao 0013, Mahmut T. Kandemir
DATE2
2010 A special-purpose compiler for look-up table and code generation for function evaluation
abstract
Elementary functions are extensively used in computer graphics, signal and image processing, and communication systems. This paper presents a special-purpose compiler that automatically generates customized look-up tables and implementations for elementary functions under user given constraints. The generated implementations include a C/C++ code that can be used directly by applications running on multicores, as well as a MATLAB-like code that can be translated directly to a hardware module on FPGA platforms. The experimental results show that our solutions for function evaluation bring significant performance improvements to applications on multicores as well as significant resource savings to designs on FPGAs.
Lanping Deng, Praveen Yedlapalli, Sai Prashanth Muralidhara, Hui Zhao 0013, Mahmut T. Kandemir, Chaitali Chakrabarti, Nikos Pitsianis, Xiaobai Sun
DATE5