VLDB 2026 Research / reviewers in the wild / expert
Yunping Zhao
dblp:156/8745
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-5600-3740ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LIRL-NoC: Long-Range Link Insertion Using Reinforcement Learning for Network-on-Chips
Yiqun Lang, Yuhan Tang, Lizhou Wu, Sheng Ma, Yunping Zhao |
ISCAS | 7 |
| 2025 | SpMARD: A Sparse-Sparse Matrix Multiplication Accelerator with Reconfigurable Dataflow for DNN WorkloadsabstractDeep learning becomes increasingly popular, and its main workload is Sparse-Sparse Matrix Multiplication (SpMSpM). Most SpMSpM accelerators usually only support a single dataflow. Different dataflows have different performance in different computing environments. Therefore, the single-dataflow accelerator cannot maintain the highest performance in all environments. Compared with single-dataflow accelerators, multi-dataflow accelerators provide flexible options for different workloads and improve the overall performance. Flexagon, Sparm, and SPADA are state-of-the-art multi-dataflow accelerators. However, the computation process of Flexagon and Sparm is not fully pipelined, and SPADA cannot support inner product dataflow. Additionally, Flexagon, Sparm, and SPADA cannot switch dataflows quickly and accurately. Inspired by these observations, we present SpMARD, a SpMSpM accelerator with reconfigurable dataflow. The computation process of SpMARD is fully pipelined, and SpMARD can support six dataflow variants simultaneously. Through the design of a Two-stage Pipeline Adder Network (TPAN) and a Position-based Psum Array (PPA), SpMARD can execute element-level merging, which can hide the merging overhead. Through the quantitative analysis of dataflows, we implement a Dataflow Switcher (DSwitcher), which can switch dataflows more efficiently. For the SpMSpM workload, the performance (GOPS) of the SpMARD we proposed is 1.27 times that of Flexagon, 1.18 times that of Sparm, and 1.22 times that of SPADA. Bo Wang 0159, Sheng Ma, Yunping Zhao, Shengbai Luo, Lizhou Wu, Dongsheng Li 0001, Zhuojun Chen |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Intra- and Inter-Layer Scheduling Exploration and Optimization for ReRAM-Based DNN AcceleratorsabstractResistive Random Access Memory (ReRAM) based architectures have shown great potential for realizing energy-efficient Deep Neural Network (DNN) acceleration. When deploying a DNN, the ReRAM-based designs need a scheduling scheme to translate massive hardware resources into actual performance. Different scheduling schemes would lead to different levels of data reuse and computational parallelism, resulting in different energy efficiency and performance. However, the ReRAM-based scheduling scheme faces the following limitations. First, current studies mainly focus on intra-layer scheduling scheme optimizations by using the Weight Stationary (WS) data flow. These studies ignore the difference between layers and limit optimization opportunities. Second, there is no systematic definition and analysis for inter-layer scheduling schemes. Third, there is no co-optimization study on intra- and inter-layer scheduling schemes. Fourth, the complex network structure leads to intricate inter-layer data dependency, making the optimization of the scheduling scheme more challenging. These limitations restrict the comprehensive understanding of the scheduling schemes.Inspired by these observations, we identify the fundamental impact of intra-layer scheduling schemes on ReRAM-based designs, including the WS and Input Stationary (IS) data flows. We also systematically define and analyze inter-layer scheduling schemes according to different combinations of data flows, including the WS-WS, IS-IS, WS-IS, and IS-WS data flows. We analyze and explore different resource allocation strategies for these schemes. We also propose the intra- and inter-layer co-optimization to further improve performance and energy efficiency. Then, we propose a method for building a hybrid scheduling scheme by flexibly combining these inter-layer scheduling schemes for complex networks. Finally, we seek the potential to improve performance and energy efficiency for hybrid scheduling schemes. For deploying the MobileNet-V1, ResNet-18, VGG-16, and AlexNet, the hybrid scheduling scheme improves performance by 10.2×~130.7×, 1.9×~16.5×, 7.0×~56×, and 1×~153.1× than the WS-WS, IS-IS, IS-WS, and WS-IS based scheduling schemes, respectively. Similarly, the power efficiency can also be increased by 15× and 26× than the WS-WS and IS-IS based scheduling schemes, respectively. Yunping Zhao, Sheng Ma, Yuhua Tang |
IEEE Trans. Computers | 1 |
| 2025 | Tradeoff Performance and Energy Efficiency by Optimizing the Data Flow for PIM ArchitecturesabstractThe processing-in-memory (PIM) architecture becomes a promising candidate for deep learning accelerators by integrating computation and memory. Most PIM-based studies improve the performance and energy efficiency by using the weight stationary (WS) data flow due to its high parallelism. However, the WS data flow has some fundamental limitations. First, the WS data flow has huge activation movements between on-chip memory and off-chip memory due to the limited memory space of the resistive random-access memory (ReRAM) array. Second, the WS data flow needs to read the input activation repeatedly according to the convolution window. These data movements decrease the energy efficiency and performance of the PIM architecture. To address these issues, the input stationary (IS) data flow stores activations instead of weights to reduce data movements. But the IS data flow faces some challenges. First, the data dependency between adjacent layers limits the performance. Second, there are huge across-array computations due to the special mapping method. Third, the previous IS data flow cannot realize the high parallelism. Fourth, the IS data flow depends on the 3-D ReRAM structure. To address these issues, we propose a novel data flow for PIM architectures. We optimize the IS data flow to decrease the activation movement and propose a parallel computing method to realize high parallelism and reduce the across-array computations. We identify and analyze the fundamental limitations and impact of different interlayer data flows, including the WS-WS, IS-IS, WS-IS, and IS-WS. We also propose a method to build a hybrid data flow by combining these interlayer data flows to tradeoff performance and energy consumption. Our experimental results and analysis demonstrate the potential of our design. The performance and energy efficiency of our design reach 0.13–1.77 TFLOPS and 61–85 TOPS/J, respectively. Compared to the state-of-the-art design, the NEBULA, our design can improve performance by$1.4\times $,$2.3\times $, and$3.5\times $for deploying the MobileNet-V1, ResNet-18, and VGG-16, and also can improve energy efficiency by$3.3\times $,$2\times $, and$2\times $, respectively. Yunping Zhao, Sheng Ma, Yuhua Tang, Hengzhu Liu, Dongsheng Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | HPA: A Hybrid Data Flow for PIM ArchitecturesabstractThe Processing- In- Memory (PIM) architecture becomes a promising candidate for deep learning acceleration by integrating computation and memory. Due to the simple mapping method and high parallelism, the Weight Stationary (WS) data flow is widely used in PIM-based studies to improve performance and energy efficiency. However, the WS data flow leads to huge activation movements, becoming the bottleneck for reducing latency and energy consumption. To address this issue, the Input Stationary (IS) data flow stores activations instead of weights in the PIM architecture to reduce data movements. However, the traditional IS data flow also faces several challenges. First, the inter-layer data dependence and imbalance workload decrease pipeline efficiency. Second, the across-array computation reduces energy efficiency and performance. Third, the traditional IS data flow relies on the 3D ReRAM structure. Inspired by these observations, we propose a Hybrid data flow for PIM Architectures, named HPA. The HPA contains novel intra-layer and inter-layer data flows, named the PP-IS data flow and the IS- WS hybrid data flow, respectively. The PP- IS data flow optimizes the data mapping strategy and computing method to reduce activation movements. In addition, the PP- IS data flow uses the parallel computing method to decrease across-array computations. Based on the novel intra-layer data flow, we propose the IS- WS hybrid data flow to trade off performance and energy efficiency. Finally, we optimize the pipeline for the hybrid data flow to mitigate data depen-dence and balance inter-layer workloads, improving pipeline efficiency. Our experimental results and analysis demonstrate the potential of the HPA. The performance and power efficiency of the HPA reaches 1.64$GFLOPS\sim 63$G F LO P Sand 2.1$TOPS/W\sim 151\ TOPS/W$, respectively. Compared to the state-of-the-art design, the NEBULA, the HPA can significantly improve power efficiency and performance by$22.1\times$and$7.8\times$, respectively, when deploying the MobileNet VI. Sheng Ma, Yunping Zhao, Yuhua Tang |
ICCD | 2 |
| 2024 | SAC: An Ultra-Efficient Spin-based Architecture for Compressed DNNsabstractDeep Neural Networks (DNNs) have achieved great progress in academia and industry. But they have become computational and memory intensive with the increase of network depth. Previous designs seek breakthroughs in software and hardware levels to mitigate these challenges. At the software level, neural network compression techniques have effectively reduced network scale and energy consumption. However, the conventional compression algorithm is complex and energy intensive. At the hardware level, the improvements in the semiconductor process have effectively reduced power and energy consumption. However, it is difficult for the traditional Von-Neumann architecture to further reduce the power consumption, due to the memory wall and the end of Moore’s law. To overcome these challenges, the spintronic device based DNN machines have emerged for their non-volatility, ultra low power, and high energy efficiency. However, there is no spin-based design that has achieved innovation at both the software and hardware level. Specifically, there is no systematic study of spin-based DNN architecture to deploy compressed networks. In our study, we present an ultra-efficient Spin-based Architecture for Compressed DNNs (SAC), to substantially reduce power consumption and energy consumption. Specifically, we propose a One-Step Compression algorithm (OSC) to reduce the computational complexity with minimum accuracy loss. We also propose a spin-based architecture to realize better performance for the compressed network. Furthermore, we introduce a novel computation flow that enables the reuse of activations and weights. Experimental results show that our study can reduce the computational complexity of compression algorithm from 𝒪( Tk 3 to 𝒪( k 2 log k ), and achieve 14× ∼ 40× compression ratio. Furthermore, our design can attain a 2× enhancement in power efficiency and a 5× improvement in computational efficiency compared to the Eyeriss. Our models are available at an anonymous link https://bit.ly/39cdtTa . Yunping Zhao, Sheng Ma, Hengzhu Liu, Libo Huang 0002 |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | SAL: Optimizing the Dataflow of Spin-based Architectures for Lightweight Neural NetworksabstractAs the Convolutional Neural Network (CNN) goes deeper and more complex, the network becomes memory-intensive and computation-intensive. To address this issue, the lightweight neural network reduces parameters and Multiplication-and-Accumulation (MAC) operations by using the Depthwise Separable Convolution (DSC) to improve speed and efficiency. Nonetheless, the energy efficiency of classical Von Neumann architectures for CNNs is limited due to the memory wall challenge. Spin-based architectures have the potential to address this challenge thanks to the integration of memory and computing with ultra-high energy efficiency. However, deploying the DSC on spin-based architectures with the traditional dataflow leads to huge activation movements and low hardware utilization. Moreover, the inter-layer data dependency of neural networks increases latency. These factors become the bottleneck of improving energy efficiency and performance. Inspired by these challenges, we propose a novel dataflow on Spin-based Architectures for Lightweight neural networks (SAL). The novel dataflow replaces convolution unrolling by selecting activations in the crossbar according to the convolution window and also realizes the inter-layer data reuse. Moreover, the novel dataflow also reduces the latency due to the data dependency between layers, realizing higher performance. To the best of our knowledge, this is the first design to use hybrid dataflow for the PIM architecture. We also optimize the structure of the spin-based crossbar and the pipeline based on the dataflow to achieve better data reuse and computational parallelism. For deploying the MobileNet V1, the novel dataflow improves the hardware utilization by 23×∼ 105× and reduces the data traffic by 1.09×∼ 18.6×. Compared with the NEBULA, a spin-based non-Von Neumann architecture, the SAL reduces the energy consumption by 4× and improves the performance by 7.3×, which are 0.32 mJ and 10.43 GOPs -1 , respectively. Moreover, the SAL improves power efficiency over 29 times more than the NEBULA. Compared with the Eyeriss, the SAL improves the energy efficiency by four orders of magnitude. Yunping Zhao, Sheng Ma, Hengzhu Liu, Dongsheng Li 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | EPHA: An Energy-efficient Parallel Hybrid Architecture for ANNs and SNNsabstractArtificial neural networks (ANNs) and spiking neural networks (SNNs) are two general approaches to achieve artificial intelligence (AI). The former have been widely used in academia and industry fields; the latter, SNNs, are more similar to biological neural networks and can realize ultra-low power consumption, thus have received widespread research attention. However, due to their fundamental differences in computation formula and information coding, the two methods often require different and incompatible platforms. Alongside the development of AI, a general platform that can support both ANNs and SNNs is necessary. Moreover, there are some similarities between ANNs and SNNs, which leaves room to deploy different networks on the same architecture. However, there is little related research on this topic. Accordingly, this article presents an energy-efficient, scalable, and non-Von Neumann architecture (EPHA) for ANNs and SNNs. Our study combines device-, circuit-, architecture-, and algorithm-level innovations to achieve a parallel architecture with ultra-low power consumption. We use the compensated ferrimagnet to act as both synapses and neurons to store weights and perform dot-product operations, respectively. Moreover, we propose a novel computing flow to reduce the operations across multiple crossbar arrays, which enables our design to conduct large and complex tasks. On a suite of ANN and SNN workloads, the EPHA is 1.6× more power-efficient than a state-of-the-art design, NEBULA, in the ANN mode. In the SNN mode, our design is 4 orders of magnitude more than the Loihi in power efficiency. Yunping Zhao, Sheng Ma, Hengzhu Liu, Libo Huang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2023 | Consistency of Multiple Kernel ClusteringabstractConsistency plays an important role in learning theory. However, in multiple kernel clustering (MKC), the consistency of kernel weights has not been sufficiently investigated. In this work, we fill this gap with a non-asymptotic analysis on the consistency of kernel weights of a novel method termed SimpleMKKM. Under the assumptions of the eigenvalue gap, we give an infinity norm bound as $\widetilde{\mathcal{O}}(k/\sqrt{n})$, where $k$ is the number of clusters and $n$ is the number of samples. On this basis, we establish an upper bound for the excess clustering risk. Moreover, we study the difference of the kernel weights learned from $n$ samples and $r$ points sampled without replacement, and derive its upper bound as $\widetilde{\mathcal{O}}(k\cdot\sqrt{1/r-1/n})$. Based on the above results, we propose a novel strategy with Nyström method to enable SimpleMKKM to handle large-scale datasets with a theoretical learning guarantee. Finally, extensive experiments are conducted to verify the theoretical results and the effectiveness of the proposed large-scale strategy. Weixuan Liang, Xinwang Liu 0002, Yong Liu 0018, Chuan Ma 0001, Yunping Zhao, Zhe Liu 0001, En Zhu |
ICML | 5 |
| 2022 | Convolutional Neural Network Accelerator for Compression Based on Simon k-meansabstractConvolutional Neural Networks (CNN) are popular models widely used in image classification, target recognition, and other fields. FPGA-based accelerators for CNN are a standard method in recent years to reduce CNN's inference time and energy efficiency. However, the limitations of on-chip storage space and computing resources introduce deep compression. Contrary to most compression algorithms that pay no attention to the underlying hardware acceleration strategy and hardware-only accelerators, this paper introduces a novel model compression scheme with software and hardware collaboration for accelerating inference. First, we propose a pre-processing algorithm named Simon k-means based on clustering to quantify trained weight to speed up inference. Next, we propose a new encoding method for the quantized weight, significantly reducing the model's storage size. Finally, we give the architecture design of the accelerator using the quantized weight to accelerate the convolution. We have evaluated many popular CNNs in image classification tasks on various data sets. Experiments show that the number of multiply-accumulate operations on the convolutional layer can be reduced 66.6% with a slight loss of precision. Yunping Zhao, Jianzhuang Lu, Zerun Li |
IJCNN | 2 |
| 2020 | Dynamic GMMU Bypass for Address Translation in Multi-GPU Systems
Jinhui Wei, Jianzhuang Lu, Yunping Zhao |
NPC | 5 |