Longbing Zhang

dblp:74/6902 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 CUTE-XS: Bottleneck Analysis and Collaborative Optimization for Integrating CUTE Into XiangShan
Junyu Yue, Chongxi Wang, Jinpeng Ye, Jianan Xie, Longbing Zhang, Fuxin Zhang
APPT8
2025 LASM: A Lightweight and General TEE Secure Monitor Framework
Baojun Wang, Huandong Wang, Changbin Xu, Longbing Zhang
APPT6
2024 High-Utilization GPGPU Design for Accelerating GEMM Workloads: An Incremental Approach
abstract
General Purpose Graphics Processing Units (GPGPUs) have been employed primarily in domains such as graphics acceleration and high-performance computing in the past. However, the rise of artificial intelligence (AI), particularly the computational demands associated with matrix multiplications in AI models, has presented formidable demands on the computational power of GPGPUs. Consequently, the design of matrix multiplication units within GPGPUs and ensuring their utilization have become key issues in optimizing AI workloads. This paper explores an incremental design approach, building upon a Single Instruction Multiple Threads (SIMT) GPGPU architecture, to facilitate General Matrix Multiply (GEMM) acceleration. This approach encompasses not only the design of matrix units within the stream processors but, more crucially, the optimization of the data path within the GPGPU to maximize the utilization of the matrix units. We present a practical demonstration of our approach through the fabrication of a GPGPU on a 12 nm CMOS process node, achieving a core clock speed of 1 GHz and INT8 peak performance of 8 TOPS with memory bandwidth limited to 32 GB/s LPDDR4-4000. Notably, this design results in only a 6.57% increase in chip area compared to the original GPGPU design. In a series of fair GEMM workload tests, the GPGPU implemented in this work outperforms the recent three generations of NVIDIA GPGPUs—V100, T4, and A100—in terms of matrix unit utilization.
Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang
ISCAS6
2023 RBGC: Repurpose the Buffer of Fixed Graphics Pipeline to Enhance GPU Cache
abstract
The limited cache size of GPU in general-purpose computing hinders the execution efficiency of thousands of concurrent threads. Several techniques have been proposed to increase the cache size per thread, such as repurposing shared memory and register files as a cache to reduce contention. However, these studies only focus on improving the general-purpose computing hardware structure, ignoring the fixed graphics pipeline hardware structure. To solve this issue, we propose repurposing the buffer of the fixed graphics pipeline as a victim cache to enhance the GPU cache. This strategy utilizes the idle fixed graphics pipeline buffer as a victim cache for general-purpose computing tasks. The victim cache returns data to the L1 data cache when a load is located at the victim cache. We also optimize the victim cache management strategy by monitoring load reusability and only allocating cache lines to loads with higher reusability. This optimization improves the victim cache efficiency. Our experimental results show that the RBGC achieves a 39.1% performance improvement with minimal hardware overhead compared to the baseline GPU architecture.
Longbing Zhang, Fuxin Zhang
ACM Great Lakes Symposium on VLSI2
2023 Randomized Testing Framework for Dissecting NVIDIA GPGPU Thread Block-To-SM Scheduling Mechanisms
abstract
NVIDIA General Purpose Graphics Processing Units (GPGPUs) have been extensively utilized in scientific computing, artificial intelligence acceleration, and numerous other fields. The compute unified device architecture (CUDA) programming interface facilitates access to the enormous computing power of NVIDIA GPGPUs. However, the thread block-to-stream multiprocessor (SM) scheduler of NVIDIA GPGPUs has remained a black box, presenting challenges for performance modeling of CUDA programs, especially those featuring concurrent kernel execution, as well as hindering further exploration of the full potential of NVIDIA GPGPUs. To address these issues, we have devised a randomized testing framework targeting thread block-to-SM scheduling. In this framework, concurrent kernel sequences are randomly generated and assessed by a self-implemented scheduling predictor along with an actual NVIDIA GPGPU. Through iteratively scrutinizing the discrepancies between predictions and actual results, we dissect the key thread block-to-SM scheduling mechanisms intrinsic to NVIDIA GPGPUs, revealing their underlying logic and considerations. This process continuously refines the scheduling predictor until it accurately aligns with real-world NVIDIA GPGPU scheduling. Furthermore, we investigate the monitoring and allocation mechanisms for distinct resources in NVIDIA GPGPUs and SMs, elucidating the essential connection between resource management and thread block scheduling. Our comprehensive and precise analysis of Nvidia GPGPU thread block-to-SM scheduling contributes significantly to constructing more accurate simulators for GPGPU performance modeling. Moreover, it facilitates optimizing CUDA programs based on the scheduling mechanisms to exploit more parallelism among thread blocks of concurrent kernels.
Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang
ICPADS6
2022 Ring-ExpLWE: A High-Performance and Lightweight Post-Quantum Encryption Scheme for Resource-Constrained IoT Devices
abstract
As the Internet of Things (IoT) expands explosively existing network connections, the transmission and processing of private data is facing more serious threats of leakage and theft. Classical public key encryption schemes are difficult to guarantee strong security protection, because the mathematically hard problems they rely on are no longer difficult to solve under the rapid development of quantum computing. Therefore, a more high-performance and quantum-resistant encryption scheme Ring-ExpLWE is proposed, in which the error vector is sampled in the exponential distribution instead of the binary distribution in the previous Ring-BinLWE. We evaluate the Ring-ExpLWE’s security level by analyzing the runtime under quantum hybrid attack and comparing the standard deviation of the noise polynomial coefficients. Compared with Ring-BinLWE, the proposed Ring-ExpLWE requires larger runtime for quantum hybrid attack and has a more discrete noise distribution. Therefore, Ring-ExpLWE can provide a higher security level under the same parameter set. Moreover, the high-performance software and hardware implementations for the Ring-ExpLWE scheme are proposed, respectively. Based on the Cortex-M3 microprocessor platform, encryption, and decryption only require 35.6 and 17.8 ms in our software implementation, respectively. Compared with the previous Ring-BinLWE schemes, while significantly improving the security level, the Area$\times $Time (AT) of our high-performance and lightweight hardware implementations is reduced by 49.2% and 49.5%, respectively, when the FPGA platform is Spartan 6.
Dongdong Xu 0002, Xiang Wang 0006, Yuanchao Hao, Haoyu Jia, Haifeng Dong, Longbing Zhang
IEEE Internet Things J.8
2019 EDOA: an efficient delay optimization approach for mixed-polarity Reed-Muller logic circuits under the unit delay model
Zhenxue He, Limin Xiao 0002, Fei Gu 0001, Zhisheng Huo, Mingfa Zhu, Longbing Zhang, Rui Liu 0007, Xiang Wang 0006
Frontiers Comput. Sci.8
2017 An Efficient Polarity Optimization Approach for Fixed Polarity Reed-Muller Logic Circuits Based on Novel Binary Differential Evolution Algorithm
Zhenxue He, Guangjun Qin, Limin Xiao 0002, Fei Gu 0001, Zhisheng Huo, Haitao Wang 0017, Longbing Zhang, Jianbin Liu, Xiang Wang 0006
NPC8
2017 An efficient and fast polarity optimization approach for mixed polarity Reed-Muller logic circuits
Zhenxue He, Limin Xiao 0002, Fei Gu 0001, Tongsheng Xia, Shubin Su, Zhisheng Huo, Longbing Zhang, Xiang Wang 0006
Frontiers Comput. Sci.8
2017 A Power and Area Optimization Approach of Mixed Polarity Reed-Muller Expression for Incompletely Specified Boolean Functions
Zhenxue He, Limin Xiao 0002, Fei Gu 0001, Zhisheng Huo, Guangjun Qin, Mingfa Zhu, Longbing Zhang, Rui Liu 0007, Xiang Wang 0006
J. Comput. Sci. Technol.8
2016 EMA-FPRMs: An efficient minimization algorithm for fixed polarity Reed-Muller expressions
abstract
Fixed polarity Reed-Muller expressions (FPRMs) are well-suited for many practical applications due to they have many excellent properties. In order to obtain an optimal FPRM with fewest product terms, we propose an efficient minimization algorithm (EMA-FPRMs) for FPRMs. The main idea behind the EMA-FPRMs is that, firstly, the incompletely specified Boolean function is transformed into the zero polarity incompletely specified fixed polarity RM expression (ISFPRM) by using the proposed ISFPRM acquisition algorithm; secondly, the polarity and allocation of don't care terms of ISFPRM is encoded as chromosome; lastly, the optimal FPRM with fewest product terms is obtained by using genetic algorithm (GA), in which the FPRM that corresponds to the given chromosome is obtained by using the proposed chromosome conversion algorithm. The experimental results on MCNC benchmark circuits show that compared with the traditional polarity optimization approach which neglects the don't care terms, the EMA-FPRMs is highly effective in minimizing the number of product terms of FPRMs. Moreover, the EMA-FPRMs is faster than the GA based minimization algorithm which also considers the don't care terms.
Zhenxue He, Limin Xiao 0002, Longbing Zhang, Fei Gu 0001, Zhisheng Huo, Mingfa Zhu, Rui Liu 0007, Xiang Wang 0006
FPT3
2015 Stable Matching Scheduler for Single-ISA Heterogeneous Multi-core Processors
Lei Wang 0015, Shaoli Liu, Longbing Zhang, Junhua Xiao
APPT4
2015 MRP: mix real cores and pseudo cores for FPGA-based chip-multiprocessor simulation
Xinke Chen, Guangfei Zhang, Huandong Wang, Ruiyang Wu 0001, Longbing Zhang
DATE6
2014 Archipelago: A Floorplan Optimized for Concurrent Multiple Applications on Network-on-Chip
abstract
In practice, a chip multiprocessor (CMP) with many nodes often runs multiple applications simultaneously and nodes allocated to different applications seldom communicate. Leveraging the non-uniformity of communication, it is possible to optimize the Network-on-Chip (NoC) performance for the case of single chip running multiple applications through reducing the physical links lengths between nodes which tend to be allocated to the same application. In this paper, we propose a novel, nearly zero-cost floor plan method called Archipelago which rearranges the lengths of physical links by flipping hard IP of nodes. With communication monitoring and dynamic thread scheduling, Archipelago significantly reduces the average latency of NoC for concurrent multiple small- and medium-size applications. Experimental results show that Archipelago averagely reduces the packet latency by 19.10% and maximum to 24.4% for three-node applications on 4 × 4 mesh NoC.
Shaoli Liu, Lei Wang 0015, Longbing Zhang, Meng Wen
NAS4
2011 Processes Scheduling on Heterogeneous Multi-core Architecture with Hardware Support
abstract
Heterogeneous Chip Multi-Processors (heter-CMP) provide suitable resources to various applications and could get more benefits on performance than homogeneous CMP. To fully develop the performance of the heter-CMP system, the applications should be scheduled to the proper cores by the scheduler of the operating system according to their behaviours. Besides, the last level cache (LLC) miss rate is often used as the metric to identify the application features. However, the LLC miss rate could not reflect the application behaviours accurately since the memory access delay, the most important application character is related to patterns of the memory access and many other issues. To improve the accuracy of scheduling on the heter-CMP system, this paper supplies performance counters for the LLC miss penalty to monitor application behaviours. Based on that, the latency-aware scheduling algorithm for asymmetry chip multi-processors (LA-ACMP) is proposed and implemented. Those hardware supports are implemented on the Godson-3 multi-core RTL and simulator and the LA-ACMP algorithm is integrated to the Linux kernel. The performance evaluation on those platforms shows that the LA-ACMP algorithm with LLC miss penalty outperforms the algorithm with LLC miss rate about 32.4%, and outperforms the original schedule algorithm about 18.4%.
Shouqing Hao, Longbing Zhang
NAS3
2009 Efficiency-Aware QoS DRAM Scheduler
abstract
For most SoCs, off-chip DRAM is an important resource that is shared by many heterogeneous function units(FU).To meet different memory access requirements by these FUs,it is crucial that the memory subsystem is capable of providing different quality of service(QoS).Due to the nature of DRAM, the available bandwidth greatly depends on the memory access sequence. However,conventional schedulers are not aware of the variable bandwidth.In this paper, a QoS scheduler is proposed by recognizing the inefficiency caused by ongoing memory accesses.The experimental results show that, the scheduler can provide bandwidth guarantee for bandwidth sensitive FUs even in the worst case scenarios. And low latency for latency sensitive FUs can be achieved when the bus is below the saturation point.
Menghao Su, Yunji Chen, Longbing Zhang
NAS5
2006 A Hybrid Hardware/Software Generated Prefetching Thread Mechanism on Chip Multiprocessors
Hou Rui, Longbing Zhang, Weiwu Hu
Euro-Par2
2004 Parallel program performance evaluation and their behavior analysis on an OpenMP cluster
abstract
The OpenMP API is an emerging standard for parallel programming on shared memory multiprocessors. In order to run an OpenMP program on a cluster, one feasible scheme is to translate the OpenMP program into a software DSM program, then execute it on the cluster. Evaluating the performance of OpenMP programs and analyzing their behavior will help support the OpenMP programming model on a software DSM cluster efficiently. In this paper, we use an experimental approach to investigate how the characteristics of the software DSM cluster and the translation together with the original program behavior determine the performance of OpenMP programs on software DSM clusters.
Shaogang Wu, Longbing Zhang, Zhiming Tang
CCGRID3