Liang Wang 0020

dblp:56/4499-20 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0002-6112-1928ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 7 first-author · 24 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A two-stage data placement strategy for cloud-edge-device collaborative environment
Runnan Shen, Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Shengyang Tan, Yuntong Li, Xiangrong Xu 0002, Liang Wang 0020
Comput. Commun.9
2026 Online detection of hardware Trojan enabled packet tampering attack on network-on-chip: A Bayesian approach
Xiaohang Wang 0001, Ge Cao, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Liang Wang 0020, Fen Guo
Integr.7
2026 Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-Design
abstract
Large language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference.In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59×, 1.38×, and 1.61×, along with energy efficiency improvements of 1.52×, 1.42×, and 1.22× over ANT, OliVe, and FineQ, respectively.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Tairan Zhang, Jinquan Wang, Yongyue Wang, Xiaojian Liao
IEEE Trans. Computers2
2025 CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
abstract
With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56 × (1.88 × on average).
Tianhao Cai, Liang Wang 0020, Limin Xiao 0001, Xiaojian Liao
DAC2
2025 FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
abstract
Large language models (LLMs) have significantly advanced the natural language processing paradigm but impose substantial demands on memory and computational resources. Quantization is one of the most effective ways to reduce memory consumption of LLMs. However, advanced single-precision quantization methods experience significant accuracy degradation when quantizing to ultra-low bits. Existing mixed-precision quantization methods are quantized by groups with coarse granularity. Employing high precision for group data leads to substantial memory overhead, whereas low precision severely impacts model accuracy. To address this issue, we propose FineQ, software-hardware co-design for low-bit fine-grained mixed-precision quantization of LLMs. First, FineQ partitions the weights into finer-grained clusters and considers the distribution of outliers within these clusters, thus achieving a balance between model accuracy and memory overhead. Then, we propose an outlier protection mechanism within clusters that uses 3 bits to represent outliers and introduce an encoding scheme for index and data concatenation to enable aligned memory access. Finally, we introduce an accelerator utilizing temporal coding that effectively supports the quantization algorithm while simplifying the multipliers in the systolic array. FineQ achieves higher model accuracy compared to the SOTA mixed-precision quantization algorithm at a close average bit-width. Meanwhile, the accelerator achieves up to 1.79x energy efficiency and reduces the area of the systolic array by 61.2%.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Xiangrong Xu 0002
DATE2
2025 Swift-Sim: A Modular and Hybrid GPU Architecture Simulation Framework
abstract
Simulation tools are critical for architects to quickly estimate the impact of aggressive new features of GPU architecture. Existing cycle-accurate GPU simulators are typically cumbersome and slow to run. We observe that it is time-consuming and unnecessary for cycle-accurate GPU simulators to perform detailed simulations for the entire GPU when exploring the design space of specific components. This paper proposes Swift-Sim, a modular and hybrid GPU simulation framework. With a highly modular design, our framework can choose appropriate modeling approaches for each component according to requirements. For components of interest to architects, we use cycle-accurate simulation to evaluate new GPU architectures. For other components, we use analytical modeling, which accelerates simulation speed with only minor and acceptable degradation in overall accuracy. Based on this simulation framework, we present two working examples of hybrid modeling that simulate the ALU pipeline and memory accesses using analytical models. We further implement two GPU performance simulators with different levels of simplification based on Swift-Sim and evaluate them using configurations from real GPUs. The results show that the two simulators achieve an 82.6x and 211.2x geometric mean speedup compared to Accel-Sim with insignificant accuracy degradation,
Xiangrong Xu 0002, Yuanqiu Lv, Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang
DATE3
2025 PointISA: ISA-Extensions for Efficient Point Cloud Analytics via Architecture and Algorithm Co-Design
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu
MICRO2
2025 Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
abstract
The quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao
MICRO2
2025 ICCG: low-cost and efficient consistency with adaptive synchronization for metadata replication
Liang Wang 0020, Jing Shang 0001, Zhiwen Xiao, Limin Xiao 0001, Bing Wei 0002, Runnan Shen, Jinquan Wang
Frontiers Comput. Sci.2
2025 Exploiting intra-chip locality for multi-chip GPUs via two-level shared L1 cache
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Xiaojian Liao
J. Syst. Archit.2
2024 FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
abstract
Point cloud analytics has become a critical workload for embedded and mobile platforms across various applications. Farthest point sampling (FPS) is a fundamental and widely used kernel in point cloud processing. However, the heavy external memory access makes FPS a performance bottleneck for real-time point cloud processing. Although bucket-based farthest point sampling can significantly reduce unnecessary memory accesses during the point sampling stage, the KD-tree construction stage becomes the predominant contributor to execution time. In this paper, we present FuseFPS, an architecture and algorithm co-design for bucket-based farthest point sampling. We first propose a hardware-friendly sampling-driven KD-tree construction algorithm. The algorithm fuses the KD-tree construction stage into the point sampling stage, further reducing memory accesses. Then, we design an efficient accelerator for bucket-based point sampling. The accelerator can offload the entire bucket-based FPS kernel at a low hardware cost. Finally, we evaluate our approach on various point cloud datasets. The detailed experiments show that compared to the state-of-the-art accelerator QuickFPS, FuseFPS achieves about $4.3\times $ and about $6.1\times $ improvements on speed and power efficiency, respectively.
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie
ASPDAC2
2024 BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds
abstract
Point cloud-based machine perception applications have achieved great success in various scenarios. In this work, we focus on point cloud k-Nearest Neighbor (kNN) search, an important kernel for point clouds. Existing kNN acceleration techniques have overlooked the operation-level optimization in the Euclidean distance computation operations, which suffer from low efficiency due to a number of unnecessary computations and various data precision requirements.We reconsider point cloud kNN search from a new bitserial computation perspective and propose BitNN, a bit-serial architecture for point cloud kNN search. BitNN supports adaptive precision processing and unnecessary computing reduction, significantly improving the performance and power efficiency of kNN search. To achieve that, we first propose a bit-serial computation method for kNN search, which derives a recursive expression to compute the Euclidean distance bit by bit. Then, the dimension-wise point cloud encoding method and point-wise data layout method are proposed to enable adaptive precision processing based on bit-serial computation. Furthermore, we present an early termination mechanism for bit-serial kNN search. By estimating the lower bound of distance based on a few bits, a number of unnecessary computations can be reduced. Finally, we design an efficient bit-serial accelerator for kNN search. The accelerator exploits the massive parallelism to improve computing efficiency.We evaluate BitNN with several widely used point cloud datasets. BitNN achieves up to $6.6 \times$ speedup and $3.6 \times$ power efficiency compared to a comparable sized architecture. Moreover, BitNN can be easily integrated into existing bit-parallel kNN accelerators. We enhance the state-of-the-art kNN accelerator, ParallelNN, with bit-serial computation techniques, achieving up to $4.4 \times$ speedup and $2.9 \times$ power efficiency
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Tianhao Cai, Xiangrong Xu 0002
ISCA2
2024 Minimizing the cost of periodically replicated systems via model and quantitative analysis
Liang Wang 0020, Limin Xiao 0001, Shixuan Jiang, Jinquan Wang, Bing Wei 0002, Guangjun Qin
Frontiers Comput. Sci.2
2024 ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
abstract
The systolic accelerator is one of the premier architectural choices for DNN acceleration. However, the conventional systolic architecture suffers from low PE utilization due to the mismatch between the fixed array and diverse DNN workloads. Recent studies have proposed flexible systolic array architectures to adapt to DNN models. However, these designs support only coarse-grained reshaping or significantly increase hardware overhead. In this study, we propose ReDas, a flexible and lightweight systolic array that supports dynamic fine-grained reshaping and multiple dataflows. First, ReDas integrates lightweight and reconfigurable roundabout data paths, which achieve fine-grained reshaping using only short connections between adjacent PEs. Second, we redesign the PE microarchitecture and integrate a set of multi-mode data buffers around the array. The PE structure enables additional data bypassing and flexible data switching. Simultaneously, the multi-mode buffers facilitate fine-grained reallocation of on-chip memory resources, adapting to various dataflow requirements. ReDas can dynamically reconfigure to up to 129 different logical shapes and 3 dataflows for a 128 × 128 array. Finally, we propose an efficient mapper to generate appropriate configurations for each layer of DNN workloads. Compared to the conventional systolic array, ReDas can achieve about 4.6× speedup and 8.3× energy-delay product (EDP) reduction.
Liang Wang 0020, Limin Xiao 0001, Tianhao Cai, Xiangrong Xu 0002
IEEE Trans. Computers2
2024 ATA-Cache: Contention Mitigation for GPU Shared L1 Cache With Aggregated Tag Array
abstract
To fully exploit the locality of GPU applications, the GPU shared L1 cache architecture, which shares L1 cache among multiple GPU cores, is a promising architecture while still suffering from high resource contentions. We present a GPU shared L1 cache architecture with an aggregated tag array that minimizes the L1 cache contentions and takes full advantage of inter-core locality. The key idea is to decouple and aggregate the tag arrays of multiple L1 caches so that the cache requests can be compared with all tag arrays in parallel to probe the replicated data in other caches. The GPU caches are only accessed by other GPU cores when replicated data exists, filtering out unnecessary cache accesses that cause high resource contentions. We also develop a two-level thread-block scheduling policy adapted for the shared L1 cache architecture to maximize the available locality. The experimental results show that GPU performance can be improved by 14.5% on average for applications with a high inter-core locality.
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Hao Liu 0107
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Accelerating zk-SNARK with Group and Zone Optimization on GPU
abstract
Zero-knowledge proof (ZKP) is a popular cryptographic strategy for building a trusted environment, which can be applied to blockchain, electronic voting, and other scenarios. However, ZKP involves a number of computationally intensive operations that limit its widespread adoption in time-sensitive practical applications. The multi-scalar multiplication (MSM) dominates the computations and takes over 70% of the total computation time. This paper proposes a GPU-based acceleration method for ZKP by designing several optimization techniques for MSM. First, this paper constructs a formal mathematical formula of the Pippenger algorithm, which provides a theoretical optimization framework for MSM. Second, by parallelizing the prefix sum, the time complexity of the bucket reduction part of MSM is reduced from $\mathcal{O}\left( {3 \times {2^C}} \right)$ to $\mathcal{O}\left( {2 \times {2^C}} \right)$. Finally, this paper also analyzes the influence of group size on the final calculation time under different data scales and gives a suitable range of group sizes. Compared to the state-of-the-art method, our method can achieve 1.01× to 1.12× for throughput.
Runnan Shen, Liang Wang 0020, Haotian Luo, Jinqian Yang, Jinquan Wang, Qiancheng Sun, Limin Xiao 0001
ICPADS2
2023 Shogun: A Task Scheduling Framework for Graph Mining Accelerators
abstract
Graph mining is an emerging application of great importance to big data analytic. Graph mining algorithms are bottle-necked by both computation complexity and memory access, hence necessitating specialized hardware accelerators to improve the processing efficiency. Current accelerators have extensively exploited task-level and fine-grained parallelism in these algorithms. However, their task scheduling still has room for optimization. They use either breadth-first search, depth-first search or a combination of both, leading to either poor intermediate data locality, low parallelism or inter-depth barriers.
Jianfeng Zhu 0001, Wenrui Wei, Longlong Chen, Liang Wang 0020, Shaojun Wei, Leibo Liu
ISCA5
2023 CFIO: A conflict-free I/O mechanism to fully exploit internal parallelism for Open-Channel SSDs
Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin
J. Syst. Archit.2
2023 QuickFPS: Architecture and Algorithm Co-Design for Farthest Point Sampling in Large-Scale Point Clouds
abstract
Point clouds have been employed extensively in machine perception applications. Farthest point sampling (FPS) is a critical kernel for point cloud processing. With the rapid growth of point cloud scale, FPS introduces a large number of memory accesses, which become the bottleneck of the large-scale point cloud processing. In this article, we present QuickFPS, an architecture and algorithm co-design of FPS in large-scale point clouds. First, we systemically analyze the characteristics of FPS and put forward a bucket-based FPS algorithm. The algorithm introduces a two-level tree data structure to organize the large-scale point cloud into multiple buckets. By using two mechanisms named merged computation and implicit computation for the buckets, the external memory accesses and compute cost are significantly reduced. Then, we design an efficient domain-specific accelerator for FPS in large-scale point clouds. The accelerator takes advantage of different forms of parallelism and further improves the accelerator’s efficiency. Finally, we evaluate QuickFPS with several widely used point cloud datasets, which include small-scale and large-scale point clouds (up to 120 000 points). Overall, QuickFPS achieves performance speedups of$43.4\times$and$12.2\times$compared to GTX 1080Ti GPU and state-of-the-art point cloud accelerator PointAcc, respectively.
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xiangrong Xu 0002, Jianfeng Zhu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 EBIO: An Efficient Block I/O Stack for NVMe SSDs With Mixed Workloads
abstract
With the advent of high-performance nonvolatile memory express (NVMe) SSD, the overhead caused by the storage software stack becomes a significant bottleneck for exploiting the potential of NVMe SSD. Recent I/O isolation approaches eliminate CPU switching, I/O interference, and lock contention for I/O queues by pinning I/O threads in isolated and dedicated I/O paths. However, they degrade the overall performance of mixed workloads with heterogeneous I/O demands. The I/O-intensive workloads issue multiple I/O requests and quickly fill up their dedicated I/O queues, resulting in I/O wait. On the contrary, the non-I/O-intensive workloads cannot deliver enough I/O requests to saturate their associated I/O queues. Moreover, frequent allocations and deallocations of I/O request objects expose a significant overhead for I/O-intensive workloads. In this article, we propose EBIO, an efficient block I/O stack for NVMe SSD, to improve the overall performance of mixed workloads. EBIO contains an on-demand queue management strategy (ODQM) and a reusable I/O management strategy (RERM). Specifically, ODQM eliminates I/O wait by dynamically adding queues for I/O-intensive workloads and leverages I/O generation time to guarantee strong sequential consistency and fairness in scheduling I/O requests. RERM reduces the overhead caused by repeatedly allocating I/O request objects for I/O-intensive workloads by reusing the allocated objects. Experimental results show that, compared to the state-of-the-art approaches, EBIO improves input/output operations per second by up to 16.44% and reduces I/O latency by up to 32.76%.
Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Dynamic two-side matching of tasks and resources in wide-area distributed computing environments
Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang
J. Supercomput.2
2022 Upward Packet Popup for Deadlock Freedom in Modular Chiplet-Based Systems
abstract
Monolithic SoCs can be decomposed into disparate chiplets that support integration with advanced pack-aging technologies. This concept is promising in reducing the manufacturing cost of large scale SoCs due to the higher yield rate and reusability of chiplets. The chiplets should be designed in a modular manner without holistic system knowledge so that they can be reused in different SoCs. However, the design modularity is a major challenge to the networks-on-chip (NoCs) of chiplets.New deadlocks may occur across both the chiplets and the interposer due to the integration, even if the NoC of each individually designed chiplet is deadlock free. However, conventional deadlock freedom approaches are unsuitable to handle such deadlocks because they require holistic knowledge and violate the modularity. Although there are several modular approaches that specifically target at integration-induced deadlocks, their routing is overly restricted and the injection control incurs additional latency. They also lack flexibility in dynamically changing topologies due to their complex software algorithm and the hard-wired components.In this paper, a key insight on the chiplet integration-induced deadlocks is gained, inspired by which a deadlock recovery framework (named UPP) is proposed. Specifically, it is verified that an integration-induced deadlock always involves a stalled upward packet moving from the interposer to the connected chiplet via the vertical link. Thus, UPP detects a deadlock by discovering the upward packet and recovers the system from deadlock by transmitting the upward packet to its destination. Hybrid flow control mechanisms are proposed to enable the upward packet to bypass the buffers and be transmitted via the normal router datapath. To guarantee the ejection of the upward packet after transmission, a lightweight protocol is proposed to reserve ejection queue entries of the network interface. Experimental results show that while adhering to design modularity, UPP provides an average runtime speedup of 3.1%∼10.3% with an area overhead of less than 4%.
Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Jianfeng Zhu 0001, Honglan Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu
HPCA2
2022 Hypergraph-partitioning-based online joint scheduling of tasks and data
Liang Wang 0020, Limin Xiao 0001, Wei Wei 0006, Rafal Scherer, Guangjun Qin, Jinquan Wang
J. Supercomput.2
2021 UPM-DMA: An Efficient Userspace DMA-Pinned Memory Management Strategy for NVMe SSDs
Jinbin Zhu, Limin Xiao 0001, Liang Wang 0020, Guangjun Qin, Zhonglin Liu
ICA3PP (1)3
2021 An enhanced planned obsolescence attack by aging networks-on-chip
Yinyuan Zhao, Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
J. Syst. Archit.4
2021 On Performance Optimization and Quality Control for Approximate-Communication-Enabled Networks-on-Chip
abstract
For many applications showing error forgiveness, approximate computing is a new design paradigm that trades application output accuracy for mitigating computation/communication effort, which results in performance/energy benefit. Since networks-on-chip (NoCs) are one of the major contributors to system performance and power consumption, the underlying communication is approximated to achieve time/energy improvement. However, performing approximation blindly causes unacceptable quality loss. In this article, first, an optimization problem to maximize NoC performance is formulated with the constraint of application quality requirement, and the application quality loss is studied. Second, a congestion-aware quality control method is proposed to improve system performance by aggressively dropping network data, which is based on flow prediction and a lightweight heuristic. In the experiments, two recent approximation methods for NoCs are augmented with our proposed control method to compare with their original ones. Experimental results show that our proposed method can speed up execution by as much as 29.42% over the two state-of-the-art works.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Liang Wang 0020, Terrence S. T. Mak
IEEE Trans. Computers5
2021 A Deflection-Based Deadlock Recovery Framework to Achieve High Throughput for Faulty NoCs
abstract
Deadlock is a critical issue in faulty Networks-on-Chips (NoCs). Existing deadlock-free approaches on faulty NoCs suffer from low throughput and poor fairness when the network becomes oversaturated. This problem hinders their practical use as oversaturation scenarios are more frequent on faulty NoCs. To address this issue, a deflection-based deadlock recovery framework is proposed for higher oversaturation performance on faulty NoCs. First, we observe the low oversaturation performance of existing deadlock recovery approaches, and analyze the positive feedback loop that can amplify the negative impact of deadlocks and congestions, which necessitate handling both deadlocks and congestions in a deadlock recovery framework. Second, we propose a novel deadlock recovery framework, which includes an accurate, timely deadlock detection and a highly efficient deadlock recovery. Both the deadlock detection and recovery reduce the average packet traversal latency, thereby improving the average oversaturation throughput. Third, we propose a distributed implementation to make the entire network enter and exit the deflection mode, which is conducted by broadcasting special messages via a bufferless subnetwork. An average oversaturation throughput improvement of 1.1 ~ 8.1× over state-of-the-art approaches is achieved. In terms of fairness, the minimal oversaturation throughput is improved from near zero to half of the peak throughput.
Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Shouyi Yin, Shaojun Wei, Leibo Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology
abstract
Future NoCs should be highly flexible to adapt to communication demands to achieve high scalability and low power consumption. However, the flexibility is still quite limited by the high complexity of reconfiguration for globally reconfigured channels. In this paper, we propose to augment a router-based buffered NoC with a reconfigurable ring architecture by exploiting cycle decomposition of a torus bufferless network. At runtime, the topologies of the rings can be reconfigured according to the workloads by choosing different cycle decompositions of the torus network. Because the shapes of the rings are restricted to a specified regular shape, the reconfiguration time can be reduced to a linear complexity with respect to network size, and the reconfiguration algorithm can be implemented in a distributed hardware. The experimental results show that the reconfigurable rings provide 54% and 26% improvements on packet latency and static power saving, respectively, for realistic workloads.
Liang Wang 0020, Leibo Liu, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei
DAC1
2020 On hardware-trojan-assisted power budgeting system attack targeting many core systems
Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Mei Yang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
J. Syst. Archit.4
2020 Aggressive Fine-Grained Power Gating of NoC Buffers
abstract
Power gating is effective for networks-on-chip (NoCs) to reduce the excessive leakage power dissipated by idle network components. Most existing NoC power-gating approaches rely on the routing algorithms to mitigate the power-gating blocking latency problem. When the network becomes faulty and fault-tolerant routing algorithms are applied, these approaches are no longer applicable or can seriously degrade the performance. Other approaches propose fine-grained buffer power gating, but they are too conservative in power saving due to the buffer backpressure flow control. To address these problems, we propose an aggressive fine-grained power gating of flit-sized buffer entries by adopting backpressureless flow control in an input-buffered network. The power-gating decisions are made based on the flit deflection rate. However, directly applying the backpressureless flow control leads to the difficulties of multiflit packet truncation and protocol deadlocks. Therefore, we modify the packet injection architecture to avoid packet truncation. This is done by chaining the local input port with a randomly chosen input port. Finally, we design a progressive recovery framework to handle both livelocks and protocol deadlocks. It does not need to truncate packets or strictly separate different message classes when the network is free of livelocks or protocol deadlocks. The experimental results show that with a hardware overhead of 9.6%, our design can save up to 59% network power consumption in both a fault-free and a faulty NoC with little zero-load latency penalty. Our design also approaches an ideal energy-proportional NoC because it can constantly reduce power consumption over a wide range of injection rates.
Leibo Liu, Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Achieving Flexible Global Reconfiguration in NoCs Using Reconfigurable Rings
abstract
The communication behaviors in NoCs of chip-multiprocessors exhibit great spatial and temporal variations, which introduce significant challenges for the reconfiguration in NoCs. Existing reconfigurable NoCs are still far from ideal reconfiguration scenarios, in which globally reconfigurable interconnects can be immediately reconfigured to provide bandwidths on demand for varying traffic flows. In this paper, we propose a hybrid NoC architecture that globally reconfigures the ring-based interconnect to adapt to the varying traffic flows with a high flexibility. The ring-based interconnect has the following advantages. First, it includes horizontal rings and vertical rings, which can be dynamically combined or split to provide low-latency channels for heavy traffic flows. Second, each combined ring connects a number of nodes, thereby improving both the utilization of each ring and the probability to reuse previous reconfigurable interconnects. Finally, the reconfiguration algorithm has a linear-time complexity and can be implemented using a low-overhead hardware design, making it possible to achieve a fast reconfiguration in NoCs. The experimental results show that compared to recent reconfigurable NoCs, the proposed NoC architecture can greatly improve the saturation throughput for synthetic traffic patterns, and reduce the packet latency over 40 percent for realistic benchmarks without incurring significant area and power overhead.
Liang Wang 0020, Leibo Liu, Jie Han 0001, Xiaohang Wang 0001, Shouyi Yin, Shaojun Wei
IEEE Trans. Parallel Distributed Syst.1
2019 A Lifetime Reliability-Constrained Runtime Mapping for Throughput Optimization in Many-Core Systems
abstract
Due to technology scaling, lifetime reliability is becoming one of the major design constraints in the performance optimization of future many-core systems. Given a lifetime reliability constraint, the existing lifetime-constrained runtime mapping schemes often lead to low throughput because of the requirement to map all applications to compact regions. In this paper, we propose a runtime application mapping scheme that exploits a borrowing strategy to improve the throughput of many-core systems given a lifetime constraint. First, we propose using different strategies for mapping communication-intensive applications and computation-intensive applications. The lifetime reliability constraint can be relaxed in the local time scale when the communication requirement is high. The throughput is improved because the communication distance of communication-intensive applications is optimized while the waiting time of computation-intensive application is reduced. Then, we propose a method to effectively classify applications depending on the communication-to-computation ratio. A dynamic threshold is determined according to the current locations of available cores. Finally, we propose an improved neighborhood allocation scheme to reduce the communication cost in the task mapping. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed mapping scheme improves the throughput of many-core systems by 26% on average for synthetic task graphs and by 20% on average for realistic task graphs while the lifetime reliability is maintained within a constraint.
Liang Wang 0020, Ping Lv, Leibo Liu, Jie Han 0001, Ho-fung Leung, Xiaohang Wang 0001, Shouyi Yin, Shaojun Wei, Terrence S. T. Mak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 A Non-Minimal Routing Algorithm for Aging Mitigation in 2D-Mesh NoCs
abstract
Due to technology scaling, aging issue is becoming one of major concerns in the design of network-on-chip (NoC). The imbalanced workload distribution and routing algorithm cause aging hotspots, where a certain group of routers have higher aging effect than others. This can possibly lead to shorter lifetime of NoC. Most existing aging-aware routing algorithms are based on minimal routing, which suffers from less degree of adaptiveness compared to non-minimal routing. Thus, they are inefficient to mitigate the aging effect of routers. In this paper, we propose to use a non-minimal routing scheme to detour the traffic away from the aging hotspots, with the objective of mitigating the aging effect for NoCs. The problem is formulated as a bottleneck shortest path problem and solved using a dynamic programming approach. Finally, the experimental results show that compared to the state-of-the-art aging-aware routing algorithm, the non-minimal routing algorithm has up to 20% lifetime improvement for hotspot traffic patterns and realistic workload traces.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Runtime task mapping for lifetime budgeting in many-core systems
abstract
Due to technology scaling, lifetime reliability is becoming one of major design constraints in the design of future many-core systems. In this paper, we propose a novel runtime mapping scheme which can dynamically map the applications given a lifetime reliability constraint. A borrowing strategy is adopted to manage the lifetime in a long-term scale, and the lifetime constraint can be relaxed in short-term scale when the communication performance requirement is high. The through-put can be improved because the communication performance of communication intensive applications is optimized, and mean-while the waiting time of computation intensive application is reduced. An improved neighborhood allocation method is proposed for the runtime mapping scheme. Moreover, we propose a method to effectively classify communication intensive applications and computation intensive applications. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed scheme has more than 20% throughput improvement in average.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
FDL1
2017 Throughput Optimization for Lifetime Budgeting in Many-Core Systems
abstract
Due to technology scaling, lifetime reliability is becoming one of major design constraints in the design of future many-core systems. In this paper, we propose a novel runtime mapping scheme which could dynamically map the applications given a lifetime reliability constraint. A borrowing strategy is adopted to manage the lifetime in a long-term scale, and the lifetime constraint could be relaxed in short-term scale when the communication performance requirement is high. The throughput could be improved because the communication performance of communication intensive applications is optimized, and meanwhile the waiting time of computation intensive application is reduced. Furthermore, an improved neighborhood allocation method is proposed for the runtime mapping scheme. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed mapping scheme could have over 20% throughput improvement.
Liang Wang 0020, Xiaohang Wang 0001, Ho-fung Leung, Terrence S. T. Mak
ACM Great Lakes Symposium on VLSI1
2016 SlideAcross: A Low-Latency Adaptive Router for Chip Multi-processor
abstract
The non-uniform distributed traffic of chip multi-processor (CMP) demands an on-chip communication infrastructure which is able to avoid congestion under high traffic conditions while possessing minimal pipeline delay at low load conditions. In this paper, we propose a low-latency adaptive router with a low-complexity single-cycle bypassing mechanism to meet the communication needs of CMPs. At low loads, this router transmits a flit using dimension-ordered routing (DoR) in the bypass datapath. When the output port required intra-dimension bypassing is not available, the packet is routed adaptively to avoid congestion. The router also has a simplified virtual channel allocation (VA) scheme that yields a non-speculative low-latency pipeline. By combining the low-complexity bypassing technique together with adaptive routing, the proposed router architecture can achieve low-latency communication under various traffic loads. Simulation shows that proposed router can reduce applications' execution time by 16.9% in average compared to low-latency router SWIFT.
Wen Zong, Liang Wang 0020, Qiang Xu 0001, Michael Opoku Agyeman
DSD2
2016 Adaptive Routing Algorithms for Lifetime Reliability Optimization in Network-on-Chip
abstract
Technology scaling leads to the reliability issue as a primary concern in Network-on-Chip (NoC) design. We observe that due to routing algorithm some routers age much faster than others which becomes a bottleneck for NoC lifetime. In this paper, lifetime is modeled as a resource consumed over time. A metric lifetime budget is associated with each router, indicating the maximum allowed workload for current period. Since the heterogeneity in router lifetime reliability has strong correlation with the routing algorithm, we define a problem to optimize the lifetime by routing packets along the path with maximum lifetime budgets. The problem is then extended for both performance and lifetime reliability optimization. The lifetime is optimized in long-term time scale while performance is optimized in short-term time scale. Two dynamic programming-based adaptive routing algorithms (lifetime aware routing and multi-objective routing) are proposed to solve the problems. In the experiments, the lifetime aware routing and multi-objective routing algorithms are evaluated with synthetic traffic and real benchmarks respectively. The experimental results show that the lifetime aware routing has around 20, 45 and 55 percent minimal lifetime improvement than XY routing, NoP routing and Oddeven routing, respectively. In addition, the multi-objective adaptive routing algorithm can effectively improve both performance and lifetime.
Liang Wang 0020, Xiaohang Wang 0001, Terrence S. T. Mak
IEEE Trans. Computers1
2014 Dynamic programming-based lifetime aware adaptive routing algorithm for Network-on-Chip
abstract
Technology scaling leads to the reliability issue as a primary concern in Network-on-Chip (NoC) design. Due to the routing algorithms, some routers may age much faster than others, which becomes a bottleneck for system lifetime. In this paper, lifetime is modeled as a resource consumed over time. A metric lifetime budget is associated with each router, indicating the maximum allowed workload for current period. Since the heterogeneity in router lifetime reliability has strong correlation with the routing algorithm, we define a problem to optimize the lifetime by routing flits along the path with maximum lifetime budgets. A dynamic programming-based lifetime aware routing algorithm is proposed based on the lifetime budget metric. The dynamic programming network approach is employed to solve this problem with linear complexity. The experimental results show that the lifetime aware routing has around 20%, 45%, 55% minimal MTTF improvement than XY routing, NoP routing, oddeven routing, respectively.
Liang Wang 0020, Xiaohang Wang 0001, Terrence S. T. Mak
VLSI-SoC1