EDBT 2026 Demo / reviewers in the wild / expert
Heng Liao
dblp:49/6267
· DBLP profile ↗
10ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0002-5992-5000ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Libra: A Hybrid-Sparse Attention Accelerator Featuring Multi-Level Workload BalanceabstractTransformers have delivered exceptional performance and are widely used across various natural language processing (NLP) tasks, owing to their powerful attention mechanism. However, the high computational complexity and substantial memory usage pose significant challenges to inference efficiency. Numerous quantization and value-level sparsification methods have been proposed to overcome these challenges. Since higher sparsity leads to greater acceleration efficiency, leveraging both value-level and bit-level sparsity (hybrid sparsity) can effectively exploit the acceleration potential of the attention mechanism. However, increased sparsity exacerbates load imbalance across compute units, potentially limiting the extent of acceleration benefits. To fully exploit the acceleration potential of hybrid sparsity, we propose Libra, an attention accelerator developed through algorithm-hardware co-design. At the algorithm level, we design the bit-group-based algorithm consisting of filtered bit-group sparsification (FBS) and dynamic bit-group quantization (DBQ) to maximize the utilization of sparsity in attention. FBS imposes structured sparsity on weights, while DBQ introduces dynamic sparsification during the computation of activations. At the hardware level, we design task pool to achieve multi-level workload balance, effectively mitigating the load imbalance among compute units induced by hybrid sparsity. Additionally, different stages in DBQ can be executed in parallel, with each stage operating at distinct bit-widths. To support this, we design an adaptive bit-width architecture that enables simultaneous computations at varying bitwidths. Our experiments demonstrate that, compared to state-of-the-art (SOTA) attention accelerators, Libra achieves up to $1.49 \times \sim 5.89 \times$ speedup and $2.65 \times \sim 10.82 \times$ enhancement in energy efficiency. Faxian Sun, Runzhou Zhang, Heng Liao, Zhinan Qin, Jianli Chen, Jun Yu 0010, Kun Wang 0005 |
DAC | 4 |
| 2025 | UB-mesh: An New Interconnection Technology for Large AI SuperNode
Heng Liao |
HCS | 1 |
| 2025 | RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingabstractMemory systems characterized by high bandwidth and/or capacity alongside high access latency are becoming increasingly critical.This trend can be observed both at the device level-for instance, in non-volatile memory-and at the system level, as seen in CXL-based memory pooling architectures.To benefit from such memory in general-purpose computing systems, it is essential to employ techniques that can tolerate high memory access latency.Although prefetching has long been recognized as a classical approach for latency tolerance, conventional prefetching techniques are typically either optimized for area efficiency or constrained by limited prefetching patterns.Consequently, they often fail to convert the abundant metadata into significant performance improvements at minimal cost.To address these challenges, we propose RICH-a prefetcher that strategically consumes memory capacity and bandwidth to reduce memory access latency.First, RICH is capable of leveraging abundant metadata to improve performance by integrating spatial prefetching with diverse region sizes and prefetch triggers.Second, RICH implements such metadata with minimal overheads by employing a hierarchical on-chip/off-chip storage mechanism, thereby avoiding both large on-chip storage and critical off-chip accesses.We propose a specific implementation of RICH and evaluate it across a wide range of workloads.With increased memory latency, RICH achieves performance improvements of 8.3% over Bingo and 6.2% over PMP.This highlights the RICH's suitability for future memory systems.In a conventional system, RICH still outperforms Bingo by 3.4%. Ningzhi Ai, Wenjian He, Hu He 0001, Heng Liao, Guowei Zhang 0002 |
MICRO | 5 |
| 2025 | The application of FCM-based computer image segmentation technology in agricultural production
Heng Liao, Huadong Huang |
Serv. Oriented Comput. Appl. | 1 |
| 2024 | MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected LayersabstractIn order to reduce the computational complexity of large language models, great efforts have been made to to improve the efficiency of transformer models such as linear attention and flash-attention. However, the model size and corresponding computational complexity are constantly scaled up in pursuit of higher performance. In this work, we present MemoryFormer, a novel transformer architecture which significantly reduces the computational complexity (FLOPs) from a new perspective. We eliminate nearly all the computations of the transformer model except for the necessary computation required by the multi-head attention operation. This is made possible by utilizing an alternative method for feature transformation to replace the linear projection of fully-connected layers. Specifically, we first construct a group of in-memory lookup tables that store a large amount of discrete vectors to replace the weight matrix used in linear projection. We then use a hash algorithm to retrieve a correlated subset of vectors dynamically based on the input embedding. The retrieved vectors combined together will form the output embedding, which provides an estimation of the result of matrix multiplication operation in a fully-connected layer. Compared to conducting matrix multiplication, retrieving data blocks from memory is a much cheaper operation which requires little computations. We train MemoryFormer from scratch and conduct extensive experiments on various benchmarks to demonstrate the effectiveness of the proposed model. Yehui Tang 0001, Haochen Qin, Zhenli Zhou, Chao Xu 0006, Kai Han 0002, Heng Liao, Yunhe Wang 0001 |
NeurIPS | 8 |
| 2021 | Ascend: a Scalable and Unified Architecture for Ubiquitous Deep Neural Network Computing : Industry Track PaperabstractDeep neural networks (DNNs) have been successfully applied to a great variety of applications, ranging from small IoT devices to large scale services in a data center. In order to improve the efficiency of processing these DNN models, dedicated hardware accelerators are required for all these scenarios. Theoretically, there exists an optimized acceleration architecture for each application. However, considering the cost of chip design and corresponding tool-chain development, researchers need to trade off between efficiency and generality. In this work, we demonstrate that it is practical to use a unified architecture, called Ascend, to support those applications, ranging from IoT devices to data-center services. We provide a lot of design details to explain that the success of Ascend relies on contributions from different levels. First, heterogeneous computing units are employed to support various DNN models. And the datapath is adapted according to the requirement of computing and data access. Second, when scaling the Ascend architecture from a single core to a cluster containing thousands of cores, it involves design efforts, such as memory hierarchy and system level integration. Third, a multi-tier compiler, which provides flexible choices for developers, is the last critical piece. Experimental results show that using accelerators based on the Ascend architecture can achieve comparable or even better performance in different applications. In addition, various chips based on the Ascend architecture have been successfully commercialized. More than 100 million chips have been used in real products. Heng Liao, Jiajin Tu, Xiping Zhou, Honghui Yuan, Yuxing Hu |
HPCA | 1 |
| 2019 | DaVinci: A Scalable Architecture for Neural Network ComputingabstractThis article consists of a collection of slides from the author's conference presentation. Heng Liao, Jiajin Tu, Xiping Zhou |
Hot Chips Symposium | 1 |
| 2006 | Parallel Switch System with QoS Guarantee for Real-Time Traffic
Wenjie Li 0002, Bin Liu 0001, Yang Xu 0010, Heng Liao |
J. Comput. Sci. Technol. | 4 |
| 1997 | Available Parallelism in Video ApplicationsabstractMost recent research in instruction-level parallelism has focused on general-purpose applications such as the SPEC benchmarks. Many quantitative experiments have been performed over the years measuring the impact of different execution models and optimization techniques on these applications. Researchers have been developing various ILP architectures for media processors in order to exploit parallelism in audio, video, and graphics applications. It has been assumed that these applications contain far more potential parallelism than general-purpose code, but there have been few attempts to quantify the available parallelism. We present a linear complexity global scheduling algorithm that can process very long traces up to 1 billion operations. Therefore, traces of video applications such as MPEG1, MPEG2, MPEG4 and H.263 encoders and decoders can be analyzed. Using an idealized execution model, speedups of over 1000 have been found in some applications. The experiment shows that eliminating currently identifiable bottlenecks can allow the exploitation of huge amounts of ILP in audio and video applications. Heng Liao, Andrew Wolfe |
MICRO | 1 |
| 1996 | DYNAMEM - A microarchitecture for improving memory disambiguation at run-time
Xianzhu Wang, Heng Liao, Sanli Li |
J. Comput. Sci. Technol. | 2 |