EDBT 2026 Demo / reviewers in the wild / expert
Sheng Liu 0001
dblp:03/5747-1
· DBLP profile ↗
17ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Ultra-Reliable Memories: A Practical Framework for Zero Mis-correction SEC-DED-DAEC Codes for Safety-Critical SystemsabstractIn safety-critical systems such as autonomous driving and aerospace, memory reliability standards are evolving from "high-reliability" to "ultra-reliability," demanding the eradication of all foreseeable, deterministic failure modes. To address the prevalent challenge of Double Adjacent Errors (DAE) induced from radiation, the design of SEC-DED-DAEC codes faces a critical dilemma: efficient but flawed Hsiao-based codes that risk miscorrection, versus correct-by-construction but costly and inflexible OLS-based codes. This trade-off between efficiency and correctness presents a key barrier to designing ultra-reliable systems. To resolve this impasse, this paper introduces MCTS-CDB, a novel Computer-Aided Design (CAD) framework. By integrating a CDCL-inspired search with Monte Carlo Tree Search (MCTS) guidance, it systematically constructs codes that achieve a zero-miscorrection guarantee within the highly-efficient Hsiao architecture. Experimental results validate our approach, showing that compared to a wide range of existing schemes, our generated codes achieve the correctness while reducing average encoding and decoding delays by 24.15% and 13.66%. This work provides a practical solution for designing the ultra-reliable memory subsystems required by next-generation safety-critical applications. Guixiang Chen, Sheng Liu 0001, Yang Guo 0003 |
DATE | 2 |
| 2026 | CACWS: Congestion-Aware Coordinated Warp Scheduler for Partitioned GPGPU
Sheng Liu 0001, Yang Guo 0003, Jianfeng Cui, Zekun Jiang |
IPDPS | 2 |
| 2025 | AICAWS: Arithmetic Intensity Based Cache-Conscious Adaptive Warp SchedulerabstractGeneral-Purpose Graphics Processing Units (GPGPUs) are crucial for parallel computing in artificial intelligence and big data with their performance heavily relying on efficient warp scheduling. Traditional schedulers, such as Round-Robin (RR) and Greedy-Then-Oldest (GTO), employ static strategies that struggle with adapting to diverse workloads, causing performance disparities across different applications. Prior work has focused on aspects like critical warps and memory access locality but has often overlooked the arithmetic intensity of workloads. Drawing inspiration from the Roofline model and recognizing that different workloads exhibit distinct computational intensities, we propose an Arithmetic Intensity based CacheConscious Adaptive Warp Scheduler (AICAWS). It operates by first analyzing the kernel's static arithmetic intensity through compiler, which serves as a baseline for the hardware. Subsequently, during warp execution, AICAWS dynamically monitors the warp's execution progress, analyzes its runtime arithmetic intensity, and adjusts warp scheduling strategies based on this. Furthermore, AICAWS considers cache locality during warp execution, enabling fine-grained classification of warps based on this. This synergistic mechanism enables AICAWS to effectively hide long-latency memory access operations. Evaluations on diverse benchmarks demonstrate that AICAWS achieves an average performance improvement of 26.3% compared to the baseline scheduler, with a peak improvement of 77.9%. Sheng Liu 0001, Zekun Jiang, Jianfeng Cui, Yang Guo 0003 |
ICCD | 2 |
| 2024 | LWECC: A Lightweight ECC Technology for HPC Accelerators Supporting Multi-granularity Memory AccessabstractHeterogeneous multi-core accelerators have exhibited remarkable computational performance, establishing their competitiveness within the realm of high-performance computing. The on-chip memory to support reliable and correct data access is a crucial element of the accelerator. With the continuous scaling down of semiconductor process nodes, memory units become more susceptible to soft errors caused by Single Event Upset (SEU), thus significantly increasing the risk of memory access errors. Additionally, data interaction between cores supports a variety of granularities, making it challenging to achieve error tolerant designs for multi-granular data with low hardware overhead. This paper proposes the Light-Weight ECC (LWECC) integrated in the 64-bit accelerator’s Scalar Memory (SM) to improve the efficiency of data protection for multi-granularity memory access. LWECC adopts a software-hardware co-design approach, effectively balancing area overhead and providing comprehensive and flexible data protection. Compared to the finest-grained ECC solution as a baseline, using LWECC can reduce redundant memory area by 80.90% and ECC circuit area by 16.87%, markedly reducing the whole hardware overhead of the SM. In addition, LWECC efficiently caters to diverse application domains, supporting not only high-performance computing applications but also accommodating fine-grained memory access applications with low operational latency. Lanting Guo, Chen Li 0015, Sheng Liu 0001 |
ISCAS | 4 |
| 2023 | UNCER: A framework for uncertainty estimation and reduction in neural decoding of EEG signals
Tiehang Duan, Zhenyi Wang 0001, Sheng Liu 0001, Yiyi Yin, Sargur N. Srihari |
Neurocomputing | 3 |
| 2022 | Mentha: Enabling Sparse-Packing Computation on Systolic ArraysabstractGeneralized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network. Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001 |
ICPP | 8 |
| 2022 | Adaptive Low-Cost Loop Expansion for Modulo Scheduling
Hongli Zhong, Zhong Liu 0003, Sheng Liu 0001, Sheng Ma, Chen Li 0015 |
NPC | 3 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 5 |
| 2021 | Sparse Matrix-Vector Multiplication Cache Performance Evaluation and Design ExplorationabstractIn this paper, we conducted a group of evaluations on the SpMV kernel with sequential implementation to investigate cache performance on single-core platforms. We verified a similar pattern inside a suite of sparse matrices covering various domains, which makes cache hit rate extraordinary inspiring in a sequential environment. This implicit regularity drove us to propose a cache space splitting approach, aiming at a better locality in dense vector accessing and utilization of large cache capacity in modern processors. Finally, we explored the design space of cache on Matrix 3000 GPDSP and proposed a group of cache parameters, based on our experimental results. Jianfeng Cui, Kai Lu 0001, Sheng Liu 0001 |
MASCOTS | 3 |
| 2021 | Advancing DSP into HPC, AI, and beyond: challenges, mechanisms, and future directions
Chen Li 0015, Chang Liu 0019, Sheng Liu 0001, Yuanwu Lei, Jian Zhang 0022, Yang Guo 0003 |
CCF Trans. High Perform. Comput. | 4 |
| 2018 | Conflict-Free Block-with-Stride Access of 2D Storage Structure
Guozhao Zeng, Sheng Liu 0001 |
ICA3PP (3) | 3 |
| 2017 | Modeling and evaluation for gather/scatter operations in Vector-SIMD architecturesabstractGather/scatter are state of the art vector memory access modes in Vector-SIMD architectures. However, because of the stochastic and complicated properties, the hardware design of gather/scatter operations lacks theoretical analysis and modeling. This paper proposes a model for gather/scatter operations on local vector memory for the first time. The model can not only give all the possible distributions of access locations, calculate the probability of access conflicts and predict the number of access conflicts, but also can provide the theoretical guidance for the performance optimization. This model is validated through experiments which can guide users to more specifically design and optimize the implementation of gather/scatter operations. Hongbing Tan, Sheng Liu 0001 |
ASAP | 3 |
| 2017 | Round-trip DRAM Access Fairness in 3D NoC-based Many-core SystemsabstractIn 3D NoC-based many-core systems, DRAM accesses behave differently due to their different communication distances and the latency gap of different DRAM accesses becomes bigger as the network size increases, which leads to unfair DRAM access performance among different nodes. This phenomenon may lead to high latencies for some DRAM accesses that become the performance bottleneck of the system. The paper addresses the DRAM access fairness problem in 3D NoC-based many-core systems by narrowing the latency difference of DRAM accesses as well as reducing the maximum latency. Firstly, the latency of a round-trip DRAM access is modeled and the factors causing DRAM access latency difference are discussed in detail. Secondly, the DRAM access fairness is further quantitatively analyzed through experiments. Thirdly, we propose to predict the network latency of round-trip DRAM accesses and use the predicted round-trip DRAM access time as the basis to prioritize the DRAM accesses in DRAM interfaces so that the DRAM accesses with potential high latencies can be transferred as early and fast as possible, thus achieving fair DRAM access. Experiments with synthetic and application workloads validate that our approach can achieve fair DRAM access and outperform the traditional First-Come-First-Serve (FCFS) scheduling policy and the scheduling policies proposed by reference [7] and [24] in terms of maximum latency, Latency Standard Deviation (LSD)1 and speedup. In the experiments, the maximum improvement of the maximum latency, LSD, and speedup are 12.8%, 6.57%, and 8.3% respectively. Besides, our proposal brings very small extra hardware overhead (<0.6%) in comparison to the three counterparts. Zhonghai Lu, Sheng Liu 0001, Shuming Chen |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2011 | A Novel Highly Scalable Architecture with Partially Distributed Pipeline and Hardware/Software Instruction EncodingabstractThe partitioning resources like pipelines and register files among clusters is proved to be an effective way to improve performance and scalability. However, improvement in scalability is limited by traditional instruction encoding schemes that quickly run out of bits in fixed-length instruction words to encode multiple register operands. Meanwhile, clustered processors may come at a cost of performance degradation, the major cause of which is the limited data locality arising from the lack of available registers and functional units. This paper introduces a highly scalable clustered architecture (HiSCA) to improve the scalability and performance of clustered processors. The pipeline of HiSCA provides high performance through in-order issuing, out-of-order execution and parallel but in-order commitment, while releasing instruction issuing from the heavy burden of dynamic scheduling. The hardware/software instruction encoding scheme of HiSCA splits instruction stream into chains of instructions (packs), and provides common information of instructions in the same packs in dedicated instruction words, thus reducing the total amount of information encoded in the instructions within the packs. HiSCA scales efficiently to 32 clusters with 1024 general purpose registers. Experiment results show that, for a 4-cluster and 8-issue configuration, HiSCA can achieve a 4.6% improvement in frequency with minimal hardware overhead, and an average of 13.3% performance speedup at the cost of 1.9% overhead to code size, compared with a traditional clustered processor with nearly the same hardware complexity. Sheng Liu 0001, Shuming Chen |
NAS | 2 |
| 2011 | Supporting Efficient Memory Conflicts Reduction Using the DMA Cache Technique in Vector DSPsabstractThis paper presents a Vector DMA Cache (VDC)scheme between the DMA bus and the Vector Memory (VM) in vector DSPs. The VDC can effectively reduce the VM access counts from the DMA requests and decrease the VM access conflicts. The VDC is specially designed for the DMA and not for the CPU, so it has some unique techniques which differ from the traditional CPU cache scheme. The main techniques of the VDC include the separate read cache and write cache, full line auto-updating policy and software cache coherence. Experimental results show the single-port VM plus the VDC can make programs reach more than 95% execution efficiency with only 44.1% and 51.4% chip area and power cost, compared with the dual-port VM scheme. And the single-port VM plus the VDC can reduce the execution cycles of programs by 3.7%~ 21.5% with only additional 7.3% and 6.3% chip area and power cost, compared with the pure single-port VM scheme. Sheng Liu 0001, Shuming Chen, Shenggang Chen |
NAS | 1 |
| 2011 | Matrix Odd-Even Partition: A High Power-Efficient Solution to the Small Grain Data ShuffleabstractThe shuffle operation is one of the bottlenecks invector DSPs. The partitioning problem of the shuffle matrix will have a great effect on the design of the shuffle unit, when dealing with the small grain data shuffle using a smaller-sized crossbar. The traditional matrix block partitioning solution will bring much chip area, power consumption and critical path delay. This paper presents a new matrix partitioning solution: the odd-even partition, which has advantages in dealing with the data going into or out of the shuffle network. It can simplify the hardware design and improve the power efficacy of the shuffle unit in vector DSPs, compared with the matrix block partitioning solution. The principles and theorems of the odd-even partition are explained and proofed. The shuffle unit based on the odd-even partition is implemented. Some contrastive experimental results show that the odd-even partitioning solution can decrease the critical path delay by 12.2%~21.6%, reduce the hardware area by 20.9%~40.8% and cut down the power consumption by 23.5%~56.8%, compared with the matrix block partitioning solution. Sheng Liu 0001, Shuming Chen, Jianghua Wan |
NAS | 1 |
| 2010 | Mapping of H.264/AVC Encoder on a Hierarchical Chip Multicore DSP PlatformabstractThe emergence of large-scale chip multicore processors makes the on-chip parallel H.264/AVC encoder with high parallelism feasible. To reduce the data reload frequency, a hierarchical chip multi-core DSP platform with overall 64 DSP cores is designed to accommodate the computation/data-intensive H.264/AVC encoder. To increase parallelism, macro block level parallelism is exploited in this paper and wave front algorithm is utilized. Centralized shared memory in super nodes of this hierarchical DSP platform affords larger local space to hold the frequently used data and reduce bandwidth requirement. Subtask level parallelism within motion estimation, intra prediction and mode decision is further exploited to keep the DSP cores in a super node busy even only one macro block are assigned to a super node. Because of lack of available macro blocks in filling and emptying stages when encoding a frame, super nodes cannot be kept busy all the time and speedups of 13, 24, 26 and 49 are achieved for QCIF, SIF, CIF and HD sequences, respectively. To further improve the speedups and make best use of the processor resources, frame level parallelism should be exploited with carefully tuned memory allocation policy. Shenggang Chen, Shuming Chen, Huitao Gu, Yaming Yin, Shuwei Sun, Sheng Liu 0001 |
HPCC | 8 |