VLDB 2026 Research / reviewers in the wild / expert
Ling Yang 0008
dblp:01/24-8
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-0838-0066ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vector Value Prediction with Element-wise Stride CompressionabstractThe increasing emphasis on vectorization and Single Instruction, Multiple Data (SIMD) processing reflects their central role in modern processors. However, as workloads in data processing, multimedia, and algorithmic operations grow in complexity, they introduce more pronounced data dependencies, leading to longer execution times compared to scalar instructions. To address these evolving challenges, we present the Vector Value TAGE predictor (VVTAGE), a novel value predictor specifically designed for vector instructions. Although value prediction has been proposed as a fundamental strategy to enhance processor performance, it has traditionally focused on predicting 64-bit scalar values to mitigate data dependencies and improve pipeline throughput. VVTAGE extends the prediction capabilities of existing scalar predictors to accommodate the wide vector registers used in contemporary processors. Our research demonstrates that VVTAGE can significantly improve processor performance, achieving performance gains of up to 20.1% and an average increase of 4.53% in the evaluated SIMD benchmarks. This innovative approach and surprising results represent a significant advancement in optimizing the performance of SIMD processors. Furthermore, to enhance the scalability of VVTAGE, we propose an element-wise stride compression method to reduce its storage overhead. Experimental results show that VVTAGE still retains 64% performance gain while reducing 15.4KB overhead. Yanmeng Huang, Ling Yang 0008, Yuanhu Cheng, Quan Deng 0003, Junbo Tie, Yongwen Wang, Hai Zhong, Libo Huang 0002 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei |
ISCA | 1 |
| 2025 | Late Breaking Results: AFS: Improving Accuracy of Quantized Mamba via Aggressive Forgetting StrategyabstractMamba overcomes the quadratic complexity problem inherent in Transformer models while maintaining comparable contextual modeling capabilities. However, Mamba-based foundation models encounter challenges in achieving efficient inference on resource-constrained devices, primarily due to their considerable size. Model compression techniques, such as linear quantization, offer a viable solution to this problem. Nevertheless, the introduction of significant outliers during Mamba's past-state forgetting process can lead to a notable decrease in accuracy when employing linear quantization. To overcome these challenges, this paper introduces the Aggressive Forgetting Strategy (AFS), an innovative and efficient algorithm designed to mitigate the quantization issues caused by outliers in the state forgetting mechanism. AFS incorporates a computation-free approach for handling outliers, facilitating both efficient and accurate linear quantization for Mamba, which is essential for applications in resource-constrained scenarios. By leveraging the AFS strategy, Mamba can perform more efficient inference, while significantly improving the accuracy by up to 21.5× compared to conventional methods. Zhouquan Liu, Libo Huang 0002, Ling Yang 0008, Gang Chen 0023, Yongwen Wang |
DATE | 3 |
| 2025 | SONet: Towards Practical Online Neural Network for Enhancing Hard-to-Predict Branches
Zhenxuan Xiong, Libo Huang 0002, Ling Yang 0008, Hui Guo 0004, Songwen Pei, Gang Chen 0023, Yongwen Wang |
Euro-Par (2) | 3 |
| 2025 | X-SA: An Efficient Configurable Systolic Array Computing Architecture for GPGPUabstractGPGPUs are pivotal for edge AI, but resource constraints demand efficient low-precision computation. Conventional GPGPUs face challenges in resource utilization, particularly with irregular matrices common in AI, and memory bandwidth limitations on edge devices. Traditional fixed-size systolic arrays often suffer from underutilization under varying workloads. This paper introduces X-SA, a configurable systolic array architecture tailored for INT8 matrix multiplication on GPGPUs in resource-constrained edge environments. X-SA distinctively employs a parameterized$2 \times N$processing element design enabling dynamic computational scaling, unlike fixed systolic arrays. It integrates an interleaved matrix buffer to alleviate memory bottlenecks and optimize dataflow. Experimental results demonstrate X-SA achieves a$2.83 \times$performance speedup over the Vortex baseline with minimal Look-Up Table overhead of 2.8% and Flip-Flops overhead of 1.4%.. It offers comparable performance to a standard$4 \times 4$systolic array but with significantly reduced area by 46.26% and power by 39.42%, and superior processing element utilization for irregular matrices. X-SA provides a approach to help improve the performance of some AI applications running on edge GPGPUs relatively in resource-constrained environments. Yingsong Wang, Zhenzhen Jia, Ling Yang 0008, Hongbing Tan, Junsheng Chang, Junbo Tie, Libo Huang 0002 |
HPCC | 3 |
| 2025 | PolyPE: An Efficient Multi-Precision Multi-Mode Floating-Point Processing Element for HPC and AIabstractIn this paper, an efficient multi-precision multimode floating-point Processing Element is designed for HPCenabled AI workloads, called PolyPE, in which Poly means multiprecision multi-mode. It supports both conventional and mixedprecision FMA operations, including single-FMA, dual-FMA, and quad-FMA modes, as well as quad-FMA-add for enhanced throughput. The supported precisions include double precision, single precision, half precision, TF32, and BF16. At each clock cycle, the processing element can perform one double-precision, two single-precision, or four half-precision operations. Compared to existing designs, it offers broader precision support, including TF32 and BF16, with higher throughput and lower hardware overhead, achieving up to 5× improvement over standard FMA. We integrated the design into an open-source GPGPU and extended its instruction set. Experimental results show up to 2.17× performance gain, with 27.2% and 41.2% reductions in LUT and FF usage, respectively, while preserving functional equivalence. Zhenzhen Jia, Hongbing Tan, Ling Yang 0008, Hui Guo 0004, Junsheng Chang, Yongwen Wang, Libo Huang 0002 |
ICCD | 3 |
| 2025 | RVAM16: a low-cost multiple-ISA processor based on RISC-V and ARM Thumb
Libo Huang 0002, Ling Yang 0008, Sheng Ma, Yongwen Wang, Yuanhu Cheng |
Frontiers Comput. Sci. | 3 |
| 2025 | Optimizing value prediction for ILP processors: A design space exploration approach
Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001 |
Integr. | 1 |
| 2024 | Cost-Effective Value Predictor for ILP processors through Design Space ExplorationabstractValue prediction is a microarchitectural technique that enhances processor performance by speculatively breaking true data dependencies. It has demonstrated improved performance in both single-threaded and multi-threaded workloads, rendering it an appealing microarchitectural approach. While high-performance value predictors can achieve impressive accuracy, they may also incur significant costs in terms of area, power consumption, and complexity. Therefore, there is a demand for lightweight value prediction techniques capable of striking a favorable balance between performance and overhead. However, designing value predictors with superior performance using limited resources presents an urgent challenge. Consequently, this work proposes a design space exploration framework for the state-of-the-art EVES value predictor, aiming to efficiently configure the design parameters of the value predictor within constrained RAM resources. Additionally, the article evaluates the performance of the explored value predictor across a wide range of workloads. The explored value predictors exhibit high efficiency across RAM sizes ranging from 2KB to 16KB while maintaining acceptable computational complexity. Furthermore, the results indicate that the explored value predictor achieves optimal efficiency under the 2KB constraint, with the highest acceleration-to-cost ratio reaching 4.02%/KB. Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | Confidence Counter Modelling for Value PredictorabstractValue prediction suffers from high penalties for mispredictions, so confidence mechanisms using saturation counters are often introduced to increase the output threshold of the value predictor. Statistics from saturation counters allow the fine-grained analysis of value prediction performance. However, for architects, traditional simulator is time-consuming and non-scalable, and for software developers, value predictors under the microarchitecture are usually invisible, which makes it difficult to optimize software. In this paper, we model saturation counters commonly used in value predictors with the Markov method, which enables offline estimation of misprediction rates. Such offline analysis can better provide information for performance estimation and compiler optimization. The final experimental results show that the difference between the misprediction rate obtained by the model and the simulator is very small, with the arithmetic average of 0.07% and 0.19% for the two commonly used counters, respectively. Ling Yang 0008, Libo Huang 0002 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | Low-Cost Multiple-Precision Multiplication Unit Design For Deep LearningabstractLow-precision formats have been proposed and applied to deep learning algorithms to speed up training and inference. This paper proposes a novel multiple-precision multiplication unit(MU) for deep learning. The proposed MU supports four types of precision for floating-point(FP) numbers-FP8-E4M3, FP8-E5M2, FP16, FP32-and 8-bit fixed-point(FIX) numbers. The MU can execute four parallel FP8 and eight parallel FIX8 multiplications simultaneously in one cycle, or four parallel FP16 multiplications fully pipelined with a latency of one, or one FP32 multiplication with a latency of one cycle. The simultaneous execution of FIX8 and FP8 can meet the requirements of the specific deep learning algorithms. Thanks to the low-precision-combination(LPC) and vectorization design method, multiplication in any precision can get 100% utilization of the multiplier resources, and the MU can adopt a lower clock delay to achieve better performance in all data types. Compared with the existing multiple-precision units designed for deep learning, this MU can support more types of low-precision formats by lower area overhead; and exhibits higher throughput at FIX8 with at least 8× improvement. Libo Huang 0002, Hongbing Tan, Ling Yang 0008, Qianming Yang |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications
Hongbing Tan, Run Yan, Ling Yang 0008, Libo Huang 0002, Liquan Xiao, Qianming Yang |
ICA3PP | 3 |
| 2022 | Optimizing Winograd Convolution on GPUs via Partial Kernel Fusion
Gan Tong, Run Yan, Ling Yang 0008, Mengqiao Lan, Yuanhu Cheng, Yashuai Lü, Sheng Ma, Libo Huang 0002 |
NPC | 3 |