EDBT 2026 Demo / reviewers in the wild / expert
Lin Peng 0001
dblp:34/1725-1
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-6828-3364ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Selection of Supervised Learning-Based Sparse Matrix Reordering Algorithms
Tao Tang 0001, Youfu Jiang, Yingbo Cui 0001, Jianbin Fang, Peng Zhang 0061, Lin Peng 0001, Chun Huang 0006 |
HiPC | 6 |
| 2025 | LLM-Guided Mutation Location Selection for Vulnerability-Aware JavaScript Engine FuzzingabstractModern JavaScript engines employ multi-tier JIT compilation for high performance, but these aggressive optimizations often introduce subtle and hard-to-detect security vulnerabilities. Existing fuzzers, whether syntax-aware or coverage-guided, lack mechanisms to focus mutations on vulnerability-relevant regions identified through semantic analysis. This paper presents LocFuzz, a JavaScript engine fuzzer that leverages Large Language Models (LLMs) to identify vulnerabilitysensitive mutation locations based on the semantic structure of historical bug samples. Unlike prior LLM-based fuzzers that primarily rely on models for input generation, LocFuzz introduces a novel mutation location guidance mechanism to direct semantic-aware mutations. It employs a temperature-controlled sampling strategy to balance exploration and precision when selecting mutation sites, and measures execution feature similarity via function address sequences to evaluate semantic consistency between original and mutated runs. Evaluations on V8 and SpiderMonkey show that LocFuzz significantly outperforms baseline fuzzers in test validity, behavioral similarity, and vulnerability activation, uncovering four new SpiderMonkey bugs. These results demonstrate the potential of LLM-guided semantic analysis to enhance both precision and efficiency in JavaScript engine fuzzing. Jizhe Li, Lin Peng 0001, Muxin Xu |
TrustCom | 4 |
| 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph ProcessingabstractGraph attention networks (GATs) have advanced performance in various application domains by introducing the attention mechanism into the graph neural networks (GNNs). The inefficiency of running GATs on CPUs or GPUs necessitates specialized hardware designs. Unfortunately, previous specialized architecture designs have focused on either the GNN architecture or the attention mechanism, resulting in limited performance and leaving ample room for improvement. This article presents Gator , a joint optimization approach with software–hardware co-designs for GAT inference. On the software level, Gator leverages degree-weighted graph partitioning and parameter-adaptive feature selection techniques to preprocess the input graph data, mining subgraph-level parallelism and mitigating the computation bottleneck of the dedicated dataflow. On the hardware level, Gator designs a unified processing engine to support various kernels by extracting a common computation pattern and a dimension-aware microarchitecture for efficient partial sum reduction. Extensive experiments show that our approach can achieve 11.5× more efficiency compared to NVIDIA RTX 4090 and provide a speedup of 3× to 9.4×, along with a 2.6× to 4.7× reduction in memory traffic, when compared to six state-of-the-art methods, with minimal accuracy loss. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zixiao Yu |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | VLASPH: Smoothed Particle Hydrodynamics on VLA SIMD Architectures
Xiaokang Fan, Zhen Ge, Tao Tang 0001, Chun Huang 0006, Lin Peng 0001, Canqun Yang |
Euro-Par (3) | 6 |
| 2024 | Parallel Optimization for Accelerating the Generation of Correctly Rounded Elementary FunctionsabstractCorrectly rounded elementary mathematical functions are crucial for numerical computations and scientific applications. Generating these functions accurately is a challenging task. The latest methods automate this process by transforming the problem of generating correctly rounded elementary mathematical functions into a linear programming problem. However, this generation process is serial, and the inefficiency of serialization hinders the creation of new elementary mathematical functions and limits the broader application of the technique. Xianglin Wang, Xin Yi 0002, Hengbiao Yu, Chun Huang 0006, Lin Peng 0001 |
ICPP | 5 |
| 2024 | Auto-tuning for HPC storage stack: an optimization perspectiveabstractAbstract Storage stack layers in high-performance computing (HPC) systems offer many tunable parameters controlling I/O behaviors and underlying file system settings. The setting of these parameters plays a decisive role in I/O performance. Nevertheless, the increasing complexity of data operations and storage architectures makes identifying a set of well-performing configurations a challenge. Auto-tuning is a promising technology. This paper presents a comprehensive survey on "Auto-tuning in HPC I/O". We expound a general storage structure based on a general storage stack and critical elements of auto-tuning, and categorize related studies according to the way of tuning. On the basis of the order in which the approaches were applied, we introduce the specific works of each approach in detail, and summarize and compare the pros and cons of these approaches. Through a comprehensive and in-depth study of existing research, we elaborate on the development history of auto-tuning technology in HPC I/O, analyze the current situation, and provide guidance for optimization technology in the future. Zhangyu Liu, Jinqiu Wang, Qingzhen Ma, Lin Peng 0001, Zhanyong Tang |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise ProductabstractSparse-dense matrix multiplication (SpMM) is the performance bottleneck of many high-performance and deep-learning applications, making it attractive to design specialized SpMM hardware accelerators. Unfortunately, existing hardware solutions do not take full advantage of data reuse opportunities of the input and output matrices or suffer from irregular memory access patterns. Their strategies increase the off-chip memory traffic and bandwidth pressure, leaving much room for improvement. We present Mentor , a new approach to designing SpMM accelerators. Our key insight is that column-wise dataflow, while rarely exploited in prior works, can address these issues in SpMM computations. Mentor is a software-hardware co-design approach for leveraging column-wise dataflow to improve data reuse and regular memory accesses of SpMM. On the software level, Mentor incorporates a novel streaming construction scheme to preprocess the input matrix for enabling a streaming access pattern. On the hardware level, it employs a fully pipelined design to unlock the potential of column-wise dataflow further. The design of Mentor is underpinned by a carefully designed analytical model to find the tradeoff between performance and hardware resources. We have implemented an FPGA prototype of Mentor . Experimental results show that Mentor achieves speedup by geomean 2.05× (up to 3.98×), reduces the memory traffic by geomean 2.92× (up to 4.93×), and improves bandwidth utilization by geomean 1.38× (up to 2.89×), compared with the state-of-the-art hardware solutions. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zidong Du, Yongwei Zhao 0001, Zheng Wang 0079 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 3 |
| 2023 | Optimizing HPC I/O Performance with Regression Analysis and Ensemble LearningabstractTo improve parallel I/O performance, it is imperative to optimize the adjustable parameters across the different layers of the I/O software stack. Finding an optimal configuration for different scenarios is hampered by the complex interaction dynamics between these parameters and the large parameter space. Previous research efforts have focused on tuning these parameters using independent algorithms; however, these approaches exhibit certain shortcomings such as unstable performance results and delayed convergence rates.This paper introduces OPRAEL, an auto-tuning approach on parallel I/O tasks by ensembles and performance modeling using regression analysis. To test its effectiveness, we applied this approach on the Tianhe-II supercomputer using one well-known I/O benchmark(IOR) and two I/O kernels(S3D-I/O, BT-I/O). Leveraging our experience in predictive modeling, we optimized the tuning of the I/O stack parameters. Our experimental results show a remarkable 10.2X improvement in write performance speedup for the optimization task with BT-I/O and a 500x500x500 input. We also compared the potential of using a single search algorithm versus using reinforcement learning search in the I/O parameter auto-optimization task. Our results show that OPRAEL outperforms the traditional approach, resulting in a maximum 8.4X improvement in write performance for the 128-process IOR optimization. Zhangyu Liu, Cheng Zhang 0007, Jianbin Fang, Lin Peng 0001, Guixin Ye, Zhanyong Tang |
CLUSTER | 5 |
| 2021 | Large-Scale Parallel Alignment Algorithm for SMRT Reads
Yingbo Cui 0001, Peng Zhang 0061, Tao Tang 0001, Lin Peng 0001, Chun Huang 0006, Canqun Yang, Xiangke Liao |
ICA3PP (2) | 7 |