Kunpeng Xie

dblp:225/3464 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ByteEye: A smart contract vulnerability detection framework at bytecode level with graph neural networks
abstract
Smart contract vulnerability detection has attracted increasing attention due to billions of economic losses caused by vulnerabilities. Existing smart contract vulnerability detection methods have high false negative and high false positive rates. To address these issues, we present ByteEye, a bytecode level smart contract vulnerability detection framework with Graph Neural Networks (GNNs). ByteEye first constructs an edge-enhanced Control Flow Graph (CFG) to maintain rich information from the low-level bytecode with low latency. ByteEye also designs and incorporates both general information and vulnerability-specific information into its detection method as bytecode level features. Furthermore, ByteEye flexibly supports machine/deep learning models, especially with graph neural networks, which can facilitate vulnerability detection precisely. The extensive experimental results highlight that ByteEye outperforms the state-of-the-art approaches on all three types of vulnerability detection. ByteEye can achieve an average of 35.29%, 43.95%, and 6.38% higher on F1 than the bytecode level best-performed baseline on reentrancy vulnerability, timestamp dependency vulnerability, and integer overflow/underflow vulnerability, respectively. Moreover, ByteEye can detect 361 new vulnerabilities in real-world smart contracts, which are reported for the first time. ByteEye enhances control flow information, designs general bytecode-level features with expert knowledge, and flexibly supports deep learning models, particularly GNNs, thus achieving high detection effectiveness.
Jinni Yang, Shuang Liu 0007, Surong Dai, Yaozheng Fang, Kunpeng Xie, Ye Lu 0004
Autom. Softw. Eng.5
2026 Beyond benchmarks: Towards robust artificial intelligence bone segmentation in socio-technical systems
abstract
Despite the advances in automated medical image segmentation, AI models still underperform in various clinical settings, posing challenges for integration into real-world workflows. In this pre-registered prospective multicenter evaluation, we analyzed 20 state-of-the-art mandibular segmentation models across 19,218 segmentations of 1,000 clinically resampled CT/CBCT scans. Our results suggest that for a given model, segmentation accuracy can vary by up to 25% in Dice score as socio-technical factors such as voxel size, bone orientation, and patient conditions (e.g., osteosynthesis or pathology) shift from favorable to adverse. Higher sharpness, isotropic smaller voxels, and neutral orientation significantly improved results, while metallic osteosynthesis and anatomical complexity led to significant degradation. Our findings challenge the common view of AI models as “plug-and-play” tools and suggest evidence-based optimization recommendations for both clinicians and developers. This will in turn boost the integration of AI segmentation tools in routine healthcare.
Kunpeng Xie, Lennart Johannes Gruber, Martin Crampen, Elias Tappeiner, Maxime Gillot, Jan Schepers, Jiangchang Xu, Tobias Pankert, Michel Beyer, Negar Shahamiri, Reinier ten Brink, Gauthier Dot, Charlotte Weschke, Niels van Nistelrooij, Pieter-Jan Verhelst, Zhibin Xu, Jonas Bienzeisler, Ashkan Rashad, Tabea Flügge, Ross Cotton, Shankeeth Vinayahalingam, Robert R. Ilesan, Stefan Raith, Dennis Madsen, Constantin Seibold, Tong Xi 0001, Stefaan Bergé, Sven Nebelung, Oldrich Kodym, Osku Sundqvist, Florian M. Thieringer, Hans Lamecker, Antoine Coppens, Thomas Potrusil, Joep Kraeima, Max J. H. Witjes, Guomin Wu, Xiaojun Chen 0003, Adriaan Lambrechts, Stefan Zachow, Alexander Hermans, Daniel Truhn, Victor Alves, Jan Egger, Rainer Röhrig, Frank Hölzle, Behrus Hinrichs-Puladi
Expert Syst. Appl.1
2024 ADS-CNN: Adaptive Dataflow Scheduling for lightweight CNN accelerator on FPGAs
Xianzhong Xie, Kunpeng Xie, Dezhi Yi, Ye Lu 0004, Keke Gai
Future Gener. Comput. Syst.4
2024 Winols: A Large-Tiling Sparse Winograd CNN Accelerator on FPGAs
abstract
Convolutional Neural Networks (CNNs) can benefit from the computational reductions provided by the Winograd minimal filtering algorithm and weight pruning. However, harnessing the potential of both methods simultaneously introduces complexity in designing pruning algorithms and accelerators. Prior studies aimed to establish regular sparsity patterns in the Winograd domain, but they were primarily suited for small tiles, with domain transformation dictating the sparsity ratio. The irregularities in data access and domain transformation pose challenges in accelerator design, especially for larger Winograd tiles. This paper introduces “Winols,” an innovative algorithm-hardware co-design strategy that emphasizes the strengths of the large-tiling Winograd algorithm. Through a spatial-to-Winograd relevance degree evaluation, we extensively explore domain transformation and propose a cross-domain pruning technique that retains sparsity across both spatial and Winograd domains. To compress pruned weight matrices, we invent a relative column encoding scheme. We further design an FPGA-based accelerator for CNN models with large Winograd tiles and sparse matrix-vector operations. Evaluations indicate our pruning method achieves up to 80% weight tile sparsity in the Winograd domain without compromising accuracy. Our Winols accelerator outperforms dense accelerator by a factor of 31.7× in inference latency. When compared with prevailing sparse Winograd accelerators, Winols reduces latency by an average of 10.9×, and improves DSP and energy efficiencies by over 5.6× and 5.7×, respectively. When compared with the CPU and GPU platform, Winols accelerator with tile size 8× 8 achieves 24.6× and 2.84× energy efficiency improvements, respectively.
Kunpeng Xie, Ye Lu 0004, Xinyu He 0002, Dezhi Yi, Huijuan Dong, Yao Chen 0008
ACM Trans. Archit. Code Optim.1
2022 Elastic Significant Bit Quantization and Acceleration for Deep Neural Networks
abstract
Quantization has been proven to be a vital method for improving the inference efficiency of deep neural networks (DNNs). However, it is still challenging to strike a good balance between accuracy and efficiency while quantizing DNN weights or activation values from high-precision formats to their quantized counterparts. We propose a new method called elastic significant bit quantization(ESB) that controls the number of significant bits of quantized values to obtain better inference accuracy with fewer resources. We design a unified mathematical formula to constrain the quantized values of the ESB with a flexible number of significant bits. We also introduce a distribution difference aligner (DDA) to quantitatively align the distributions between the full-precision weight or activation values and quantized values. Consequently, ESB is suitable for various bell-shaped distributions of weights and activation of DNNs, thus maintaining a high inference accuracy. Benefitting from fewer significant bits of quantized values, ESB can reduce the multiplication complexity. We implement ESB as an accelerator and quantitatively evaluate its efficiency on FPGAs. Extensive experimental results illustrate that ESB quantization consistently outperforms state-of-the-art methods and achieves average accuracy improvements of 4.78%, 1.92%, and 3.56% over AlexNet, ResNet18, and MobileNetV2, respectively. Furthermore, ESB as an accelerator can achieve 10.95 GOPS peak performance of 1k LUTs without DSPs on the Xilinx ZCU102 FPGA platform. Compared with CPU, GPU, and state-of-the-art accelerators on FPGAs, the ESB accelerator can improve the energy efficiency by up to 65, 11, and 26, respectively.
Ye Lu 0004, Kunpeng Xie, Zongming Jin, Tao Li 0022, Yanzhi Wang 0001
IEEE Trans. Parallel Distributed Syst.3
2019 Random Inception Module and Its Parallel Implementation
Yingqi Gao, Kunpeng Xie, Song Guo 0002, Kai Wang 0001, Hong Kang, Tao Li 0022
APPT2
2019 An Efficient Log Parsing Algorithm Based on Heuristic Rules
Xueshuo Xie, Kunpeng Xie, Zhi Wang 0014, Ye Lu 0004, Yujun Zhang 0001
APPT3
2019 LHC: A Low-Power Heterogeneous Computing Method on Neural Network Accelerator
abstract
Accelerators can achieve high performance and low energy consumption in training or inference of neural networks. If the Non-Neural Network (Non-NN) algorithms with large amount of computation could make full use of the accelerators, it is possible to speed up its implementation, reduce energy consumption, and achieve load balancing, especially on mobile devices equipped with accelerators. However, accelerators are dedicated to neural network calculations, so that other Non-NN algorithms have difficulty in using their advantages. Furthermore, many hardware-specific restrictions have become the obstacles, such as constrained precision of operands and limited computation scale. In this paper, we propose a method named Low-power Heterogeneous Computing (LHC) to bridge the gap between Non-NN algorithms and NN accelerators. Firstly, we analyze the general principle of the accelerator and reveal the calculation model of the accelerator. To hide the details of the underlying neural network library, we extract some operators from the limited number of types of neural network computation they support. We encapsulate the low-level library, extract operators suitable for general algorithms, and implement some more advanced operators that can adapt to the constrained hardware conditions. These operators could facilitate programmers to implement some Non-NN algorithms. In the aspect of the algorithm, we extract the computationally intensive parts of the Non-NN algorithm and deploy these computational tasks on the accelerator by calling the operators. To verify our method, we implement three Non-NN algorithms by using operators and adjusting these algorithms, include Grid-based Motions Statistics, k-Nearest Neighbors, and k-Means, on a specific accelerator, Cambricon-1A. The experimental results show that the energy consumption of calculation is reduced by up to 5.4x, compared with the CPU baseline. Our method can be further applied to other similar accelerators.
Fangxin Liu, Kunpeng Xie, Shusheng Liu, Ye Lu 0004, Tao Li 0022
ICPADS2