Chengpeng Xia

dblp:236/2965 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-9520-0229ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BITLUME: Precision-Flexible Photonic Computing for Ultra-Fast and Energy-Efficient DNN Acceleration
abstract
As deep learning expands across emerging domains, computational demands are pushing traditional electronic accelerators to their limits. Silicon photonics has emerged as a promising technology for accelerating deep learning workloads, but precision remains a challenge due to noise and non-idealities. In this paper, we present BITLUME, a novel photonic computing unit that enables multiplications beyond 8-bit precision through a precision-flexible scheme. We further propose an optimized round-truncation algorithm and data mapping strategy for BITLUME to reduce optoelectronic conversions, enhance data reuse, and maintain computational accuracy. A hybrid optoelectronic architecture integrating BITLUME is developed and validated using a prototype built with FPGA, RF, and photonic components, achieving 3.7× lower end-to-end latency than the A100 GPU in dot product. Simulations of training seven DNN models at FP32 show that BITLUME achieves up to 3.35× and 10.78× speedup, and 1.53× and 4.12× energy savings, compared to the state-of-the-art photonic accelerator and A100 GPU, respectively.
Chengpeng Xia, Haibo Zhang 0001, Hao Zhang 0058, Yawen Chen 0001, Amanda S. Barnard
ICCAD1
2025 ROCKET: An RNS-based Photonic Accelerator for High-Precision and Energy-Efficient DNN Training
abstract
In recent years, the rapid development of Deep Neural Networks (DNNs) has posed significant challenges in terms of training duration and costs. High-frequency, low-power photonic computing has emerged as a highly promising solution. However, the substantial cost of data conversion and the limitations introduced by noise in photonic devices continue to hinder the realization of high-precision and energy-efficient DNN training. To address this challenge, we propose a novel photonic accelerator, ROCKET, based on the Residue Number System (RNS). RNS is based on modular arithmetic and enables support for high-precision computation through parallel multi-path low-precision operations. First, we leverage specialized lookup tables to enable high-throughput, low-latency conversions between high-precision and low-precision numerical representations. Next, we design a low-power photonic accelerator architecture utilizing intensity modulators, which minimizes the number of computational components while maximizing data reuse. Subsequently, we propose a hybrid photonic-electronic pipelined dataflow to maximize parallelism within the photonic-electronic computation path. Finally, we develop a high-frequency (4.096 GHz) hybrid photonic-electronic prototype using FPGA, Radio Frequency (RF), and photonic components to validate the feasibility of the ROCKET. Our large-scale simulations on seven mainstream DNN models show that, compared to the A100 GPU, TPU v4, and the state-of-the-art photonic accelerator Mirage, ROCKET achieves speedups of 33×, 243×, and 198×, respectively, while saving energy by factors of 64×, 204×, and 142×.
Hao Zhang 0058, Haibo Zhang 0001, Chengpeng Xia, Zhiyi Huang 0001, Yawen Chen 0001, Amanda S. Barnard
ICS3
2023 STADIA: Photonic Stochastic Gradient Descent for Neural Network Accelerators
abstract
Deep Neural Networks (DNNs) have demonstrated great success in many fields such as image recognition and text analysis. However, the ever-increasing sizes of both DNN models and training datasets make deep leaning extremely computation- and memory-intensive. Recently, photonic computing has emerged as a promising technology for accelerating DNNs. While the design of photonic accelerators for DNN inference and forward propagation of DNN training has been widely investigated, the architectural acceleration for equally important backpropagation of DNN training has not been well studied. In this paper, we propose a novel silicon photonic-based backpropagation accelerator for high performance DNN training. Specifically, a general-purpose photonic gradient descent unit named STADIA is designed to implement the multiplication, accumulation, and subtraction operations required for computing gradients using mature optical devices including Mach-Zehnder Interferometer (MZI) and Mircoring Resonator (MRR), which can significantly reduce the training latency and improve the energy efficiency of backpropagation. To demonstrate efficient parallel computing, we propose a STADIA-based backpropagation acceleration architecture and design a dataflow by using wavelength-division multiplexing (WDM). We analyze the precision of STADIA by quantifying the precision limitations imposed by losses and noises. Furthermore, we evaluate STADIA with different element sizes by analyzing the power, area and time delay for photonic accelerators based on DNN models such as AlexNet, VGG19 and ResNet. Simulation results show that the proposed architecture STADIA can achieve significant improvement by 9.7× in time efficiency and 147.2× in energy efficiency, compared with the most advanced optical-memristor based backpropagation accelerator.
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Jigang Wu
ACM Trans. Embed. Comput. Syst.1
2023 Comparing the performance of multi-layer perceptron training on electrical and optical network-on-chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hao Zhang 0058, Chengpeng Xia
J. Supercomput.6
2022 Three-stage auction scheme for computation offloading on mobile blockchain with edge computing
abstract
Summary Blockchain has been applied in wide range of fields to guarantee security. However, it has been very challenging for blockchain to flourish in mobile environment with limited resources. Existing studies mainly assume that single mobile user can buy the whole resources from edge servers in mobile blockchain. This paper formulates the problem of maximizing the social welfare for computation offloading in mobile blockchain. A three‐stage auction scheme with approximation ratio of based on group‐buying mechanism is proposed to allocate edge server resources for mobile blockchain applications. In the first stage, the miners are divided into groups, and a Vickrey–Clarke–Groves based auction is proposed to determine the bid of each group for each edge server. In the second stage, a matching algorithm is proposed to match edge servers and Access Points for maximizing the profit of edge servers. In the third stage, the edge server resources are allocated to mobile users for mining base on the results in the above stages. We prove that our auction scheme guarantees truthfulness, individual rationality and budget balance. Simulation results show that, the social welfare of our scheme is improved by 33.78%, 21.84%, 19.69%, and 6.69% for 1000 miners, compared with the existing works.
Chengpeng Xia, Yalan Wu, Long Chen 0006, Yawen Chen 0001, Jigang Wu
Concurr. Comput. Pract. Exp.1
2021 Photonic Computing and Communication for Neural Network Accelerators
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Hao Zhang 0058, Jigang Wu
PDCAT1
2021 Combinatorial Double Auction for Resource Allocation in Mobile Blockchain Network
Xuelian Liu, Jigang Wu, Long Chen 0006, Chengpeng Xia, Yidong Li
Wirel. Networks4
2018 ETRA: Efficient Three-Stage Resource Allocation Auction for Mobile Blockchain in Edge Computing
abstract
Blockchain technology is emerging in various fields, to guarantee security of digital currency and internet of things. In this paper, we provide incentive to encourage edge servers to serve mobile users for the mobile blockchain application. We formulate the problem as a resource allocation problem, then we propose a three-stage auction to implement resource allocation specially designed for mobile blockchain, and introduce the group-buying mechanism to motivate mobile users. We prove that our auction scheme is truthful, individual rationality, and computational efficiency. We compare proposed scheme with TACD and HAF mechanisms, and simulation results show that the social welfare achieved by our scheme is higher than that of TACD and HAF mechanisms.
Chengpeng Xia, Xuelian Liu, Jigang Wu, Long Chen 0006
ICPADS1