Zhuojun Chen

dblp:184/0392 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-0431-8852ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 Evaluation of Thermal and Power integrity and its Impact on Performance for 3D Memory-on-Logic CPUs with FSPDN and BSPDN
abstract
While three-dimensional (3D) Memory-on-Logic integration benefits high-performance computing (HPC), it faces critical bottlenecks in power delivery and thermal management. This paper presents a comprehensive power, performance, area, and thermal (PPAT) evaluation of a 3D Memory-on-Logic CPU utilizing Frontside Power Delivery Network (FSPDN) and Backside Power Delivery Network (BSPDN). Our analysis reveals a fundamental trade-off: while BSPDN significantly improves power integrity by reducing logic IR drop by 7.7 × (vs. 3D FSPDN CPU) and 12× (vs. 2D CPU), the extreme substrate thinning required for backside connectivity severely impedes lateral heat dissipation, raising peak temperatures by ~8°C (vs. 3D FSPDN CPU) and ~12°C (vs. 2D CPU). By incorporating thermal-electrical coupling into a spatial-temperature-aware timing analysis, we demonstrate that unlike 3D FSPDN which yields negligible gains over 2D case due to through-silicon via bottlenecks, the superior power integrity of BSPDN decisively outweighs thermal penalties, achieving a net ~30% performance improvement over the 2D counterpart.
Xincheng Liu, Linqiu Wang, Haolan Yang, Zhuojun Chen, Lianmao Peng, Rongmei Chen
DATE6
2026 Architecture, Design and Technology Co-optimization for 3D ICs with Advanced BSPDN Considering Power & Thermal Integrity Impact
abstract
This paper presents a comprehensive power and thermal integrity analysis of a commercial IP based 7nm 3D CPU with a much larger SRAM area compared to its logic section. We systematically investigate the impact of different 3D stacking architectures—Memory-on-Logic (MoL) and Logic-on-Memory (LoM)—combined with both front-side and back-side power delivery networks (FSPDN/BSPDN). A key contribution is a novel lightweight IR drop modeling tool developed in-house, which enables supper fast and highly accurate power integrity estimation at early physical design stages—far before signoff—significantly reducing design iteration time caused by IR violations. This tool also fills a critical gap in commercial EDA support for advanced 3D integration and BSPDN evaluation. Using this tool alongside multi-physics thermal simulations, we compare four 3D design scenarios. Results show that the MoL architecture with BSPDN achieves an optimal balance between power and thermal integrity: it reduces worst-case IR drop in the logic die to just one-fourth of the 2D reference, and lowers peak temperature by over 15°C compared to a LoM counterpart. Further improvements, 50% in IR drop decrease and 14°C temperature reduction, are attainable through TSV optimization and high-thermal-conductivity material integration. This study provides essential 3D architeture, design and technology cooptimization methodologies for future high-perfermance 3D CPUs of advanced technology nodes.
Haolan Yang, Xingcheng Liu, Linqiu Wang, Feifan Xie, Zhuojun Chen, Lianmao Peng, Rongmei Chen
DATE7
2026 CIMA: An Energy-Efficient In-Memory Accelerator Architecture based on Non-Volatile Capacitive Crossbar Array for Random Forest Inference
Jinrong Zhou, Shaoan Yan, Zhuojun Chen
ISCAS5
2025 Traffic Anomaly Detection through Generative Modeling of Multi-Agent Interactions in Traffic Flow
Zhuojun Chen, Tacitus Hui, Xinghua Zhu, Dongzhe Su
AAMAS1
2025 SpMARD: A Sparse-Sparse Matrix Multiplication Accelerator with Reconfigurable Dataflow for DNN Workloads
abstract
Deep learning becomes increasingly popular, and its main workload is Sparse-Sparse Matrix Multiplication (SpMSpM). Most SpMSpM accelerators usually only support a single dataflow. Different dataflows have different performance in different computing environments. Therefore, the single-dataflow accelerator cannot maintain the highest performance in all environments. Compared with single-dataflow accelerators, multi-dataflow accelerators provide flexible options for different workloads and improve the overall performance. Flexagon, Sparm, and SPADA are state-of-the-art multi-dataflow accelerators. However, the computation process of Flexagon and Sparm is not fully pipelined, and SPADA cannot support inner product dataflow. Additionally, Flexagon, Sparm, and SPADA cannot switch dataflows quickly and accurately. Inspired by these observations, we present SpMARD, a SpMSpM accelerator with reconfigurable dataflow. The computation process of SpMARD is fully pipelined, and SpMARD can support six dataflow variants simultaneously. Through the design of a Two-stage Pipeline Adder Network (TPAN) and a Position-based Psum Array (PPA), SpMARD can execute element-level merging, which can hide the merging overhead. Through the quantitative analysis of dataflows, we implement a Dataflow Switcher (DSwitcher), which can switch dataflows more efficiently. For the SpMSpM workload, the performance (GOPS) of the SpMARD we proposed is 1.27 times that of Flexagon, 1.18 times that of Sparm, and 1.22 times that of SPADA.
Bo Wang 0159, Sheng Ma, Yunping Zhao, Shengbai Luo, Lizhou Wu, Dongsheng Li 0001, Zhuojun Chen
ACM Trans. Archit. Code Optim.9
2025 A 1.31-ppm/ ° C CMOS Voltage Reference With Second-Order and Sub-Ranging Compensation
abstract
This article presents a high-order compensation voltage reference with a low temperature coefficient (TC) over a wide temperature range for the high-precision signal processing circuits. By using a shunt resistor in the current-mode bandgap reference, a nonlinear PTAT current for second-order temperature compensation is achieved. Besides, the sub-ranging high-order compensation is applied in the voltage reference to obtain a lower TC. The proposed voltage reference was designed and fabricated in a standard 180-nm CMOS technology. The proposed voltage reference achieves an average TC of 1.31-ppm/∘C after one-point batch trimming from -40∘C to 90∘C. It consumes 28 μA at 27∘C and occupies an area of 0.075 mm2.
Wenzhao Lv, Yilun Cai, Zhuojun Chen
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Stacking Deep Set Networks and Pooling by Quantiles
abstract
We propose Stacked Deep Sets and Quantile Pooling for learning tasks on set data. We introduce Quantile Pooling, a novel permutation-invariant pooling operation that synergizes max and average pooling. Just like max pooling, quantile pooling emphasizes the most salient features of the data. Like average pooling, it captures the overall distribution and subtle features of the data. Like both, it is lightweight and fast. We demonstrate the effectiveness of our approach in a variety of tasks, showing that quantile pooling can outperform both max and average pooling in each of their respective strengths. We also introduce a variant of deep set networks that is more expressive and universal. While Quantile Pooling balances robustness and sensitivity, Stacked Deep Sets enhances learning with depth.
Zhuojun Chen, Xinghua Zhu, Dongzhe Su, Justin C. I. Chuang
ICML1
2024 Improving Radiation Reliability of SRAM-Based Physical Unclonable Function With Self-Healing and Pre-Irradiation Masking Techniques
abstract
Physical unclonable function (PUF) is an innovative primitive used for key generation and device authentication, which has promising applications for resource-limited scenarios such as satellite communication. However, the reliability of traditional PUF circuits is low and the power consumption is high. Maintaining reliability also requires high costs, which limits its practical application. This article proposes a multimode SRAM PUF based on a 55-nm CMOS process, which has a self-healing feature. By using a voltage tilt preselection mechanism, the unstable PUF cells can be detected and most of them can be healed by mode switching, thereby improving the reliability of the PUF and reducing the costs required for golden key screening tests. However, the total ionizing dose (TID) effect poses a threat to the PUF, which can significantly increase the bit error rate (BER). In this article, the radiation effect of the PUF is characterized and its mechanism is analyzed. Besides, by means of the self-healing and preirradiation temperature and voltage masking techniques, the radiation reliability of the proposed PUF can be improved avoiding traditional destructive testing. The experimental results demonstrate that the BER can be reduced to as low as 0.0183%, which is$345\times $lower than the raw BER, after irradiation up to 100 krad(Si). The proposed technique presents great potential for communication security in the aerospace environment.
Zhuojun Chen, Jinghang Chen, Zujun Wang, Ding Ding 0003
IEEE Trans. Very Large Scale Integr. Syst.1
2023 PUF-CIM: SRAM-Based Compute-In-Memory With Zero Bit-Error-Rate Physical Unclonable Function for Lightweight Secure Edge Computing
abstract
With the rapid development of the Internet of Things (IoT), compute-in-memory (CIM) is a promising candidate for edge computing, which eliminates the need for frequent data transfer between the memory and processing units. On one hand, CIM provides an efficient approach to achieve deep neural network (DNN) inference; on the other hand, its security issues such as model leakage are ignored in most of state-of-the-art studies. In the resource-limited conditions, it is essential to develop a lightweight secure scheme for CIM without significantly sacrificing its performance. In this work, a physical unclonable function (PUF)-CIM macro based on static random access memory (SRAM) is proposed for lightweight model protection. An SRAM-based PUF extracts unique and stable keys from the mismatch of multirow discharge rate and further achieves zero bit error rate (BER) under temperature and voltage fluctuations by the on-chip masking technique. Besides, we propose a 10T SRAM bit cell to perform XOR operation and multiply-and-accumulate (MAC) operation in the same cell, which avoids weight movement during encryption and decryption. The whole process including key generation, XOR encryption, and ciphertext MAC operation has been realized. The test chip is designed and fabricated in a 55-nm CMOS process. Experimental results show that the proposed PUF with on-chip masking can provide zero unstable bits and BER under temperature and voltage variations. The proposed PUF-CIM achieves 96.20% inference accuracy in MNIST. Besides, it presents 93.35–121.38-TOPS/W energy efficiency and 1324.24-GOPS throughput. Compared with the normal CIM, the proposed PUF-CIM can enhance AI model security, only with accuracy loss of 1% and energy efficiency degradation of 16.1%.
Zhuojun Chen, Renlong Li, Jinzhe Tan, Ding Ding 0003
IEEE Trans. Very Large Scale Integr. Syst.1
2022 DA PUF: dual-state analog PUF
abstract
Physical unclonable function (PUF) is a promising lightweight hardware security primitive that exploits process variations during chip fabrication for applications such as key generation and device authentication. Reliability of the PUF information plays a vital role and poses a major challenge for PUF design. In this paper, we propose a novel dual-state analog PUF (DA PUF) which has been successfully fabricated in 55nm process. The 40,960 bits generated by the fabricated DA PUF pass the NIST randomness test with reliability over 99.99% for working environment of -40 ~ 125° C (temperature) and 0.96 ~ 1.44V (voltage), outperforming the two state-of-the-art analog PUFs reported in JSSC 2016 and 2021.
Jiliang Zhang 0002, Zhuojun Chen, Wenshang Li, Gang Qu 0001
DAC3
2022 STT-MRAM-Based Reliable Weak PUF
abstract
In recent years, micro-nano device characteristics like ferroelectrics and resistive switching are being used to build important security primitives such as Physical Unclonable Function (PUF). The micro-nano device-based hardware security primitives, although with higher security, energy efficiency, and integration density, suffer from serious reliability issues caused by process scaling. To mitigate this issue, this paper introduces a reconfigurable weak PUF based on spin-transfer torque magnetoresistive random-access memory (STT-MRAM), which adopts the crossing switches implemented with simple demultiplexes (DEMUXs) to improve the flexibility and reliability. Moreover, two algorithms,neighboring bit linesandtop-$n$n, are proposed to enlarge the gap between two parallel reading currents, thus further enhancing the reliability of PUF responses. Experimental results demonstrate that the proposed PUF scheme achieves good uniqueness (50.64 percent), uniformity (50.02 percent), and bit-aliasing ($\approx$49.80%). Particularly, the proposed method significantly improves the PUF reliability, achieving low bit error rate (BER$\leq$2.13%) within the range of -20$^\circ$C to 90$^\circ$C.
Yupeng Hu 0004, Linjun Wu, Zhuojun Chen, Xiaolin Xu 0001, Keqin Li 0001, Jiliang Zhang 0002
IEEE Trans. Computers3
2021 Competitive Neural Network Circuit Based on Winner-Take-All Mechanism and Online Hebbian Learning Rule
abstract
In this article, we design a memristive competitive neural network circuit based on the winner-take-all (WTA) mechanism and the online Hebbian learning rule. Each synapse of the network contains two memristors whose terminals of signal inputs are opposite. However, only one memristor participates in the calculation each time, and that one is determined by the original input signal. The competitive neural network circuit includes two parts: forward calculation and weight update. In this article, the forward calculation part of the circuit is designed based on the WTA mechanism. The combination of the leaky-integrate-and-fire (LIF) model and pMOS realizes the lateral inhibition of neurons. The design of the weight updating part is based on Hebbian learning rules. In each cycle, only synaptic memristors connected to the winner output neuron in forward calculation can be adjusted. The voltage used for synaptic memristor adjustment comes from the membrane voltage of the winner output neuron. The whole neural network circuit does not need the participation of a central processing unit (CPU) or a field-programmable gate array (FPGA) and really realizes parallel calculation, the saving of area, power consumption, and a certain extent computing-in-memory. Based on the circuit designed in PSPICE, we simulated the classification of$5\times3$pixel pictures. The changing trend of weights in the training phase and the high recognition accuracy in the recognition phase prove that the network can learn and recognize different patterns. The competitive neural network can be applied to the neuromorphic system of visual pattern recognition.
Zhuojun Chen, Judi Zhang, Shuangchun Wen, Ya Li 0008, Qinghui Hong
IEEE Trans. Very Large Scale Integr. Syst.1
2020 RankPose: Learning Generalised Feature with Rank Supervision for Head Pose Estimation
Donggen Dai, Wangkit Wong, Zhuojun Chen
BMVC3
2020 Deep Density-Aware Count Regressor
abstract
We seek to improve crowd counting as we perceive limits of currently prevalent density map estimation approach on both prediction accuracy and time efficiency. We leverage multilevel pixelation of density map as it helps improve SNR of training data and therefore, reduce prediction error. To achieve a better model, we introduce multilayer gradient fusion for training a density-aware global count regressor. More specifically, on training stage, a backbone network receives gradients from multiple branches to learn the density information, whereas those branches are to be detached to accelerate inference. By taking advantages of such method, our model improves benchmark results on public datasets and exhibits itself to be a new solution to crowd counting problems in practice.
Zhuojun Chen, Yuchen Yuan, Dongping Liao, Jiancheng Lv 0001
ECAI1