Zhiyuan Pan

dblp:195/1277 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 A side-channel attack (SCA)-resistant and reconfigurable cryptographic engine design for multiple hash algorithms
Jinghe Wang, Wenrui Liu 0002, Jiafeng Cheng, Nengyuan Sun, Zhiyuan Pan, Zhaoyi Niu, Jianghong Li, Linhan Wang, Kangning Song, Haoxiang Yu, Weize Yu
Integr.5
2026 A random modular-reduction (RMR)-based ASIC design of CRYSTALS-Kyber engine against side-channel attacks
Jinghe Wang, Zhiyuan Pan, Nengyuan Sun, Zhaoyi Niu, Wenrui Liu 0002, Jiafeng Cheng, Jianghong Li, Linhan Wang, Kangning Song, Yuzhu Wu, Weize Yu
Integr.2
2026 An Area-Efficient and Low-Latency ASIC Design of Deflate Data Compressor for SSD Applications
abstract
In this brief, a high-speed [multiway parallel (MWP)] hardware-implemented deflate data compressor (DDC) is proposed for reducing the storage of solid-state drives (SSDs). To minimize the area of the DDC, registers instead of static random access memories (SRAMs) are utilized for building hash tables because multiway data within the DDC are able to access a register-based hash table simultaneously. To further reduce the area of the DDC, the output data of indefinite length are concatenated with a tree-type hardware architecture for reducing the overall concatenation complexity. Moreover, a solid mathematical foundation is established for optimizing the latency values of Lempel–Ziv (LZ)77 circuit, the Huffman encoding circuit, and the output data concatenation circuit within the MWP DDC. The results show that the proposed MWP DDC is capable of achieving a 12.1-Gb/s throughput and a 1.76 compression ratio (CR) with a 1.17-mm2area and 0.103-$\mu $s latency, under the synthesis of SMIC 55-nm process design kits (PDKs). Hence, the proposed DDC satisfies the SSD compression requirement for a universal serial bus (USB) 3.2 connector.
Nengyuan Sun, Jianghong Li, Zhaoyi Niu, Jinghe Wang, Zhiyuan Pan, Jiafeng Cheng, Wenrui Liu 0002, Linhan Wang, Kangning Song, Haoxiang Yu, Weize Yu
IEEE Trans. Very Large Scale Integr. Syst.6
2025 Reasoning Runtime Behavior of a Program with LLM: How Far are We?
abstract
Large language models for code (i.e., code LLMs) have shown strong code understanding and generation capabilities. To evaluate the capabilities of code LLMs in various aspects, many benchmarks have been proposed (e.g., HumanEval and ClassEval). Code reasoning is one of the most essential abilities of code LLMs (i.e., predicting code execution behaviors such as program output and execution path), but existing benchmarks for code reasoning are not sufficient. Typically, they focus on predicting the input and output of a program, ignoring the evaluation of the intermediate behavior during program execution, as well as the logical consistency (e.g., the model should not give the correct output if the prediction of execution path is wrong) when performing the reasoning. To address these problems, in this paper, we propose a framework, namely$\boldsymbol{\mathcal{R}}\mathbf{Eval}$, for evaluating code reasoning abilities and consistency of code LLMs with program execution. We utilize existing code benchmarks and adapt them to new benchmarks within our framework. A large-scale empirical study is conducted and most LLMs show unsatisfactory performance on both Runtime Behavior Reasoning (i.e., an average accuracy of 44.4%) and Incremental Consistency Evaluation (i.e., an average IC score of 10.3). Evaluation results of current code LLMs reflect the urgent need for the community to strengthen the code reasoning capability of code LLMs. Our code, data and$\boldsymbol{\mathcal{R}}\mathbf{Eval}$leaderboard are available at https://r-eval.github.io.
Junkai Chen, Zhiyuan Pan, Xing Hu 0008, Zhenhao Li 0002, Ge Li 0001, Xin Xia 0001
ICSE2
2025 Safety in DRL-Based Congestion Control: A Framework Empowered by Expert Refinement
abstract
Deep reinforcement learning (DRL) has been used in congestion control algorithms (CCAs) for its ability to adapt to different network environments. However, its effectiveness is often hindered by the limited availability of training data and constrained training scales. While it has been proved that combining rule-based (expert) CCAs as a guide for DRL (namely hybrid CCAs) can address this limitation, we show through experimental measurements that rule-based CCAs potentially restrict action exploration of DRL models and may cause the DRL models to overly rely on them for higher reward gains. To address this gap, this paper proposes Marten, a framework that improves the effectiveness of rule-based CCAs for DRL. Marten’s key innovations include an entropy-based dynamic exploration scheme that expands the exploration of DRL, and a reward adjustment scheme to prevent the DRL models’ over-reliance on experts in hybrid CCAs. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. Experimental results in both emulated and production networks demonstrate Marten can improve throughput by 0.31% and reduce latency by 12.69% on average compared to the state-of-the-art hybrid CCAs. Compared to BBR, Marten achieves a 2.79% increase in throughput and an 11.73% reduction in latency on average.
Jianer Zhou, Zhiyuan Pan, Zhenyu Li 0001, Gareth Tyson, Weichao Li 0001, Xinyi Qiu, Xinyi Zhang 0004, Gaogang Xie
IEEE Trans. Netw.2
2024 PPT4J: Patch Presence Test for Java Binaries
abstract
The number of vulnerabilities reported in open source software has increased substantially in recent years. Security patches provide the necessary measures to protect software from attacks and vulnerabilities. In practice, it is difficult to identify whether patches have been integrated into software, especially if we only have binary files. Therefore, the ability to test whether a patch is applied to the target binary, a.k.a. patch presence test, is crucial for practitioners. However, it is challenging to obtain accurate semantic information from patches, which could lead to incorrect results.
Zhiyuan Pan, Xing Hu 0008, Xin Xia 0001, Xian Zhan, David Lo 0001, Xiaohu Yang 0001
ICSE1
2024 A low-overhead and high-reliability physical unclonable function (PUF) for cryptography
Wenrui Liu 0002, Jiafeng Cheng, Nengyuan Sun, Heng Sha, Hongyang Zhao, Zhiyuan Pan, Jinghe Wang, Selçuk Köse, Weize Yu
Integr.7
2023 Marten: A Built-in Security DRL-Based Congestion Control Framework by Polishing the Expert
abstract
Deep reinforcement learning (DRL) has been proved to be an effective method to improve the congestion control algorithms (CCAs). However, the lack of training data and training scale affect the effectiveness of DRL model. Combining rule-based CCAs (such as BBR) as a guide for DRL is an effective way to improve learning-based CCAs. By experiment measurement, we find that the rule-based CCAs limit the action exploration and even cause DRL’s excessive dependence to gain higher DRL’s reward gain. To overcome the constraints, we propose Marten, a framework which improves the effectiveness of rule-based CCAs for DRL. Marten uses entropy as the degree of exploration and uses it to expand the exploration of DRL. Furthermore, Marten introduces the shielding mechanism to avoid wrong DRL actions. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. The experimental results in production network demonstrate Marten can improve throughput by 0.36% and reduce latency by 14.89% on average compared with Eagle, and improve throughput by 2.79% and reduce latency by 11.73% on average compared with BBR.
Zhiyuan Pan, Jianer Zhou, Xinyi Qiu, Weichao Li 0001
INFOCOM1
2018 Asynchronous Value Iteration Network
Zhiyuan Pan, Zongzhang Zhang
ICONIP (2)1
2017 Weighted Double Q-learning
abstract
Q-learning is a popular reinforcement learning algorithm, but it can perform poorly in stochastic environments due to overestimating action values. Overestimation is due to the use of a single estimator that uses the maximum action value as an approximation for the maximum expected action value. To avoid overestimation in Q-learning, the double Q-learning algorithm was recently proposed, which uses the double estimator method. It uses two estimators from independent sets of experiences, with one estimator determining the maximizing action and the other providing the estimate of its value. Double Q-learning sometimes underestimates the action values. This paper introduces a weighted double Q-learning algorithm, which is based on the construction of the weighted double estimator, with the goal of balancing between the overestimation in the single estimator and the underestimation in the double estimator. Empirically, the new algorithm is shown to perform well on several MDP problems.
Zongzhang Zhang, Zhiyuan Pan, Mykel J. Kochenderfer
IJCAI2