EDBT 2026 Demo / reviewers in the wild / expert
Xin Zheng 0001
dblp:13/6922-1
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-4931-9664ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Efficiency Bidirectional Translator between SystemC and VerilogabstractThe SystemC language, with its higher level of abstraction, plays a critical role in facilitating hardware/software co-design and architecture exploration. However, as most hardware models are predominantly written in Verilog and translating between SystemC and Verilog remains a challenge, an efficient and reliable tool for translating between these two languages is essential to streamline system development. This article proposes SCAV, a bidirectional translator between SystemC and Verilog, which breaks these limitations. SCAV provides a fully automated solution for translating both SystemC to Verilog and Verilog to SystemC, leveraging a translation framework with front-end/back-end separation. Additionally, SCAV incorporates an Abstract Syntax Tree (AST) filter, optimizing the translation process by filtering out invalid content. The experimental results demonstrate that SCAV achieves a 100% adaptation rate for Verilog and a 98% adaptation rate for SystemC, with 100% accuracy in both directions. Furthermore, SCAV outperforms existing tools, delivering a minimum speedup of 18% across various test cases. Xin Zheng 0001, Yongfeng Zhong, Shaofen Zeng, Huaien Gao, Shuting Cai, Xiaoming Xiong |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2026 | FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V CoresabstractArea efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available athttps://github.com/LX-IC/FPUltra Xian Lin, Jiahao Lan, Xin Zheng 0001, Huanxin Zhuang, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | FLALM: A Flexible Low Area-Latency Montgomery Modular Multiplication on FPGAabstractMontgomery Modular Multiplication (MMM) is widely used in many public key cryptography systems. This paper presents a Flexible Low Area-Latency MMM (FLALM) implementation, which supports Generic Montgomery Modular Multiplication (GMM) and Square Montgomery Modular Multiplication (SMM) operations. A new SMM schedule for the Finely Integrated Product Scanning (FIPS) GMM algorithm is proposed to accelerate SMM with tiny additional design. Furthermore, a new FIPS dual-schedule is proposed to solve the data hazards of this algorithm. Finally, we explore the trade-off between area and latency, and present the FLALM to accelerate GMM and SMM. The FLALM is implemented on FPGA (Virtex-7 platform). The results show that the area*latency (AL) value of FLALM (wordsize$w$=128) is 38.1% and 44.7% better than the previous state-of-art scalable references when performing 1024-bit and 2048-bit GMM, respectively. Moreover, when computing SMM, the advantage of AL value is raised to 73.7% and 86.3% respectively. Yujun Xie 0001, Yuan Liu 0022, Xin Zheng 0001, Bohan Lan, Dengyun Lei, Dehao Xiang, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Computers | 3 |
| 2025 | Efficient Design Space Exploration for the BOOM Using SAC-Based Reinforcement LearningabstractDesign space exploration (DSE) is crucial for optimizing the performance, power, and area (PPA) of CPU microarchitectures ($\mu $-archs). While various machine learning (ML) algorithms have been applied to the$\mu $-arch DSE problem, the potential of reinforcement learning (RL) remains underexplored. In this article, we propose a novel RL-based approach to address the reduced instruction set computer V (RISC-V) CPU$\mu $-arch DSE problem. This approach enables dynamic selection and optimization of$\mu $-arch parameters without relying on predefined modification sequences, thus significantly enhancing exploration flexibility. To address the challenges posed by high-dimensional action spaces and sparse rewards, we use a discrete soft actor-critic (SAC) framework with entropy maximization to promote efficient exploration. In addition, we integrate multistep temporal-difference (TD) learning, an experience replay (ER) buffer, and return normalization to improve sample efficiency and learning stability during training. Our method further aligns optimization with user-defined preferences by normalizing PPA metrics relative to baseline designs. Experimental results on the Berkeley out-of-order machine (BOOM) demonstrate that the proposed approach achieves superior performance compared with state-of-the-art methods, showcasing its effectiveness and efficiency for$\mu $-arch DSE. Our code is available athttps://github.com/exhaust-create/SAC-DSE. Mingjun Cheng, Xin Zheng 0001, Xian Lin, Huaien Gao, Shuting Cai, Xiaoming Xiong, Bei Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Quantization-aware Optimization Approach for CNNs Inference on CPUsabstractData movements through the memory hierarchy are a fundamental bottleneck in the majority of convolutional neural network (CNN) deployments on CPUs. Loop-level optimization and hybrid bitwidth quantization are two representative optimization approaches for memory access reduction. However, they were carried out independently because of the significantly increased complexity of design space exploration. We present QAOpt, a quantization-aware optimization approach that can reduce the high complexity when combining both for CNN deployments on CPUs. We develop a bitwidth-sensitive quantization strategy that can perform the trade-off between model accuracy and data movements when deploying both loop-level optimization and mixed precision quantization. Also, we provide a quantization-aware pruning process that can reduce the design space for high efficiency. Evaluation results demonstrate that our work can achieve better energy efficiency under acceptable accuracy loss. Jiasong Chen, Zeming Xie, Weipeng Liang, Bosheng Liu, Xin Zheng 0001, Jigang Wu, Xiaoming Xiong |
ASPDAC | 5 |
| 2024 | FPUx: High-Performance Floating-Point Support for Cost-Constrained RISC-V CoresabstractIn the Internet of Things (IoT) field, cloud and fog computing dramatically increase the complexity of floating-point (FP) calculations. Cost-constrained microcontrollers (MCUs) urgently need more efficient FP computing methods, such as integrated FP units (FPUs). To this end, this brief proposes FPUx, a high-performance FPU designed through a hybrid pipeline and state-machine approach. The FPUx is integrated into E203 for implementation (E203-FPUx). Furthermore, the Easy-lite is proposed to reduce handshake delay and a range of single-precision FP (FP32) arithmetic IPs are designed to customize FPUs. Compared with E203-FPnew and E203, the performance of E203-FPUx is improved by$1.5\times $and$36\times $, and the total energy consumption is saved by 36% and 1430% on average, respectively. Xian Lin, Heming Liu, Xin Zheng 0001, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | BSSE: Design Space Exploration on the BOOM With Semi-Supervised LearningabstractWith the rising prominence of RISC-V-based microprocessors in processor design, the challenge of exploring the vast and complex RISC-V microarchitecture design space has become increasingly apparent. We propose the Berkeley Out-of-Order Machine Semi-Supervised Explorer (BSSE)—a novel framework leveraging the semi-supervised learning method and parallel emulation to speed up and make tradeoffs on the RISC-V microarchitecture design space exploration (DSE). BSSE constructs the initial training dataset with the microarchitecture experimental design sampling (MEDS) method and then employs the cotraining-style k-nearest neighbors (Co-KNN) model to fit the microarchitecture features to the architectural metric value space. The trained Co-KNN model assists in searching a Pareto-optimal set with parallel emulation. Finally, a distance-based method is proposed to select a designer-preferred microarchitecture from the identified Pareto-optimal set. Extensive experiments on the Berkeley Out-of-Order Machine (BOOM) show that our proposed BSSE method can search for a better Pareto-optimal set with less time consumption compared to the state-of-the-art methods and can find microarchitectures that are equivalent to or even better than the existing manually designed BOOM microarchitectures. Xin Zheng 0001, Mingjun Cheng, Jiasong Chen, Huaien Gao, Xiaoming Xiong, Shuting Cai |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | High-performance Reconfigurable DNN Accelerator on a Bandwidth-limited Embedded SystemabstractDeep convolutional neural networks (DNNs) have been widely used in many applications, particularly in machine vision. It is challenging to accelerate DNNs on embedded systems because real-world machine vision applications should reserve a lot of external memory bandwidth for other tasks, such as video capture and display, while leaving little bandwidth for accelerating DNNs. In order to solve this issue, in this study, we propose a high-throughput accelerator, called reconfigurable tiny neural network accelerator (ReTiNNA), for the bandwidth-limited system and present a real-time object detection system for the high-resolution video image. We first present a dedicated computation engine that takes different data mapping methods for various filter types to improve data reuse and reduce hardware resources. We then propose an adaptive layer-wise tiling strategy that tiles the feature maps into strips to reduce the control complexity of data transmission dramatically and to improve the efficiency of data transmission. Finally, a design space exploration (DSE) approach is presented to explore design space more accurately in the case of insufficient bandwidth to improve the performance of the low-bandwidth accelerator. With a low bandwidth of 2.23 GB/s and a low hardware consumption of 90.261K LUTs and 448 DSPs, ReTiNNA can still achieve a high performance of 155.86 GOPS on VGG16 and 68.20 GOPS on ResNet50, which is better than other state-of-the-art designs implemented on FPGA devices. Furthermore, the real-time object detection system can achieve a high object detection speed of 19 fps for high-resolution video. Xianghong Hu 0001, Hongmin Huang, Xueming Li 0001, Xin Zheng 0001, Qinyuan Ren, Jingyu He, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | An Efficient Image Categorization Method With Insufficient Training SamplesabstractImage classification is an important part of pattern recognition. With the development of convolutional neural networks (CNNs), many CNN methods are proposed, which have a large number of samples for training, which can have high performance. However, there may exist limited samples in some real-world applications. In order to improve the performance of CNN learning with insufficient samples, this article proposes a new method called the classifier method based on a variational autoencoder (CFVAE), which is comprised of two parts: 1) a standard CNN as a prior classifier and 2) a CNN based on variational autoencoder (VAE) as a posterior classifier. First, the prior classifier is utilized to generate the prior label and information about distributions of latent variables; and the posterior classifier is trained to augment some latent variables from regularized distributions to improve the performance. Second, we also present the uniform objective function of CFVAE and put forward an optimization method based on the stochastic gradient variational Bayes method to solve the objective model. Third, we analyze the feasibility of CFVAE based on Hoeffding's inequality and Chernoff's bounding method. This analysis indicates that the latent variables augmentation method based on regularized latent variables distributions can generate samples fitting well with the distribution of data such that the proposed method can improve the performance of CNN with insufficient samples. Finally, the experiments manifest that our proposed CFVAE can provide more accurate performance than state-of-the-art methods. Luyue Lin, Bo Liu 0002, Xin Zheng 0001, Yanshan Xiao |
IEEE Trans. Cybern. | 3 |
| 2022 | A Latent Variable Augmentation Method for Image Categorization with Insufficient Training SamplesabstractOver the past few years, we have made great progress in image categorization based on convolutional neural networks (CNNs). These CNNs are always trained based on a large-scale image data set; however, people may only have limited training samples for training CNN in the real-world applications. To solve this problem, one intuition is augmenting training samples. In this article, we propose an algorithm called Lavagan ( La tent V ariables A ugmentation Method based on G enerative A dversarial N ets) to improve the performance of CNN with insufficient training samples. The proposed Lavagan method is mainly composed of two tasks. The first task is that we augment a number latent variables (LVs) from a set of adaptive and constrained LVs distributions. In the second task, we take the augmented LVs into the training procedure of the image classifier. By taking these two tasks into account, we propose a uniform objective function to incorporate the two tasks into the learning. We then put forward an alternative two-play minimization game to minimize this uniform loss function such that we can obtain the predictive classifier. Moreover, based on Hoeffding’s Inequality and Chernoff Bounding method, we analyze the feasibility and efficiency of the proposed Lavagan method, which manifests that the LV augmentation method is able to improve the performance of Lavagan with insufficient training samples. Finally, the experiment has shown that the proposed Lavagan method is able to deliver more accurate performance than the existing state-of-the-art methods. Luyue Lin, Xin Zheng 0001, Bo Liu 0002, Wei Chen 0103, Yanshan Xiao |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | Subgraph feature extraction based on multi-view dictionary learning for graph classification
Xin Zheng 0001, Shouzhi Liang, Bo Liu 0002, Xiaoming Xiong, Xianghong Hu 0001, Yuan Liu 0022 |
Knowl. Based Syst. | 1 |
| 2020 | An efficient multi-label learning method with label projection
Luyue Lin, Bo Liu 0002, Xin Zheng 0001, Yanshan Xiao, Zhijing Liu |
Knowl. Based Syst. | 3 |
| 2020 | A multi-task transfer learning method with dictionary learning
Xin Zheng 0001, Luyue Lin, Bo Liu 0002, Yanshan Xiao, Xiaoming Xiong |
Knowl. Based Syst. | 1 |
| 2020 | The Software/Hardware Co-Design and Implementation of SM2/3/4 Encryption/Decryption and Digital Signature SystemabstractThe security of smart devices is facing great challenges. This article presents a new hybrid cipher framework suitable for such devices. Using software/hardware (SW/HW) co-design method, an efficient encryption/decryption and digital signature scheme based on SM2, SM3, and SM4 algorithms is implemented. First, the framework is partitioned into software and hardware parts based on the analysis result of a pure software solution and the hardware overhead under the conditions of design constraints. Second, the flow of the whole scheme and embedded algorithms are realized by SW/HW co-design. Finally, an improved implementation of SM2/3/4 algorithms is proposed to achieve higher efficiency. In this implementation, some SW/HW modules are parallelized to reduce the running time and enhance the performance. The proposed design is safe and can resist simple power analysis (SPA) attacks. Especially, the AHB bus interface IP and software scheduling approach are adopted in data transfer and manipulation. The design is taped-out on a silicon chip with SMIC 110-nm technology process. The chip uses about 199K logic gates and 1-mm2areas. The operating frequency of the design is 36 MHz and the chip consumes 23-mW power. The comparison with similar previous works shows our proposed design is more efficient with speed increasing of more than 10%. Xin Zheng 0001, Chongyao Xu, Xianghong Hu 0001, Yun Zhang 0001, Xiaoming Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |