EDBT 2026 Demo / reviewers in the wild / expert
Xuecheng Zou
dblp:90/3590
· DBLP profile ↗
23ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-6404-5270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel PFM Control Chip with Model-Based Duty Ratio Prediction and vds-Sensed Fine Tuning for Optimal ZVS in VHF Resonant SEPIC Converters
Desheng Zhang 0003, Run Min, Qiaoling Tong, Jianming Lei, Xuecheng Zou |
ISCAS | 6 |
| 2026 | A Multimode Built-In Self-Test Circuit for Sequential Cell Timing Characteristics Based on Digital-to-Time ConverterabstractTo address the growing challenges in high-precision timing characterization, where conventional measurement methods face limitations due to increasing process variations and multicharacteristic testing demands, this work presents an innovative multimode built-in self-test (BIST) circuit based on a digital-to-time converter (DTC). This circuit comprises a multimode design under test (DUT) module, a DTC, and a controller, enabling the simultaneous measurement of critical timing characteristics (setup/hold times and CK–Q/QN delays) across multiple sequential cells with varying trigger types and clock edges. To enhance the measurement performance, we propose a DTC architecture featuring a digitally controlled multilevel tunable delay cell within a two-stage “coarse–fine” delay chain and a multilevel switching mechanism. The BIST circuit achieves a minimum resolution of 18.7 ps, with the DTC offering a 10.5-ns dynamic range and high linearity, occupying 0.03 mm2, while the total BIST area is 0.58 mm2. Experimental results in a 180-nm process show setup/hold time measurement errors of −21 to 27 and −27 to 25 ps, respectively, and a CK–Q(QN) delay error range of −22.3 to 30.7 ps. Wenwen Cai, Yanhui Zhao, Zhengrui Chen, Xuecheng Zou, Cheng Zhuo, Li Zhang 0021 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2026 | TRO-Based Dual-Domain Voltage Co-Regulation of Digital Logic and SRAM in SoCsabstractConventional low-power system-on-chips (SoCs) commonly regulate digital logic using adaptive voltage and frequency scaling (AVFS), while SRAM voltage is managed by an independent scaling policy. This separation leaves cross-domain SRAM-related critical paths over-margined and prevents system-level energy optimization. This brief presents a unified voltage co-regulation framework that jointly tunes the digital and SRAM supply rails to minimize total SoC energy. A tunable replica oscillator (TRO) is repurposed from a digital timing monitor into a dual-mode delay allocator: its programmable level sets the digital timing slack, and an all-digital AVFS loop adjusts the digital supply to lock the target frequency. The released cycle budget is then converted into SRAM voltage reduction, determined by a cache-based canary test. Domain-level power is profiled on-chip to enable a measurement-driven search for the dual-domain minimum-energy point (DD-MEP) without relying on PVT-dependent model parameters. Silicon results from a taped-out 110-nm Cortex-M3 SoC demonstrate up to 11.4% total energy reduction compared with digital-only AVFS, with only 0.05% area overhead. Zhaoxu Wang, Mingyang Gong, Zhenglin Liu, Xuecheng Zou |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector CoreabstractVector processing has evolved from early systems like the CDC STAR-100 and Cray-1 to modern ISAs like ARM's Scalable Vector Extension (SVE) and RISC-V Vector (RVV) extensions.However, scaling vector processing for contemporary workloads presents challenges due to overheads in traditional architectures.We introduce Titan-I (T1), an out-of-order (OoO) RVV architecture designed Jiuyang Liu, Qinjun Li, Yunqian Luo, Jiongjia Lu, Shupei Fan, Jianhao Ye, Yanqi Yang, Zewen Ye, Yuhang Zeng, Wei Cong, Xuecheng Zou, Mingyu Gao 0001 |
MICRO | 16 |
| 2023 | ML-Accelerated Yield Analysis Framework Using Regularization for Sparsity in High-Sigma and High-Dimensional ScenariosabstractHighly repetitive structures in IC, such as SRAM cells typically require extremely low failure ratio, making traditional Monte Carlo analysis extremely time consuming. Furthermore, the “curse of dimensionality” has become a major challenge for existing high-sigma yield analysis techniques. Thus, we propose a “sampling-training-substitution-verification” (STSV) yield analysis framework, which utilizes machine learning (ML) techniques to accelerate yield analysis in high-sigma and high-dimensional scenarios, effectively addresses the “curse of dimensionality.” In our framework, least absolute shrinkage and selection operator (Lasso) regression is adopted to substitute the mapping from process parameters to circuit performance, achieving high accuracy, and generalization. The model is adaptive for both low- and high-dimensional scenarios since the dimensional sparsity is achieved by$l1$regularization. In addition, important process parameters can be identified by sparse feature weights of the Lasso model, which is of assistance for yield optimization. Compared with existing yield analysis techniques, the Lasso-based STSV framework offers great saving in a simulation program with integrated circuit emphasis (SPICE) cost, is attractive in high-dimensional demands. Haoran Fan, Bo Jiang 0018, Jianfei Chen 0003, Qiaoling Tong, Xuecheng Zou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | A Flexible and High-Performance Lattice-Based Post-Quantum Crypto Secure CoprocessorabstractProgress of quantum computing technology seriously threaten the industrial information security based on traditional public-key cryptosystem. Thus, the cryptosystem with anti-quantum attack characteristics is gradually becoming a significant research in the security field. In this article, a flexible and high-performance secure coprocessor is designed for security in industrial processes, which can execute the post-quantum cryptographic algorithm Saber efficiently. Custom instruction set and arithmetic accelerators are proposed to effectively optimize the flexibility of system architecture, and improve the performance of calculation. The hardware implementation results show that the maximum operating frequency of the coprocessor can reach 345 MHz. Compared with related state-of-the-art works, it achieves the highest operating frequency on the same Xilinx UltraScale+ FPGA platform, performing the encryption and decryption operations within 13.5 and 15.4μs, respectively. Meanwhile, this article achieves 1.7/3.1/5.9× area-time product improvements in look-up table flip-flop block memory storage with good flexibility. Dongsheng Liu 0001, Xiang Li 0220, Xingjie Liu, Jiahao Lu 0002, Xuecheng Zou, Ang Hu, Tianming Ni |
IEEE Trans. Ind. Informatics | 8 |
| 2022 | An Energy-Efficient SIFT Based Feature Extraction Accelerator for High Frame-Rate Video ApplicationsabstractVisual feature extraction is a key technology of computer vision for intelligent video processing. Efficient feature extraction is a fundamental problem in computer vision applications. Scale-Invariant Feature Transform (SIFT) is one of the most popular feature extraction algorithms because SIFT features are invariant to image scale and rotation and robust to changes in illumination and noise. However, SIFT is a computationally-intensive and power-hungry algorithm, which needs to be accelerated by efficient hardware design to achieve both high-speed feature extraction and high energy efficiency for many high frame-rate video applications at Artificial-intelligent Internet of Things edges. In this work, an energy-efficient SIFT based feature extraction accelerator is proposed. In the Gaussian pyramid and Differences of Gaussian (DoG) pyramid construction process, three design methods are proposed to reduce power consumption and improve information fidelity: a fast and slow dual clock domain design method with a reconfigurable design strategy is proposed to reduce the computation resources; a partial sum reuse design method is proposed to further reduce the computation resources and the amount of computation; a dynamic padding design method is proposed to solve the problem of information loss at image edges and corners after convolution operation. In the keypoint descriptor generation process, an optimized algorithm using circular region and polar coordinates is proposed to parallelize the main orientation assignment and descriptor generation to achieve high-speed processing, while maintaining a comparable matching accuracy with the state-of-the-art designs. The experiment results show that the proposed SIFT hardware accelerator is able to extract features by up to 162 frames per second ($640\times 480$pixels) under 100 MHz, with the power consumption of 364.26 mW and energy efficiency of 2.25 mJ/frame based on 180 nm technology, which is suitable for many high frame-rate AIoT applications including autonomous driving cars and unmanned aerial vehicles. Bingqiang Liu, Zehua Yin, Xvpeng Zhang, Xiaofeng Hu, Guoyi Yu, Yuanjin Zheng, Chao Wang 0096, Xuecheng Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2022 | An Efficient Unstructured Sparse Convolutional Neural Network Accelerator for Wearable ECG Classification DeviceabstractConvolution neural network (CNN) with pruning techniques has shown remarkable prospects in electrocardiogram (ECG) classification. However, efficiently deploying the existing pruned neural network to wearable devices for ECG classification is a great challenge due to the limited hardware resource and randomly distributed sparse weights. To address this issue, an efficient unstructured sparse CNN accelerator is proposed in this paper. A tile-first dataflow with compressed data storage format is presented to skip zero weight multiplications and increase the computing efficiency during inference of small-scale model with large sparsity. The two-level weight index matching structure in the dataflow exploits shifting operation to select valid data pairs and maintain the fully-pipelined calculation process. A configurable processing element (PE) array with 32-bit instruction control is proposed to increase the flexibility of the accelerator. Verified in FPGA and post-synthesis simulations in SMIC 40nm process, the proposed sparse CNN accelerator consumes$3.93~\mu $J/classification at 2MHz clock frequency and it achieves an averaged ECG classification accuracy of 98.99%. A computing efficiency of 118.75% is realized which is improved by 48% compared to the dense baseline. In brief, the proposed efficient CNN accelerator is especially suitable for wearable ECG classification device. Jiahao Lu 0002, Dongsheng Liu 0001, Ang Hu, Xuecheng Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Towards Efficient Hardware Implementation of NTT for Kyber on FPGAsabstractKyber is a promising lattice-based post-quantum cryptography (PQC) for key encapsulation mechanisms. Since number theoretic transform (NTT) is the most computationally expensive operation in Kyber, this paper focuses on novel optimization techniques for efficient hardware implementation of NTT for Kyber. Benefiting from the proposed fast modular multiplication method and a doubled bandwidth ping-pong memory access scheme, our NTT architecture can complete an NTT operation in Kyber in 490 cycles using only 609 LUTs, 640 FFs, 2 DSPs on a Xlinx Artix-7 FPGA. The proposed NTT architecture is 3.95 times faster than the state-of-the-art design for Kyber while achieves an improvement of more than 1.5 times in the area time product (ATP). Compared with the state-of-the- art NTT designs for other algorithms, our NTT architecture reduces 24% FFs and 50% DSPs and ranks second smallest in ATPs, which can also confirm the high efficiency of our design. Dongsheng Liu 0001, Xingjie Liu, Xuecheng Zou, Guangda Niu, Bo Liu 0063, Quming Jiang |
ISCAS | 4 |
| 2021 | Controlled nano-cracking actuated by an in-plane voltage
Xuecheng Zou, Long You |
Sci. China Inf. Sci. | 5 |
| 2021 | A 0.20-2.43 GHz fractional-N frequency synthesizer with optimized VCO and reduced current mismatch CPabstractA 0.20–2.43 GHz fractional- N frequency synthesizer is presented for multi-band wireless communication systems, in which the scheme adopts low phase noise voltage-controlled oscillators (VCOs) and a charge pump (CP) with reduced current mismatch. VCOs that determine the out-band phase noise of a phase-locked loop (PLL) based frequency synthesizer are optimized using an automatic amplitude control technique and a high-quality factor figure-8-shaped inductor. A CP with a mismatch suppression architecture is proposed to improve the current match of the CP and reduce the PLL phase errors. Theoretical analysis is presented to investigate the influence of the current mismatch on the output performance of PLLs. Fabricated in a TSMC 0.18-µm CMOS process, the prototype operates from 0.20 to 2.43 GHz. The PLL synthesizer achieves an in-band phase noise of −96.8 dBc/Hz and an out-band phase noise of −122.8 dBc/Hz at the 2.43-GHz carrier. The root-mean-square jitter is 1.2 ps under the worst case, and the measured reference spurs are less than −65.3 dBc. The current consumption is 15.2 mA and the die occupies 850 µm×920µm. Daming Ren, Xuecheng Zou |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2021 | Efficient Hardware Architecture of Convolutional Neural Network for ECG Classification in Wearable Healthcare DeviceabstractNowadays, with the increasing shortage of traditional medical resources, the existing portable monitoring healthcare device is no longer satisfactory. Thus, wearable healthcare device with diagnostic capability is becoming much more desirable. However, the design of wearable healthcare device faces the challenge of limited hardware resource and high diagnostic accuracy. In this paper, an efficient hardware architecture is proposed to implement a 1-D CNN with global average pooling (GAP) specially for embedded electrocardiogram (ECG) classification. The GAP is implemented by substituting division into shifting operation without extra computing resource consumption and it can largely reduce the parameters of the network. The fully pipelined processing unit (PU) array is designed to increase computing efficiency. A sign bit based dynamic activation strategy is developed for removing redundant multiplications and resource consumption of ReLU. The proposed efficient hardware architecture is implemented on Xilinx Zynq ZC706 board and achieves an average performance of 25.7 GOP/s under 200-MHz with resource consumption of 1538 LUT, which makes resource efficiency improved by more than 3× compared with non-optimized case. The averaged classification accuracy of five ECG beats classes is 99.10%. In brief, the proposed efficient hardware design is prospective for wearable healthcare device especially in ECG classification area. Jiahao Lu 0002, Dongsheng Liu 0001, Xuecheng Zou, Bo Liu 0063 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2021 | A SC PUF Standard Cell Used for Key Generation and Anti-Invasive-Attack ProtectionabstractBy using metal blocks as the protective coating, placing the sensitive signals in last but second metal (LSM), integrating a low-cost one-time programming (OTP) cell in each PUF unit, the proposed switched-capacitor (SC) PUF can both provide sensitive anti-invasive-attack protective coating and stable key for the security chip. Moreover, the circuit parameters and the layout implementation of the SC PUF unit are all compatible with other digital standard cells, which greatly facilitates the integration of SC PUF unit in the security chip by using digital design flow when its function, timing, power, and layout views are characterized using commercial timing and layout extraction tools. The anti-invasive-attack ability, stability, and digital design flow compatibility of the proposed SC PUF standard cell are verified in a security chip by using a standard 0.18- μm CMOS process. The measured bit error rate, bias, average intra-die HD, and average inter-die HD of output keys after OTP is-4, 46.72%, 0%, and 50.38% respectively. Finally, the failed probing and destruction attack attempts to the coating also verify the invasive-attack-resistant property of the proposed SC PUF standard cell. With the help of SC PUF standard cell, the whole security chip can easily obtain stable keys and sensitive anti-invasive-attack ability by using digital design flow. Zhangqing He, Meilin Wan, Jiuyang Liu, Haoshuang Gu, Xuecheng Zou |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2019 | Superposed Compensation Strategy to Optimize Load/Line Transient Response and Reference Tracking for Discontinuous Conduction Mode Boost ConverterabstractFor boost converter operating in the discontinuous conduction mode (DCM), feedback and feedforward compensators are widely used to improve the converter performance. However, it is relatively difficult to optimize load transient response (LoTR), line transient response (LiTR), and reference tracking speed (RTS) simultaneously, since the optimizations require different compensators that are incompatible. In order to solve the issue, a superposed compensation strategy is proposed in this paper, which consists of a feedback compensator and two feedforward compensators. Each compensator is tuned according to an objective transfer function, which optimizes LoTR, LiTR, and RTS. The outputs are summed as duty cycle according to the linear superposition principle. Compatibility of the compensators is improved by designing the feedforward compensators to adapt to the feedback compensator. Furthermore, based on the closed-loop model, design rules for the objective transfer functions are given to minimize the influences of the sample-and-hold effect and calculation delay, which are intrinsic in a digital controller. Finally, converter's LoTR, LiTR, and RTS are simultaneously optimized, which is proven by closed-loop magnitude-frequency plots, state trajectory analyses, and experimental results. Run Min, Dian Lyu, Linkai Li, Qiaoling Tong, Xuecheng Zou, Zhenglin Liu |
IEEE Trans. Ind. Informatics | 6 |
| 2018 | A Time-Division-Multiplexing Scheme for Simultaneous Wavelength Locking of Multiple Silicon Micro-RingsabstractThis paper presents a time-division-multiplexing (TDM) scheme for simultaneous wavelength locking of multiple silicon micro-rings by exploiting the speed mismatch between the heater and the controller. This scheme could reduce the overall chip size significantly without reducing the wavelength lock speed. It also avoids cross-channel coupling using the least number of ADCs and DACs. The simplest TDM scheme involving two micro-rings is experimentally verified using board-level circuits. Theoretically, this approach can be scaled to even hundreds of micro-rings, pointing out a way towards large-scale integrated optoelectronics, which is required by many important applications, such as wavelength division multiplexing for chip-to-chip optical I/O. Zhicheng Wang 0008, Yu Yu 0005, Xi Xiao 0004, Miaofeng Li, Xuecheng Zou, Dingshan Gao, Min Tan 0004 |
ISCAS | 5 |
| 2017 | Chaotic Encrypted Polar Coding Scheme for General Wiretap ChannelabstractA wiretap channel is an important model for wireless communication. By applying an extended multiblock polar coding scheme, recent literature has achieved the secrecy capacity of a general wiretap channel (not necessary degraded or symmetric). However, this secure polar coding scheme of physical layer also limits the transmission rate of the main channel, which may fail to meet the demand of high transmission rate and strong transmission security for practical wireless transmission. In order to obtain a higher secrecy transmission rate than the physical layer coding scheme over a general wiretap channel, a cross-layer encryption and coding scheme is proposed in this paper. In the proposed scheme, an onetime-pad encryption and a secure key transmission is constructed by combining a chaos stream cipher with the extended multiblock polar coding scheme. As proved, the proposed scheme has achieved a high secrecy transmission rate than the former physical layer coding scheme under the constraints of reliability and strong security for a general wiretap channel. Yizhi Zhao, Xuecheng Zou, Zhaojun Lu, Zhenglin Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A Novel Thyristor-Based Silicon Physical Unclonable FunctionabstractThis paper describes a new silicon physical unclonable function (PUF) using thyristor components that can be fabricated on a standard CMOS process. Our proposed design is built using thyristor-based sensors, time difference amplifier (TDA), time difference comparator, voting mechanism, and diffusion algorithm circuit. Multiple identical thyristor-based sensors are fabricated on the same chip. Due to the manufacturing process variations, each sensor produces two slightly different delay values that can be compared in order to create a digital identification for the chip. Diffusion algorithm circuit further ensures that the proposed PUF is able to effectively identify a population of integrated circuits. We also improve the stability of PUF design with respect to temporary environmental variations, such as temperature and supply voltage with the introduction of TDA and voting mechanism. The thyristor-based PUF is fabricated in a 0.18-$\mu \text{m}$ CMOS technology. Experimental results show that the PUF has a good output statistical characteristic of a uniform distribution and a high stability of 96.8% with respect to temperature variation from -40 °C to 100 °C, and supply voltage variation from 1.7 to 1.9 V. Chuang Bai, Xuecheng Zou, Kui Dai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Hardware IP Protection through Gate-Level ObfuscationabstractHardware Intellectual Property (IP) cores have emerged as an integral part of modern System-on-Chip (SoC) designs. However, recent trends of reverse engineering pose major threat to IP-based SOC design flow. The paper proposes a novel approach for hardware IP protection using gate-level obfuscation, which could make design less intelligible in order to neutralize or weaken the effect of reverse engineering. The basic idea is to hide the original logic function by using Physical Unclonable Function (PUF), multiplexer and configurable logic, so that it is difficult for reverse engineering attackers to get complete information of circuit net list. The design methodology could be applied in combinational logic and sequential logic. Simulation results on several IP cores show that we can achieve high levels of security through a well-formulated obfuscation scheme at less than 10% area overhead under delay constraint. Wenchao Liu 0003, Xuecheng Zou, Zhenglin Liu |
CAD/Graphics | 3 |
| 2014 | A Compact Hardware Implementation of SM3 Hash FunctionabstractWith mobile and wireless devices becoming pervasive, low-cost hardwares of security functions are being desired. A compact hardware implementation of the SM3 hash algorithm is presented in this paper. A SRAM is used to do message expansion function instead of shift registers which are used in common hardware implementations, and the values of A~H and V0~V7 registers are updated in the serial shift way when they are initialized and updated. The computation units are saved as much as possible. Compared with traditional designs, the store resources for message expansion function can be shared with other modules to reduce the cost of a system. The Synopsys' DC synthesis results show that the area of the compact SM3 is approximate 8277 GEs while its throughput can be as high as 276 Mbps. If the SRAM is shared with other modules, only 6904 GEs are required to implement the SM3 hardware module in the system. The compact architecture can be accommodated to resource-constrained systems for its advantages of low-cost and low-power. Tianyong Ao, Zhangqing He, Jinli Rao, Kui Dai, Xuecheng Zou |
TrustCom | 5 |
| 2011 | Implementation and evaluation of parallel FFT on Engineering and Scientific Computation Accelerator (ESCA) architectureabstractThe fast Fourier transform (FFT) is a fundamental kernel of many computation-intensive scientific applications. This paper deals with an implementation of the FFT on the accelerator system, a heterogeneous multicore architecture to accelerate computation-intensive parallel computing in scientific and engineering applications. The Engineering and Scientific Computation Accelerator (ESCA) consists of a control unit and a single instruction multiple data (SIMD) processing element (PE) array, in which PEs communicate with each other via a hierarchical two-level network-on-chip (NoC) with high bandwidth and low latency. We exploit the architecture features of ESCA to implement a parallel FFT algorithm efficiently. Experimental results show that both the proposed parallel FFT algorithm and the ESCA architecture are scalable. The 16-bit fixed-point parallel FFT performance of ESCA is compared with a published work to prove the superiority of the mapping algorithm and the hardware architecture. The floating-point parallel FFT performances of ESCA are evaluated and compared with those of the IBM Cell processor and GPU to demonstrate the computing power of the ESCA system for high performance applications. Dan Wu 0009, Xuecheng Zou, Kui Dai, Jinli Rao, Zhaoxia Zheng |
J. Zhejiang Univ. Sci. C | 2 |
| 2010 | A High Efficient On-Chip Interconnection Network in SIMD CMPs
Dan Wu 0009, Kui Dai, Xuecheng Zou, Jinli Rao |
ICA3PP (1) | 3 |
| 2010 | A Methodology for Design of Unbuffered Router Microarchitecture for S-Mesh NoC
Feifei Cao, Dongsheng Liu 0001, Xuecheng Zou |
NPC | 4 |
| 2007 | On the Ability of AES S-Boxes to Secure Against Correlation Power Analysis
Zhenglin Liu, Xuecheng Zou |
ISPEC | 5 |