VLDB 2026 Research / reviewers in the wild / expert
Zheng Wang 0027
dblp:w/ZhengWang27
· DBLP profile ↗
24ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0003-2855-9570ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NMSAcc: An Arch-Decoupled Non-Maximum Suppression Accelerator on Edge Devices
Zheng Wang 0027, Qiawei Zheng, Jinghan Zhou, Chao Chen 0022 |
ISCAS | 2 |
| 2026 | Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoCabstractWhile recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation (TM)—centered on high-volume data movement with minimal computation—remains underexplored. This work addresses that gap by introducing the TM unit (TMU): a reconfigurable, near-memory hardware block designed to execute data-movement-intensive (DMI) operators efficiently. The TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling support for both a wide range of coarse- and fine-grained tensor transformations. The proposed architecture integrates the TMU alongside a TPU within a high-throughput AI system-on-chip (SoC), leveraging double buffering and output forwarding to improve pipeline utilization. The TMU, synthesized under the SMIC 40-nm standard cell library, occupies only$0.019~\mathrm {\text {mm}^{2}}$while supporting over 10 representative DMI operators. Benchmarking shows that the TMU alone achieves up to$82.42\times $and$11.06\times $operator-level latency reduction over ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 22.89% reduction in end-to-end inference latency, demonstrating the effectiveness in reducing inference latency and the scalability of the TMU architecture across diverse tensor operators. Weiyu Zhou, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Zhuoyu Wu, Anupam Chattopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody DesignerabstractAntibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery. Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001 |
AAAI | 8 |
| 2025 | ProtCLIP: Function-Informed Protein Multi-Modal LearningabstractMulti-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model. Hanjing Zhou, Mingze Yin, Wei Wu 0045, Kun Fu 0002, Jintai Chen, Jian Wu 0001, Zheng Wang 0027 |
AAAI | 8 |
| 2025 | De2r: Unifying DVFS and Early-Exit for Embedded AI Inference via Reinforcement LearningabstractExecuting neural networks on resource-constrained embedded devices faces challenges. Efforts have been made at the application and system levels to reduce the execution cost. Among them, the early-exit networks reduce computational cost through intermediate exits, while Dynamic Voltage and Frequency Scaling (DVFS) offers system energy reduction. Existing works strive to unify early-exit and DVFS for combined benefits on both timing and energy flexibility, yet limitations exist: 1) varying time constraints that make different exit points become more, or less, important in terms of inference accuracy, are not taken care of, and 2) the optimal decisions of unifying DVFS and early-exit as a multi-objective optimization problem are not achieved due to the large configuration space. To address these challenges, we propose Dr2r, a reinforcement learning-based framework that jointly optimizes early-exit points and DVFS settings for continuous inference. In particular, Dr2r includes a cross-training mechanism that fine-tunes the early-exit network to accommodate dynamic time constraints and system conditions. Experimental results demonstrate that Dr2r achieves up to 22.03% energy reduction and 3.23% accuracy gain compared to contemporary techniques. Yuting He 0002, Jingjin Li, Chengtai Li, Qingyu Yang 0004, Zheng Wang 0027, Heshan Du, Jianfeng Ren, Heng Yu 0001 |
DATE | 5 |
| 2025 | TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionabstractRapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications.However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by the quadratic computational complexity of attention.These issues limit LLMs in long-context scenarios.In this paper, we propose Dynamic Token-Level KV Cache Selection (TokenSelect), a training-free method for efficient and accurate long-context inference.TokenSelect builds upon the observation of non-contiguous attention sparsity, using QK dot products to measure per-head KV Cache criticality at tokenlevel.By per-head soft voting mechanism, To-kenSelect selectively involves a few critical KV cache tokens in attention calculation without sacrificing accuracy.To further accelerate To-kenSelect, we design the Selection Cache based on observations of consecutive Query similarity and implemented the efficient Paged Dot Product Kernel, significantly reducing the selection overhead.A comprehensive evaluation of To-kenSelect demonstrates up to 23.84× speedup in attention computation and up to 2.28× acceleration in end-to-end latency, while providing superior performance compared to state-of-theart long-context inference methods. Wei Wu 0045, Zhuoshi Pan, Kun Fu 0002, Chao Wang 0086, Liyi Chen 0001, Yunchu Bai, Tianfu Wang 0002, Zheng Wang 0027, Hui Xiong 0001 |
EMNLP | 8 |
| 2025 | AttenPU: An Area Efficient Attention Processor with Reconfigurable FP8 Precision and DataflowabstractEfficient numerical representation is crucial for deep learning accelerators, especially for large language models (LLMs). The 8-bit-floating-point (FP8) data representation achieves higher precision and fewer quantization efforts than integer, which has been proven inevitable in attention-based accelerators for LLMs. Therefore, area-efficient design techniques for FP8 play a central role in lowering LLMs chip’s budget. This paper presents AttenPU, which is built upon reconfigurable FP8 units and supports E4M3 for inference and E5M2 for training. Bidirectional dataflow is exploited to enable AttenPU to interact with FP32 coprocessor to reduce latency. The design achieves a low FP8-to-INT8 area ratio of 1.63×, an area efficiency of 193.5 GFLOPS/mm2, with an 87.05% reduction in the latency of RTX3090 GPU. Qiawei Zheng, Zheng Wang 0027, Zhuoyu Wu, Zhihao Du, Chao Chen 0022, Yongkui Yang, Wenqi Fang, Anupam Chattopadhyay |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | PDR-KAN: Pipeline-Driven Reconfigurable Accelerator for Kolmogorov-Arnold Networks with Cross-Mode Sparsity SupportabstractThe commonly used Multi-layer Perceptrons (MLPs) in modern AI applications significantly limit the real-time performance due to their intensive memory access operation. Recently, Kolmogorov–Arnold Networks (KANs) have gained attention for offering a similar structure to MLPs but with a more efficient parameter utilization. However, the lack of customized support in conventional hardware restricts the performance gain from its algorithmic superiority. What’s more, the early-stage development of KAN raises questions about the necessity of incurring additional hardware costs for it. In this work, we present PDR-KAN, a reconfigurable accelerator that features two distinct operating modes, one for KANs and one for MLPs. Apart from its pipeline mode and sparsity encoder (SE) for efficient KAN processing, PDR-KAN can also improve the throughput of MLPs with cross-mode sparsity support. Experiments on a real-world dataset show that PDR-KAN provides a 3.95× acceleration and 13% reduction in accuracy loss by simply replacing MLPs with KANs. For a more accurate KAN model version with 3.33× parameter scaling, the latency overhead on PDR-KAN is only 1.16× compared to the baseline KAN model. Additionally, PDR-KAN achieves a 6.73× speed-up in KAN inferences and 20.89× in energy efficiency, against Quad-core ARM Cortex-A72 CPU. Wenhui Ou, Zhuoyu Wu, Alexandra Geciova, Zheng Wang 0027, C. Patrick Yue |
ISCAS | 5 |
| 2025 | ZergPPU: an evolutionary multi-mode post processor for vision neural networksabstractWith the widespread adoption of AI in various industries, there is increasing demand for edge devices to efficiently utilize AI, particularly in object detection and related applications. Current studies aim to accelerate specific computations but struggle to keep up with rapid advancements in neural network optimization techniques like quantization and pruning. So the goal of this work is to propose a microarchitecture approach that can flexibly support multiple generations of visual neural networks and dynamic data shapes.This work presents the Zerg Post Processing Unit (ZergPPU), which determines its computational architecture based on data information, partitioning into regions for efficient processing. With software support, ZergPPU adapts to evolving AI algorithms through multi-mode processing (e.g., from YOLOv3 to YOLOv9). By incorporating a path prediction unit and dynamic address generation, it supports a wide range of neural networks and optimizations. To minimize area, ZergPPU uses selective reuse executors with automatic software optimizations, resulting in an area of 0.064mm2. Experiments show that this design achieves 49.3% lower computational latency compared to the baseline, making it well-suited for edge AI vision applications. Weilun Wang, Chen Chao, Zheng Wang 0027 |
ISCAS | 3 |
| 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMsabstractProteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge. Wei Wu 0045, Chao Wang 0086, Liyi Chen 0001, Mingze Yin, Yiheng Zhu 0002, Kun Fu 0002, Jieping Ye, Hui Xiong 0001, Zheng Wang 0027 |
KDD (2) | 9 |
| 2025 | Less is More: Improving LLM Alignment via Preference Data SelectionabstractDirect Preference Optimization (DPO) has emerged as a promising approach for aligning large language models with human preferences. While prior work mainly extends DPO from the aspect of the objective function, we instead improve DPO from the largely overlooked but critical aspect of data selection. Specifically, we address the issue of parameter shrinkage caused by noisy data by proposing a novel margin-maximization principle for dataset curation in DPO training. To further mitigate the noise in different reward models, we propose a Bayesian Aggregation approach that unifies multiple margin sources (external and implicit) into a single preference probability. Extensive experiments in diverse settings demonstrate the consistently high data efficiency of our approach. Remarkably, by using just 10\% of the Ultrafeedback dataset, our approach achieves 3\% to 8\% improvements across various Llama, Mistral, and Qwen models on the AlpacaEval2 benchmark. Furthermore, our approach seamlessly extends to iterative DPO, yielding a roughly 3\% improvement with 25\% online data, revealing the high redundancy in this presumed high-quality data construction manner. These results highlight the potential of data selection strategies for advancing preference optimization. Rui Ai 0004, Fuli Feng, Zheng Wang 0027, Xiangnan He 0001 |
NeurIPS | 5 |
| 2024 | PerFT-N: Low-overhead Permanent Fault-Tolerance Mechanism for Neural Processing UnitsabstractThe reliability of neural network (NN) accelerators is one key to ensuring inference accuracy. Relying solely on the robustness of NN can only tolerate transient faults, and once the circuit encounters permanent faults, it will lead to a serious accuracy drop. We, by employing processing elements (PEs) checking and re-scheduling of threads of functional units, propose a new fault-tolerant mechanism named PerFT-N for neural processing units (NPUs). Compared to previous NPU fault-tolerant methods, PerFT-N recovers the execution facing permanent faults with minimal hardware overhead. Specifically, we utilize the existing resources to achieve fault detection and localization, while achieving recovery by re-linking unfailed threads. Furthermore, an instruction-based programmable permanent fault emulation scheme is deployed on the FPGA platform for fast verification. Experimental results demonstrate that the PerFT-N architecture works in the extreme case with 98.5% of failed functional units while incurring the physical overheads of 2.7% in area and 3.6% in power consumption under 40nm SMIC standard cell library. Haojie Jian, Chao Chen 0022, Zheng Wang 0027 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | A3S: A General Active Clustering Method with Pairwise ConstraintsabstractActive clustering aims to boost the clustering performance by integrating human-annotated pairwise constraints through strategic querying. Conventional approaches with semi-supervised clustering schemes encounter high query costs when applied to large datasets with numerous classes. To address these limitations, we propose a novel Adaptive Active Aggregation and Splitting (A3S) framework, falling within the cluster-adjustment scheme in active clustering. A3S features strategic active clustering adjustment on the initial cluster result, which is obtained by an adaptive clustering algorithm. In particular, our cluster adjustment is inspired by the quantitative analysis of Normalized mutual information gain under the information theory framework and can provably improve the clustering quality. The proposed A3S framework significantly elevates the performance and scalability of active clustering. In extensive experiments across diverse real-world datasets, A3S achieves desired results with significantly fewer human queries compared with existing methods. Junlong Liu, Fuli Feng, Chen Shen 0003, Xiangnan He 0001, Jieping Ye, Zheng Wang 0027 |
ICML | 8 |
| 2024 | Dual-Branch StarNet with Mutual Attention and U-Net Denoising for Simultaneously Recognizing Keywords and Speakers
Yuting He 0002, Chengtai Li, Heng Yu 0001, Jianfeng Ren, Zheng Wang 0027, Heshan Du, Yinshui Xia |
ICONIP (5) | 5 |
| 2024 | Low-latency Buffering for Mixed-precision Neural Network Accelerator with MulTAP and FQPipeabstractPrevious work has proposed precision scalable accelerators to handle mixed-precision neural network (NN) inferences on the edge, which focus on designing reconfigurable MAC arrays while leaving the issue of time-costly data buffering procedure less discussed. Besides, integer-only inference is incapable of handling emerging NN models with various non-linear activation functions. In this work, we propose a mixed-precision NN accelerator supporting int8, int16 and fp32 arithmetic with two buffering techniques namely MulTAP and FQPipe, which jointly facilitate low-latency data movement. Experiment results show that MulTAP and FQPipe boost the baseline NN accelerator with 7.7 × and 1.5 × in speed respectively, which leads to the application performance of 473.9 (int8) and 252.5 (int16) inferences per second (IPS) on YOLOv3-Tiny. Post-layout netlist with SMIC 40nm standard-cell technology demonstrates a design with an area of 26.96mm2and a power estimate of 1.83W. Zheng Wang 0027, Wenhui Ou, Weiyu Zhou, Yongkui Yang, Chao Chen 0022 |
ISCAS | 2 |
| 2023 | COMPACT: Co-processor for Multi-mode Precision-adjustable Non-linear Activation FunctionsabstractNon-linear activation functions imitating neuron behaviors are ubiquitous in machine learning algorithms for time series signals while also demonstrating significant gain in precision for conventional vision-based deep learning networks. State-of-the-art implementation of such functions on GPU-like devices incurs a large physical cost, whereas edge devices adopt either linear interpolation or simplified linear functions leading to degraded precision. In this work, we design COMPACT, a co-processor with adjustable precision for multiple non-linear activation functions including but not limited to exponent, sigmoid, tangent, logarithm, and mish. Benchmarking with state-of-the-arts, COMPACT achieves a 26% reduction in the absolute error on a 1.6x widen approximation range taking advantage of the triple decomposition technique inspired by Hajduk's formula of Padé approximation. A SIMD-ISA-based vector co-processor has been implemented on FPGA which leads to a 30% reduction in execution latency but the area overhead nearly remains the same with related designs. Furthermore, COMPACT is adjustable to 46% latency improvement when the maximum absolute error is tolerant to the order of 1E-3. Wenhui Ou, Zhuoyu Wu, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang |
DATE | 3 |
| 2023 | Multi-Attention Feature Fusion Network for Accurate Estimation of Finger Kinematics From Surface Electromyographic SignalsabstractSimultaneous and proportional control (SPC) based on surface electromyographic (sEMG) signals has led to a broad range of applications. However, due to the limitation in the generalization and stability of current machine learning algorithms, these methods can only estimate less than 15 simultaneuous and proportional (SP) categories of finger movement. In this article, a novel deep learning algorithm, named multiattention feature fusion network (MAFN), is proposed to estimate comprehensive finger movement (up to 28 categories SP movements) from sEMG signals. MAFN is based on the multihead attention mechanism, which adaptively extracts essential features for analyzing the joint angles from the extracted sEMG features. Furthermore, a real-time exponential smoothing algorithm is designed for further improvement of the prediction stability. MAFN was evaluated on 28 finger movements of 38 subjects in the Ninapro_db2 dataset, and benchmarked with the state-of-the-art methods, such as temporal convolutional network (TCN) and long short term memory network (LSTM). The results demonstrated that the average Pearson correlation coefficient, root mean squared error of MAFN (0.84 ± 0.03,0.09 ± 0.01) were significantly higher than those of TCN (0.52 ± 0.06,pppp< 0.001). These improvements led to more stable and accurate movement predictions. Additionally, the time delay and power consumption of MAFN when applied to sEMG signals on a portable device are only 83.4 ms and 3 W, which implies prospective commercial applications. Weiyu Guo, Ning Jiang 0001, Dario Farina, Jingyong Su, Zheng Wang 0027, Chuang Lin 0001, Hui Xiong 0001 |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2022 | Colorization for in situ Marine Plankton Images
Guannan Guo, Zhenghui Feng, Zheng Wang 0027 |
ECCV (38) | 5 |
| 2022 | Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002 |
ECCV (32) | 4 |
| 2021 | OR-ML: Enhancing Reliability for Machine Learning Accelerator with Opportunistic RedundancyabstractReliability plays a central role in deep sub-micron and nanometre IC fabrication technology and has recently been reported to be one of the key issues affecting the inference phase of neural networks. State-of-the-art machine learning (ML) accelerators exploit massively computing parallelism observed in neural networks to achieve high energy efficiency. The topology of ML engines' computing fabric, which constitutes large arrays of processing elements (PEs), has been increasing dramatically to incorporate the huge size and heterogeneity of the rapid evolving ML algorithm. However, it is commonly observed that activations of zero value lead to reduced PE utilization. In this work, we present a novel and low-cost approach to enhance the reliability of generic ML accelerators by Qpportunistically exploring the chances of runtime Redundancy provided by neighbouring PEs, named as OR-ML. In contrast to conventional redundancy techniques, the proposed technique introduces no additional computing resources, therefore significantly reduces the implementation overhead and achieves obvious level of protection. The design prototype is evaluated using emulated fault injection on FPGA, executing mainstream neural networks for objectionclassification and detection. Zheng Wang 0027, Wenxuan Chen, Chao Chen 0022, Yongkui Yang, Zhibin Yu 0001 |
DATE | 2 |
| 2021 | CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window FilteringabstractMemory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering. Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | A Current Mirror Cross Bar Based 2.86-TOPS/W Machine Learner and PUF with <2.5% BER in 65nm CMOS for IoT ApplicationabstractEnergy-efficient machine-learning and physical unclonable function (PUF) becomes popular in Internet-of-Things (IoT) applications for saliency detection and privacy protection at sensor node. A machine-learning and PUF engine for IoT applications is presented in this work with a current mirror cross-bar (CMCB) array being a shared core, reducing silicon area. A novel dimension expansion technique is proposed to increase weight matrix dimension beyond the physically implemented array with small hardware and energy overhead. A signed multiply-and-accumulation is realized in CMCB with differential current path and 2-phase conversion. The proposed engine achieves an error rate of 6.34% on MNIST digit recognition task with an energy efficiency of 2.86 TOPS/W. The PUF achieves a native BER of 2.3% across corners and extremely low area/CRP of 4.17×10-59μm2/CRP. Yi Chen 0012, Zheng Wang 0027, Aakash Patil, Arindam Basu |
ISCAS | 2 |
| 2019 | Detecting Fault Injection Attacks Based on Compressed Sensing and Integer Linear ProgrammingabstractCryptographic ICs have been widely applied to numerous security-critical environments nowadays. Fault injection has become a serious attack on cryptographic IC, especially soft-errors or single event upsets (SEUs) by fine-resolution fault injection attacks. Detection and tamper evidence of these attacks become important. Traditional SEU diagnose methods usually require special sensors embedded into the circuits. However, these methods require non-trivial design and test effort, and usually just yield statistic results. In this paper, we formulate the detection fault injection attacks as a compressed sensing problem, due to sparsity of soft errors. Besides, due to the binary characteristic of the coefficient matrix and the variables, integer linear programming is adopted to reconstruct the soft error signals. Simulation results on a cryptographic IC demonstrate that the proposed method is capable to accurately detect the locations of soft-errors caused by fault injection attacks with negligible hardware overhead. The abnormal test output of scan-chains can be tamper evidence of the fault injection attacks. Huiyun Li, Cuiping Shao, Zheng Wang 0027 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2017 | Current mirror array: A novel lightweight strong PUF topology with enhanced reliabilityabstractIn this work, we present a novel design of physical unclonable function (PUF) based on the topology of current mirror array (CMA). The proposed strong PUF exploits the randomness in the current mirror transistors and generates the response bit by comparing the accumulated current values. Thresholding and reference current are used to increase the reliability of the proposed PUF without requiring additional hardware resources. The proposed PUF structure is also analysed in terms of its difficulty of model building for measurement-prediction attack. Measurement results on 0.35μm test chips demonstrate that the proposed PUF outperforms other state-of-the-art designs with smaller area/bit of 9 × 10−36μm2 and lower native bit error rate (BER) of 0.16%. Zheng Wang 0027, Yi Chen 0012, Aakash Patil, Chip-Hong Chang, Arindam Basu |
ISCAS | 1 |