VLDB 2026 Research / reviewers in the wild / expert
He Li 0008
dblp:05/4746-8
· DBLP profile ↗
34ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 7 first-author · 25 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HAMG: A Hierarchical Automated MemTile-based GEMM Accelerator for Versal AIE-MLabstractEfficient GEMM on AMD Versal AIE-ML requires careful coordination among AI Engines, MemTiles, and PL-side BRAM/URAM under tight off-chip bandwidth limits. We present HAMG, an automated generator that constructs a three-level memory hierarchy with replay-aware array mapping. On XCVE2302 FPGAs, HAMG reaches 2.29 TOPS for BF16 and 5.68 TOPS for INT8, delivering up to 2.8× speedup over a reproduced GAMA-style baseline. Kai Shao, Erwei Wang, He Li 0008 |
FCCM | 3 |
| 2026 | Fine-grained data integration for high throughput and bandwidth-efficient computation on FPGAs
Jiyuan Liu 0006, Baoping Wang, Yongming Tang, He Li 0008 |
Integr. | 4 |
| 2026 | Diff-Acc: An Efficient FPGA Accelerator for Unconditional Diffusion ModelsabstractThe diffusion model has achieved remarkable success in the era of Artificial Intelligence Generated Content (AIGC) across various tasks, such as image, video, text, material modeling, and molecular design. However, the diffusion model is computational intensive due to the long iteration of the reverse denoising process, which hinders its further advancement. Therefore, there is an urgent need to accelerate the diffusion model, especially in edge scenarios that require real-time computation. While researchers have made efforts to accelerate the diffusion model at the algorithm level using either efficient sampling or model quantization, they still suffer from accuracy degradation. More importantly, they have overlooked the hardware-level acceleration challenge. This work aims to bridge the gap by introducing Diff-Acc , the first FPGA accelerator for unconditional diffusion models with a novel step-wise quantization method that requires minimal calibration data to achieve the state-of-the-art (SOTA) PTQ quantization accuracy. Additionally, we adopt several hardware-oriented optimizations to reduce the computational overhead. At the architecture level, we fully analyze the computation flow of diffusion models and propose a novel architecture with group-wise parallelism to tackle the long iteration challenge. Besides, we decouple the data dependencies and adopt proper computational transformations at the micro-architecture level. Experiments on two unconditional diffusion models (DDIM and DDPM) with two image datasets (CIFAR-10 and ImageNet) demonstrate that our quantization method achieves the substantial improvements in image quality (FID: 6.67, sFID: 11.24) under 8-bit PTQ quantization. Compared with both server-based (Tesla V100 and Intel Xeon) and edge-based (Raspberry Pi 4 and Jetson Nano) platforms, Diff-Acc implemented on the Zynq UltraScale+ XCZU9EG FPGA demonstrates an up-to 12.5× energy efficiency. Particularly versus edge-based platforms, Diff-Acc achieves up to 10.26× and 1.97× performance improvements over CPU and GPU, respectively. Shidi Tang, Ruiqi Chen 0001, Yuxuan Lv, Pengwei Zheng, He Li 0008 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | CLASS: A Controller-Centric Layout Synthesizer for Dynamic Quantum CircuitsabstractLayout Synthesis for Quantum Computing (LSQC) is a critical component of quantum design tools. Traditional LSQC studies primarily focus on optimizing for reduced circuit depth by adopting a device-centric design methodology. However, these approaches overlook the impact of classical processing and communication time, thereby being insufficient for Dynamic Quantum Circuits (DQC).To address this, we introduce CLASS, a controller-centric layout synthesizer designed to reduce inter-controller communication latency in a distributed control system. It consists of a two-stage framework featuring a hypergraph-based modeling and a heuristic-based graph partitioning algorithm. Evaluations demonstrate that CLASS effectively reduces communication latency by up to 100% with only a 2.10% average increase in the number of additional operations. Yilun Zhao 0002, Bing Li 0017, He Li 0008, Mengdi Wang 0004, Yinhe Han 0001, Ying Wang 0001 |
ICCAD | 4 |
| 2025 | Scalable and Real-Time Power System Simulation Based on Heterogeneous CPU-FPGA Co-operationabstractWith the increasing integration of renewable energy devices, modern power systems have become complex, exhibiting diverse circuit topologies. Existing FPGA-based accelerators are optimized for fast single-topology simulation but lack the flexibility to handle multiple topology simulations. This paper proposes a scalable heterogeneous CPU-FPGA system designed for efficient simulation of diverse power system topologies. The proposed system enables real-time simulation for variable-scale power systems, distinguishing itself from conventional simulators by leveraging flexible matrix decomposition and a topology-aware approach. Our design specifically achieves the minimal latency of 100ns for both the boost and three-phase voltage source converter (VSC) models, highlighting its superior performance across diverse topologies. Introduced by dynamic fixed-point quantization, our proposal reduces 30.8% LUTs, 26.8% FFs, 22.2% BRAMs and 43.8% DSPs versus a fixed-point implementation with the same computational accuracy. Hangyu Yang, Jiyuan Liu 0006, Mingwang Xu, Yongming Tang, He Li 0008 |
ISCAS | 6 |
| 2025 | FPGA Accelerated Adaptive LDPC-based Quantum Error Correction by Bitwise Pipeline ParallelismabstractQuantum networks are at the forefront of next-generation communication technologies, providing unparalleled security by leveraging the fundamental properties of quantum mechanics. However, noise and interference adversely affect information sharing within quantum networks, causing quantum bit errors and diminishing the quality of quantum keys. To address this issue, high-speed quantum information error correction is necessitated and this work proposes an FPGA-accelerated error correction system based on low-density parity check (LDPC), incorporating algorithm-hardware co-optimizations. The proposed hardware architecture is applicable to generic LDPC soft decision algorithms and can adaptively select LDPC matrices with different code rates based on the quantum bit error rates. Experiment results demonstrate that the system exhibits high performance in error correction acceleration, achieving a throughput of 728 Mbps and up-to 9.4× speedup compared to existing FPGA-based LDPC error correction implementations. Bingze Ye, Jiyuan Liu 0006, He Li 0008 |
ISCAS | 3 |
| 2024 | Towards Fault-tolerant Design of Quaternary Quantum ArithmeticabstractQudit has emerged as a promising system for next-generation quantum computers because of its significant advantages in multi-phase problems and quantum error correction, while multi-qudit operations are yet to be implemented. As one of the basic operations, quantum addition has been widely applied in number factorization and discrete logarithms. Meanwhile, the physical constraints of noisy intermediate-scale quantum (NISQ) devices necessitate the development of fault-tolerant quantum adders and the exploration of quaternary operations on binary devices is still in its infancy. In this work, we propose the first approach for implementing quaternary quantum addition algorithms by employing primitive quantum gates. A library of quaternary quantum gates and quaternary quantum full adders (Q2FA) able to produce carry-first results, along with lower depth and fewer T-gates optimizations are proposed and evaluated, where all circuits are implemented on IBM Qiskit SDK. Extensive experiments show that our proposed Q2FA design, together with the optimization techniques, reduces T-depth by up to 1.4× and T-count by 1.7× compared with baseline quantum circuits without depth and T-gate optimizations. Meanwhile, the scalability of the proposed Q2FA is demonstrated by constructing quantum carry-ripple adders. Under noisy conditions, our proposed design can achieve an overall fidelity increase by 1.4×. Yunchen Zhu, Ruixuan Yang, Yuhang Gu, Fangtian Gu, Lingyi Kong, He Li 0008 |
ITC-Asia | 6 |
| 2024 | LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy PhysicsabstractThis work presents a novel reconfigurable architecture for Low Latency Graph Neural Network (LL-GNN) designs for particle detectors, delivering unprecedented low latency performance. Incorporating FPGA-based GNNs into particle detectors presents a unique challenge since it requires sub-microsecond latency to deploy the networks for online event selection with a data rate of hundreds of terabytes per second in the Level-1 triggers at the CERN Large Hadron Collider experiments. This article proposes a novel outer-product based matrix multiplication approach, which is enhanced by exploiting the structured adjacency matrix and a column-major data layout. In addition, we propose a custom code transformation for the matrix multiplication operations, which leverages the structured sparsity patterns and binary features of adjacency matrices to reduce latency and improve hardware efficiency. Moreover, a fusion step is introduced to further reduce the end-to-end design latency by eliminating unnecessary boundaries. Furthermore, a GNN-specific algorithm-hardware co-design approach is presented which not only finds a design with a much better latency but also finds a high accuracy design under given latency constraints. To facilitate this, a customizable template for this low latency GNN hardware architecture has been designed and open-sourced, which enables the generation of low-latency FPGA designs with efficient resource utilization using a high-level synthesis tool. Evaluation results show that our FPGA implementation is up to 9.0 times faster and achieves up to 13.1 times higher power efficiency than a GPU implementation. Compared to the previous FPGA implementations, this work achieves 6.51 to 16.7 times lower latency. Moreover, the latency of our FPGA design is sufficiently low to enable deployment of GNNs in a sub-microsecond, real-time collider trigger system, enabling it to benefit from improved accuracy. The proposed LL-GNN design advances the next generation of trigger systems by enabling sophisticated algorithms to process experimental data efficiently. Zhiqiang Que, Hongxiang Fan, Marcus Loo, He Li 0008, Michaela Blott, Maurizio Pierini, Alexander D. Tapper, Wayne Luk |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | When Monte-Carlo Dropout Meets Multi-Exit: Optimizing Bayesian Neural Networks on FPGAabstractBayesian Neural Networks (BayesNNs) have demonstrated their capability of providing calibrated prediction for safety-critical applications such as medical imaging and autonomous driving. However, the high algorithmic complexity and the poor hardware performance of BayesNNs hinder their deployment in real-life applications. To bridge this gap, this paper proposes a novel multi-exit Monte-Carlo Dropout (MCD)-based BayesNN that achieves well-calibrated predictions with low algorithmic complexity. To further reduce the barrier to adopting BayesNNs, we propose a transformation framework that can generate FPGA-based accelerators for multi-exit MCD-based BayesNNs. Several novel optimization techniques are introduced to improve hardware performance. Our experiments demonstrate that our auto-generated accelerator achieves higher energy efficiency than CPU, GPU, and other state-of-the-art hardware implementations. Our code is publicly available at: https://github.com/os-hxfan/MCME_FPGA_Acc.git Hongxiang Fan, Liam Castelli, Zhiqiang Que, He Li 0008, Kenneth Long, Wayne Luk |
DAC | 5 |
| 2023 | MSBF-LSTM: Most-significant Bit-first LSTM Accelerators with Energy Efficiency OptimisationsabstractLong short-term memory (LSTM) recurrent networks are frequently applied to sequence processing problems such as speech recognition and video classification. In this paper, we propose a novel LSTM network implementation based on most-significant bit-first (MSBF) arithmetic on FPGAs, called MSBF-LSTM, to improve the energy efficiency of LSTM inference engines. Furthermore, an LSTM model compression strategy with incremental network quantization is proposed to achieve high accuracy with low-precision weights. Hardware implementations conducted on the Xilinx UltraScale+ Zynq xczu9eg FPGA demonstrate that MSBF-LSTM achieves 1.51× better energy efficiency compared with the state-of-the-art FPGA-based LSTM designs. Sige Bian, He Li 0008, Changjun Song, Yongming Tang |
FCCM | 2 |
| 2023 | MSDF-SGD: Most-Significant Digit-First Stochastic Gradient Descent for Arbitrary-Precision TrainingabstractStochastic gradient descent has been a widely used machine learning algorithm, and interest in low-precision SGD is growing because it improves throughput and keeps efficient convergence. We propose MSDF-SGD, a novel approach allowing SGD to support arbitrary-precision training on FPGAs by employing most-significant digit-first arithmetic. MSDF-SGD is the first architecture that supports arbitrary-precision data, models, and intermediates at the same time. MSDF-SGD is evaluated via training linear classifiers on representative datasets. MSDF-SGD delivers a 1.6× speedup over state-of-the-art low-precision hardware implementations and converges up to 8.6× faster than cutting-edge implementations on CPUs. Finally, we provide a programming interface that permits building a custom arbitrary-precision training accelerator, making MSDF-SGD support more complicated, multi-layered and nonlinear models. Changjun Song, Yongming Tang, Jiyuan Liu 0006, Sige Bian, Danni Deng, He Li 0008 |
FPL | 6 |
| 2023 | Full State Quantum Circuit Simulation Beyond Memory LimitabstractQuantum circuit simulation (QCS) is essential in the noisy intermediate scale quantum (NISQ) era when real quantum computers are scarce. However, fully tracking the states of a quantum system in QCS is highly challenging due to the exponential memory growth that significantly limits the computational reach of classical systems for QCS. Though it is straightforward to leverage secondary storage to extend the scale of QCS, excessive data movement between memory and storage dominates the simulation time, making this solution unrealistic. To tackle this challenge, we identify an intrinsic property of QCS and implement an open-source framework to effectively reduce data movement by >116x. We evaluate the framework on various benchmarks and demonstrate 4x memory reduction with only <20% overhead. On a memory constrained system, we show that it extends the scale of QCS to 32 qubits (64 GB memory requirement) while existing simulators are bounded to 28 qubits (4 GB memory requirement). Our implementation can be accessed via https://github.com/Zhaoyilunnn/qdao. Yilun Zhao 0002, He Li 0008, Ying Wang 0001, Bingmeng Wang, Bing Li 0017, Yinhe Han 0001 |
ICCAD | 3 |
| 2023 | Design Space Exploration for Efficient Quantum Most-Significant Digit-First ArithmeticabstractQuantum computing has been considered as an emerging approach in addressing problems which are not easily solvable using classical computers. In parallel to the physical implementation of quantum processors, quantum algorithms have been actively developed for real-life applications to show quantum advantages, many of which benefit from quantum arithmetic algorithms and their efficient implementations. As one of the most important operations, quantum addition has been adopted in Shor's algorithm and quantum linear algebra algorithms. Although various least-significant digit-first quantum adders have been introduced in previous work, interest in investigating the efficient implementation of most-significant digit-first addition is growing. In this work, we propose a novel design method for most-significant digit-first addition with several quantum circuit optimisations to reduce the number of quantum bits (i.e. qubits), quantum gates, and circuit depth. An open-source library of different arithmetic operators based on our proposed method is presented, where all circuits are implemented on IBM Qiskit SDK. Extensive experiments demonstrate that our proposed design, together with the optimisation techniques, reduces T-depth by up-to 4.0×, T-count by 3.5×, and qubit consumption by 1.2×. He Li 0008, Hongxiang Fan, Yongming Tang |
IEEE Trans. Computers | 1 |
| 2023 | Remarn: A Reconfigurable Multi-threaded Multi-core Accelerator for Recurrent Neural NetworksabstractThis work introduces Remarn, a reconfigurable multi-threaded multi-core accelerator supporting both spatial and temporal co-execution of Recurrent Neural Network (RNN) inferences. It increases processing capabilities and quality of service of cloud-based neural processing units (NPUs) by improving their hardware utilization and by reducing design latency, with two innovations. First, a custom coarse-grained multi-threaded RNN/Long Short-Term Memory (LSTM) hardware architecture, switching tasks among threads when RNN computational engines meet data hazards. Second, the partitioning of this hardware architecture into multiple full-fledged sub-accelerator cores, enabling spatially co-execution of multiple RNN/LSTM inferences. These innovations improve the exploitation of the available parallelism to increase runtime hardware utilization and boost design throughput. Evaluation results show that a dual-threaded quad-core Remarn NPU achieves 2.91 times higher performance while only occupying 5.0% more area than a single-threaded one on a Stratix 10 FPGA. When compared with a Tesla V100 GPU implementation, our design achieves 6.5 times better performance and 15.6 times higher power efficiency, showing that our approach contributes to high performance and energy-efficient FPGA-based multi-RNN inference designs for datacenters. Zhiqiang Que, Hiroki Nakahara, Hongxiang Fan, He Li 0008, Jiuxi Meng, Kuen Hung Tsoi, Xinyu Niu, Eriko Nurvitadhi, Wayne Luk |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2022 | Algorithm and Hardware Co-design for Reconfigurable CNN AcceleratorabstractRecent advances in algorithm-hardware co-design for deep neural networks (DNNs) have demonstrated their potential in automatically designing neural architectures and hardware designs. Nevertheless, it is still a challenging optimization problem due to the expensive training cost and the time-consuming hardware implementation, which makes the exploration on the vast design space of neural architecture and hardware design intractable. In this paper, we demonstrate that our proposed approach is capable of locating designs on the Pareto frontier. This capability is enabled by a novel three-phase co-design framework, with the following new features: (a) decoupling DNN training from the design space exploration of hardware architecture and neural architecture, (b) providing a hardware-friendly neural architecture space by considering hardware characteristics in constructing the search cells, (c) adopting Gaussian process to predict accuracy, latency and power consumption to avoid time-consuming synthesis and place-and-route processes. In comparison with the manually-designed ResNet101, InceptionV2 and MobileNetV2, we can achieve up to 5% higher accuracy with up to$3\times$speed up on the ImageNet dataset. Compared with other state-of-the-art co-design frameworks, our found network and hardware configuration can achieve 2% (~ 6% higher accuracy,$2\times\sim 26\times$smaller latency and$8.5\times$higher energy efficiency. Hongxiang Fan, Martin Ferianc, Zhiqiang Que, He Li 0008, Shuanglong Liu, Xinyu Niu, Wayne Luk |
ASP-DAC | 4 |
| 2022 | A Configurable Floating-Point Multiple-Precision Processing Element for HPC and AI Converged ComputingabstractThere is an emerging need to design configurable accelerators for the high-performance computing (HPC) and artificial intelligence (AI) applications in different precisions. Thus, the floating-point (FP) processing element (PE), which is the key basic unit of the accelerators, is necessary to meet multiple-precision requirements with energy-efficient operations. However, the existing structures by using high-precision-split (HPS) and low-precision-combination (LPC) methods result in low utilization rate of the multiplication array and long multiterm processing period, respectively. In this article, a configurable FP multiple-precision PE design is proposed with the LPC structure. Half precision, single precision, and double precision are supported. The 100% multiplier utilization rate of the multiplication array for all precisions is achieved with improved speed in the comparison and summation process. The proposed design is realized in a 28-nm process with 1.429-GHz clock frequency. Compared with the existing multiple-precision FP methods, the proposed structure achieves 63% and 88% area-saving performance for FP16 and FP32 operations, respectively. The$4\times $and$20\times $maximum throughput rates are obtained when compared with fixed FP32 and FP64 operations. Compared with the previous multiple-precision PEs, the proposed one achieves the best energy-efficiency performance with 975.13 GFLOPS/W. Wei Mao 0002, Kai Li 0024, Liuyao Dai, Xinang Xie, He Li 0008, Longyang Lin, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2021 | Real-Time Super-Resolution System of 4K-Video Based on Deep LearningabstractVideo super-resolution (VSR) technology excels in reconstructing low-quality video, avoiding unpleasant blur effect caused by interpolation-based algorithms. However, vast computation complexity and memory occupation hampers the edge of deplorability and the runtime inference in real-life applications, especially for large-scale VSR task. This paper explores the possibility of real-time VSR system and designs an efficient and generic VSR network, termed EGVSR. The proposed EGVSR is based on spatio-temporal adversarial learning for temporal coherence. In order to pursue faster VSR processing ability up to 4K resolution, this paper tries to choose lightweight network structure and efficient upsampling method to reduce the computation required by EGVSR network under the guarantee of high visual quality. Besides, we implement the batch normalization computation fusion, convolutional acceleration algorithm and other neural network acceleration techniques on the actual hardware platform to optimize the inference process of EGVSR network. Finally, our EGVSR achieves the real-time processing capacity of [email protected]. Compared with TecoGAN, the most advanced VSR network at present, we achieve 85.04% reduction of computation density and 7.92× performance speedups. In terms of visual quality, the proposed EGVSR tops the list of most metrics (such as LPIPS, tOF, tLP, etc.) on the public test dataset Vid4 and surpasses other state-of-the-art methods in overall performance score. Yanpeng Cao, Changjun Song, Yongming Tang, He Li 0008 |
ASAP | 5 |
| 2021 | Joint Sparsity with Mixed Granularity for Efficient GPU Implementation
Chuliang Guo, Xingang Yan, Yufei Chen 0007, He Li 0008, Xunzhao Yin, Cheng Zhuo |
DATE | 4 |
| 2021 | A Reconfigurable Multiple-Precision Floating-Point Dot Product Unit for High-Performance ComputingabstractThere is an emerging need to optimize floating-point (FP) dot product units (DPU) for high-performance scientific computing as well as training deep learning models. Due to different precision requirements of applications, a reconfigurable multiple-precision DPU operation can largely reduce the cost of area and power. However, the existing methods could result in redundant bits for unit multipliers, but also leave idle hardware resources for the operations in different precisions. In this paper, a reconfigurable multiple-precision FP DPU design is proposed for high-performance computing (HPC) applications. The FP DPU can be reconfigured as follows. A bit-partitioning method is provided to minimize the redundant bits with a configurable mixed-precision multiplier for three-mode operations: 20 half-precision Dot Product (DP), 5 single-precision DP, and 1 double-precision DP operations. Any of the modes can be executed in two successive clock cycles without idle hardware resources. The proposed design is realized by using the UMC 55-nm process with simulation results. Compared with the existing multiple-precision FP methods, the proposed DPU achieves 88.9% and 35.8% area-saving performance for FP16 and FP32 operations, respectively. Moreover, when using benchmarked HPC applications where multiple precisions can be used, the proposed reconfigurable DPU can accelerate up to 4× and 20× maximum throughput rates when compared with fixed FP32 and FP64 operations, respectively. Wei Mao 0002, Kai Li 0024, Xinang Xie, Shirui Zhao, He Li 0008, Hao Yu 0001 |
DATE | 5 |
| 2021 | Reconfigurable Synthesizable Synchronization FIFOsabstractWe present a reconfigurable high-throughput synthesizable synchronization FIFO for crossing between asynchronous and synchronous timing domains. This FIFO is composed of reconfigurable mix-and-match components. The FIFO is fully synthesizable using standard-cell libraries and a standard design flow. A post-layout design throughput exceeds 1.2 G-transfer/sec when implemented in a 65nm CMOS process. Ameer Abdelhadi, He Li 0008 |
FCCM | 2 |
| 2021 | A General Video Processing Framework on Edge Computing FPGAsabstractDigital video processing needs high bandwidth transmission from source to host, which poses a massive challenge to existing technologies. As an appropriate solution, edge computing can provide immediate process with low latency, high transmission bandwidth and memory usage. In this paper, we propose a general video processing framework on edge computing FPGAs (GVPF-E), consisting of in-out buffers and an update-feedback mechanism. For general video processing algorithms, GVPF-E extracts multiple video frames' inter-correlation and updates calculation results into output feature buffers, so as to we can optimize the video processing quality via update-feedback mechanism. Our illustrative hardware implementations on a CNN-based video filtering algorithm achieve an up-to 3.87TMACS performance under 8-bit quantization. We also obtain 2.45GB/s bandwidth and 94.4% peak throughput utilization under Xilinx XCZU15EG embedded computing FPGAs. Feng Yu 0006, He Li 0008, Rongshi Dai, Yongming Tang |
FCCM | 2 |
| 2021 | Enabling Mixed-Timing NoCs for FPGAs: Reconfigurable Synthesizable Synchronization FIFOsabstractWe present an architecture of a reconfigurable high-throughput synthesizable synchronization FIFO for crossing between asynchronous and synchronous timing domains. This FIFO is composed of reconfigurable mix-and-match components. The input and output interfaces are interchangeable for edge-triggered synchronous communication and for the asP* asynchronous pulse-based handshake protocol. The FIFO capacity, data width, synchronizer latency, and interface protocols are independent design parameters. The FIFO is fully synthesizable using widely available standard-cell libraries and a standard ASIC design flow. Our post-layout design can operate at speeds greater than 1.2 giga-transfer per second under worst-case conditions when implemented in a 65nm CMOS process. Ameer Abdelhadi, He Li 0008 |
FPL | 2 |
| 2021 | Security Enhancements for Approximate Machine LearningabstractApproximate computing techniques for error-tolerant machine learning applications are gaining interest, promising a energy-accuracy balance for modern digital computing systems. As an ubiquitous step in machine learning, iterative solvers have been widely used for training neural networks and accelerating feedforward computations. In this paper, we provide three novel information hiding techniques that use properties of redundant number systems, most-significant digit-first arithmetic and the forward error analysis of stationary iterative methods, to secure approximate computing systems. We demonstrate that three different security signatures are encoded through redundant representation, function-equivalence arithmetic replacement and algorithmic optimisation. Our illustrative security enhancement countermeasures can be used to prevent potential attacks, such as privacy leakage, out-of-control systematic error and error injection in approximate computing. He Li 0008, Yaru Pang, Jiliang Zhang 0002 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | Digit Stability Inference for Iterative Methods Using Redundant Number RepresentationabstractIn our recent work on iterative computation in hardware, we showed that arbitrary-precision solvers can perform more favorably than their traditional arithmetic equivalents when the latter's precisions are either under- or over-budgeted for the solution of the problem at hand. Significant proportions of these performance improvements stem from the ability to infer the existence of identical most-significant digits between iterations. This technique uses properties of algorithms operating on redundantly represented numbers to allow the generation of those digits to be skipped, increasing efficiency. It is unable, however, to guarantee that digits will stabilize, i.e., never change in any future iteration. In this article, we address this shortcoming, using interval and forward error analyses to prove that digits of high significance will become stable when computing the approximants of systems of linear equations using stationary iterative methods. We formalize the relationship between matrix conditioning and the rate of growth in most-significant digit stability, using this information to converge to our desired results more quickly. Versus our previous work, an exemplary hardware realization of this new technique achieves an up-to 2.2× speedup in the solution of a set of variously conditioned systems using the Jacobi method. He Li 0008, Ian McInerney, James J. Davis 0001, George A. Constantinides |
IEEE Trans. Computers | 1 |
| 2021 | On-Chip Trust Evaluation Utilizing TDC-Based Parameter-Adjustable Security PrimitiveabstractField-programmable gate arrays (FPGAs) are integrated circuits (ICs) that can be reconfigured to the desired functionalities, without manufacturing dedicated chips. Due to their programmable nature, FPGAs have been prevalent in the large majority of modern systems. This raises high demands for verifying the security of circuit implementations on FPGAs, since they are vulnerable to hardware trojans (HTs) that can be inserted through modified configuration files. In this article, we propose an on-chip security framework to ensure the trustworthiness of circuit implementations on FPGAs at runtime. The core of the framework is a time-to-digital converter (TDC)-based hardware security primitive that can be predeployed on FPGAs to verify whether the FPGA-based designs are tampered with or corrupted by HTs. The parameter-adjustable TDC sensor, which is the primary component of the primitive, is carefully designed, adjusted, and implemented, thus the TDC sensor can monitor the transient voltage fluctuations within FPGAs with a high resolution. Versus statistical data analysis, tiny abnormal variations introduced by the Trojan insertion and activation are distinguished. Experimental results on Xilinx Spartan-6 FPGAs demonstrate the effectiveness of the proposed TDC-based on-chip trust evaluation framework and HT detection method. Haocheng Ma, Jiaji He 0001, Yanjiang Liu, Jun Kuai, He Li 0008, Leibo Liu, Yiqiang Zhao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | FTDL: A Tailored FPGA-Overlay for Deep Learning with High ScalabilityabstractFast inference is of paramount value to a wide range of deep learning applications. This work presents FTDL, a highly-scalable FPGA overlay framework for deep learning applications, to address the architecture and hardware mismatch faced by traditional efforts. The FTDL overlay is specifically optimized for the tiled structure of FPGAs, thereby achieving post-place-and-route operating frequencies exceeding 88 % of the theoretical maximum across different devices and design scales. A flexible compilation framework efficiently schedules matrix multiply and convolution operations of large neural network inference on the overlay and achieved over 80 % hardware efficiency on average. Taking advantage of both high operating frequency and hardware efficiency, FTDL achieves 402.6 and 151.2 FPS with GoogLeNet and ResNet50 on ImageNet, respectively, while operating at a power efficiency of 27.6 GOPS/W, making it up to 7.7× higher performance and 1.9× more power-efficient than the state-of-the-art. Runbin Shi, Yuhao Ding, Xuechao Wei, He Li 0008, Hang Liu 0001, Hayden Kwok-Hay So, Caiwen Ding |
DAC | 4 |
| 2020 | A novel method for malware detection on ML-based visualization technique
Xinbo Liu, Yaping Lin, He Li 0008, Jiliang Zhang 0002 |
Comput. Secur. | 3 |
| 2020 | architect: Arbitrary-Precision Hardware With Digit Elision for Efficient Iterative ComputeabstractMany algorithms feature an iterative loop that converges to the result of interest. The numerical operations in such algorithms are generally implemented using finite-precision arithmetic, either fixed- or floating-point, most of which operate least-significant digit first. This results in a fundamental problem: if, after some time, the result has not converged, is this because we have not run the algorithm for enough iterations or because the arithmetic in some iterations was insufficiently precise? There is no easy way to answer this question, so users will often over-budget precision in the hope that the answer will always be to run for a few more iterations. We propose a fundamentally new approach: with the appropriate arithmetic able to generate results from most-significant digit first, we show that fixed compute-area hardware can be used to calculate an arbitrary number of algorithmic iterations to arbitrary precision, with both precision and approximant index increasing in lockstep. Consequently, datapaths constructed following our principles demonstrate efficiency over their traditional arithmetic equivalents where the latter's precisions are either under- or over-budgeted for the computation of a result to a particular accuracy. Use of most-significant digit-first arithmetic additionally allows us to declare certain digits to be stable at runtime, avoiding their recalculation in subsequent iterations and thereby increasing performance and decreasing memory footprints. Versus arbitrary-precision iterative solvers without the optimizations we detail herein, we achieve up-to 16× performance speedups and 1.9× memory savings for the evaluated benchmarks. He Li 0008, James J. Davis 0001, John Wickerson, George A. Constantinides |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | ATMPA: attacking machine learning-based malware visualization detection methods via adversarial examplesabstractSince the threat of malicious software (malware) has become increasingly serious, automatic malware detection techniques have received increasing attention, where machine learning (ML)-based visualization detection methods become more and more popular. In this paper, we demonstrate that the state-of-the-art ML-based visualization detection methods are vulnerable to Adversarial Example (AE) attacks. We develop a novel Adversarial Texture Malware Perturbation Attack (ATMPA) method based on the gradient descent and L-norm optimization method, where attackers can introduce some tiny perturbations on the transformed dataset such that ML-based malware detection methods will completely fail. The experimental results on the MS BIG malware dataset show that a small interference can reduce the accuracy rate down to 0% for several ML-based detection methods, and the rate of transferability is 74.1% on average. Xinbo Liu, Jiliang Zhang 0002, Yaping Lin, He Li 0008 |
IWQoS | 4 |
| 2018 | Digit Elision for Arbitrary-accuracy Iterative ComputationabstractWe recently proposed the first hardware architecture enabling the iterative solution of systems of linear equations to accuracies limited only by the amount of available memory. This technique, named ARCHITECT, achieves exact numeric computation by using online arithmetic to allow the refinement of results from earlier iterations over time, eschewing rounding error. ARCHITECT has a key drawback, however: often, many more digits than strictly necessary are generated, with this problem exacerbating the more accurate a solution is sought. In this paper, we infer the locations of these superfluous digits within stationary iterative calculations by exploiting online arithmetic's digit dependencies and using forward error analysis. We demonstrate that their lack of computation is guaranteed not to affect the ability to reach a solution of any accuracy. Versus ARCHITECT, our illustrative hardware implementation achieves a geometric mean 20.1× speedup in the solution of a set of representative linear systems through the avoidance of redundant digit calculation. For the computation of high-precision results, we also obtain an up-to 22.4 × memory requirement reduction over the same baseline. Finally, we demonstrate that solvers implemented following our proposals can show superiority over conventional arithmetic implementations by virtue of their runtime-tunable precisions. He Li 0008, James J. Davis 0001, John Wickerson, George A. Constantinides |
ARITH | 1 |
| 2017 | architect: Arbitrary-precision constant-hardware iterative computeabstractMany algorithms feature an iterative loop that converges to the result of interest. The numerical operations in such algorithms are generally implemented using finite-precision arithmetic, either fixed or floating point, most of which operate least-significant digit first. This results in a fundamental problem: if, after some time, the result has not converged, is this because we have not run the algorithm for enough iterations or because the arithmetic in some iterations was insufficiently precise? There is no easy way to answer this question, so users will often over-budget precision in the hope that the answer will always be to run for a few more iterations. We propose a fundamentally new approach: armed with the appropriate arithmetic able to generate results from most-significant digit first, we show that fixed compute-area hardware can be used to calculate an arbitrary number of algorithmic iterations to arbitrary precision, with both precision and iteration index increasing in lockstep. Thus, datapaths constructed following our principles demonstrate efficiency over their traditional arithmetic equivalents where the latter's precisions are either under- or over-budgeted for the computation of a result to a particular accuracy. For the execution of 100 iterations of the Jacobi method, we obtain a 1.60× increase in frequency and 15.7× LUT and 50.2× flip-flop reductions over a 2048-bit parallel-in, serial-out traditional arithmetic equivalent, along with 46.2× LUT and 83.3× flip-flop decreases versus the state-of-the-art online arithmetic implementation. He Li 0008, James J. Davis 0001, John Wickerson, George A. Constantinides |
FPT | 1 |
| 2016 | A survey of hardware Trojan threat and defense
He Li 0008, Qiang Liu 0011, Jiliang Zhang 0002 |
Integr. | 1 |
| 2015 | A Survey of Hardware Trojan Detection, Diagnosis and PreventionabstractHardware Trojans (HTs) can be implanted in security-weak parts of a chip with various means to steal the internal sensitive data or modify original functionality, which may lead to huge economic losses and great harm to society. Therefore, it is very important to perform hardware Trojan detection and diagnosis, find potential safety hazards and apply protection techniques in the whole IC design cycle, in order to enhance the security of chips. In this paper, we elaborate an IC market model, and describe the potential HT threats faced by the parties involved in the model. Then we survey the recent research advances in the countermeasures against HT attacks, which are classified into HT detection, diagnosis and prevention. Finally, the challenges and prospects for HT defense are illuminated. He Li 0008, Qiang Liu 0011, Jiliang Zhang 0002, Yongqiang Lyu 0001 |
CAD/Graphics | 1 |
| 2014 | Hardware Trojan detection acceleration based on word-level statistical properties managementabstractHardware Trojan insertion has raised serious concerns to semiconductor industry and government agencies. Hardware Trojan is usually activated under rare conditions associated with low transition bits in a circuit. The damage includes circuit functional failure or important information leakage. Previous research on hardware Trojan detection is mainly based on side-channel analysis and Trojan activation. Long activation time is a major concern during the detection process. In this paper, we propose a novel approach for efficiently accelerating Trojan activation by increasing the transition activity of rare bits. In particular, the proposed approach increases the bit-level transition activity by controlling signal word-level statistical properties, such as changing the variance and autocorrelation of the signal. In addition, by analyzing the signal propagation statistical properties through various digital signal processing (DSP) operators such as adders and multipliers, the proposed approach can control the statistical properties of internal signals and then enhance the internal bit transition activity from the primary input of the circuit. The proposed approach is evaluated on several circuits. The results show that the transition activity of rare bits can be dramatically increased by up to 166.7 times and Trojan activation time can be reduced by up to 121 times. He Li 0008, Qiang Liu 0011 |
FPT | 1 |