EDBT 2026 Demo / reviewers in the wild / expert
Jie Gu 0001
dblp:20/6440-1
· DBLP profile ↗
25ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-2912-7294ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLA: Enhancing Security and Privacy for Generative Models with Logic-Locked AcceleratorsabstractWe introduce LLA, an effective intellectual property (IP) protection scheme for generative AI models. LLA leverages the synergy between hardware and software to defend against various supply chain threats, including model theft, model corruption, and information leakage. On the software side, it embeds key bits into neurons that can trigger outliers to degrade performance and applies invariance transformations to obscure the key values. On the hardware side, it integrates a lightweight locking module into the AI accelerator while maintaining compatibility with various dataflow patterns and toolchains. An accelerator with a pre-stored secret key acts as a license to access the model services provided by the IP owner. The evaluation results show that LLA can withstand a broad range of oracle-guided key optimization attacks, while incurring a minimal computational overhead of less than 0.1% for 7,168 key bits. You Li 0008, Guannan Zhao, Yuhao Ju, Yunqi He, Jie Gu 0001, Hai Zhou 0001 |
AAAI | 5 |
| 2026 | A Differentiable Simulator for Optimizing Time-Domain Analog CNN Accelerators
Mark Horton, Changwoo Park, Tergel Molom-Ochir, William McGarry, Jie Gu 0001, Yiran Chen 0001 |
ISLPED | 5 |
| 2026 | A Physics-Informed Neural Network Surrogate for Runtime PDN and Dynamic Droop Prediction in 2.5-D Chiplet Integration
Xi Chen 0099, Yuhao Ju, Seda Ogrenci Memik, Jie Gu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Humanoid Robot Control: A Mixed-Signal Footstep Planning SoC with ZMP Gait Scheduler and Neural Inverse KinematicsabstractWith the rapid expansion of autonomous robotic systems in recent years, humanoid robots are also gaining considerable attention. However, their motion control presents more complex challenges than wheeled mobile robots. For the first time, this work presents a complete footstep planning SoC chip for humanoid robots. It includes a novel time-domain graph search engine for 3D footstep planning, along with a mixed-signal zero-moment point (ZMP) gait scheduler enhanced by neural inverse kinematics for efficient motion control. This work is demonstrated in-situ on a fully assembled robot using the 65nm system-on-chip (SoC), achieving the best-in-class performance and power including an 18.4x enhancement on energy efficiency for motion control and a 2.7x energy saving in graph search compared to previous works. The demo video is provided in https://youtu.be/kBe-aRnzmG4. Qiankai Cao, Juin Chuen Oh, Jie Gu 0001 |
ASP-DAC | 4 |
| 2025 | Headset-Integrated Brain-Machine Interface for Mind Imagery and Control in VR/MR ApplicationsabstractVirtual Reality and Mixed Reality systems have revolutionized consumer electronics, driving innovations in the metaverse. However, traditional VR headsets lack brain activity integration for feedback and control. This work introduces the first 65nm SoC for in situ mind imagery-based brain machine interface, seamlessly integrated into VR/MR headsets, with reduced energy consumption and AI support. The digital core of the SoC achieves state-of-the-art energy consumption <1μJ/class for computation-intensive CNN operations thanks to a novel teacher-student low-power scheme, general instruction set architecture for general programming and system-level optimizations of the design. The demonstration video for the whole system is provided at https://youtu.be/WuGlcMSSQzY. Yijie Wei, Lance Christopher Go, Jie Gu 0001 |
ASP-DAC | 5 |
| 2025 | Modeling, Design and In-situ Demonstration of Bio-inspired Central Pattern Generator and Neuromorphic Computing Circuits for Complex Kinematic Control of Quadruped Robots
Qiankai Cao, Yuhao Ju, Zhengyu Chen 0002, Jie Gu 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Development of a Physics-Informed Neural Network Model for Rapid Power Integrity Analysis in Die-Level and Die-Package Co-Design for 2.5-D Chiplet SolutionsabstractThis work presents a novel power distribution network (PDN) analysis using the emerging physics-informed neural network (PINNs) model. Different from conventional solver-based analysis, PINN allows rapid analysis and prediction while maintaining the physics compliance for high-fidelity analysis. An adaptive multi-objective training strategy is introduced, incorporating an interconnection matrix and layer labeling to accelerate convergence across complex PDN structures. An embedded workload vector and a transient modulator with linear superposition method extend the model’s applicability to a wide range of power scenarios. The developed method is applied to both chip-level PDN analysis and Chiplet 2.5-D chip-package co-analysis, showing high accuracy and fast runtime compared with conventional methods. The approach captures both steady-state IR droop and dynamic transient supply droop, including IR and L•di/dt noise from package and on-die PDN. Experiments on 2.5-D Chiplets with RISC-V processors and CNN accelerators show that the proposed PINN-based method achieves a 299x and 7x reduction in runtime compared to conventional EDA tools or prior work and saves up to 80% of training data than traditional neural networks models. Xi Chen 0099, Yuhao Ju, Jie Gu 0001 |
ISLPED | 3 |
| 2024 | LLM-MARK: A Computing Framework on Efficient Watermarking of Large Language Models for Authentic Use of Generative AI at Local DevicesabstractAs generative AI such as ChatGPT rapidly evolves, the increasing incidence of data misconduct such as the proliferation of counterfeit news or unauthorized use of Large Language Models (LLMs) presents a significant challenge for consumers to obtain authentic information. While new watermarking schemes are recently being proposed to protect the intellectual property (IP) of LLM, the computation cost is unfortunately too high for the targeted real-time execution on local devices. In this work, a specialized hardware-efficient watermarking computing framework is proposed enabling model authentication at local devices. By employing the proposed hardware hashing for fast lookup and pruned bitonic sorting network acceleration, the developed architecture framework enables fast and efficient watermarking of LLM on the small local devices. The proposed architecture is evaluated on Xilinx XCZU15EG FPGA, demonstrating 30x computing speed-up, making this architecture highly suitable for integration into local mobile devices. The proposed algorithm to architecture codesign framework offers a practical solution to the immediate challenges posed by LLM misuse, providing a feasible hardware solution for Intellectual Property protection in the era of generative AI. Yuhao Ju, Xi Chen 0099, Jie Gu 0001 |
DAC | 4 |
| 2023 | Development of Tropical Algebraic Accelerator with Energy Efficient Time-Domain Computing for Combinatorial Optimization and Machine LearningabstractTropical algebra solves complex problems with only sum and min/max operations replacing expensive multiplication and addition in linear algebra. Due to the low computing cost, tropical algebra has recently gained significant attention in a broad range of areas such as combinatorial optimization, scheduling, machine learning, etc. In this paper, we propose a generic hardware accelerator architecture for tropical algebra supporting a wide range of applications. Novel time-domain (TD) computing accelerators with special mapping, precision expansion and, unrolling techniques are proposed to further improve hardware efficiency. Test results on various tropical calculations including linear regression, dynamic programming, and neural network are shown to demonstrate an energy saving from 1.5X to 2.1X, latency saving from 2.6X to 5.2X, or an overall energy-delay-product (EDP) improvement from 3.9X-10.5X compared with conventional digital implementation manifesting the promise of the algebraic solution on low power edge devices. Qiankai Cao, Xi Chen 0099, Jie Gu 0001 |
ISLPED | 3 |
| 2022 | Human emotion based real-time memory and computation management on resource-limited edge devicesabstractEmotional AI or Affective Computing has been projected to grow rapidly in the upcoming years. Despite many existing developments in the application space, there has been a lack of hardware-level exploitation of the user's emotions. In this paper, we propose a deep collaboration between user's affects and the hardware system management on resource-limited edge devices. Based on classification results from efficient affect classifiers on smartphone devices, novel real-time management schemes for memory, and video processing are proposed to improve the energy efficiency of mobile devices. Case studies on H.264 / AVC video playback and Android smartphone usages are provided showing significant power saving of up to 23% and reduction of memory loading of up to 17% using the proposed affect adaptive architecture and system management schemes. Yijie Wei, Jie Gu 0001 |
DAC | 3 |
| 2020 | Exploration of Design Space and Runtime Optimization for Affective Computing in Machine Learning Empowered Ultra-Low Power SoCabstractThe incorporation of artificial intelligence into the rapidly growing IoT devices demands a high level of built-in intelligence, e.g. machine learning capability at the device level. Affective computing offers a new degree of cognitive intelligence into edge processing IoT devices by inferring human emotion, stress levels for intelligent human assistance. This work explores the design space and runtime optimization opportunity for affective computing at the system-on-chip (SoC) level. A design optimization methodology for the neural network classifier and runtime power management schemes are proposed to achieve high energy efficiency on embedded low power devices. A test chip based on a 65nm CMOS process was used to demonstrate the proposed methodology on emotion and stress classification for affective computing. An average power saving of 45% is achieved with a peak power savings of 60% from the proposed emotion-driven adaptive power management scheme. Yijie Wei, Kofi Otseidu, Jie Gu 0001 |
DAC | 3 |
| 2020 | NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceabstractMachine learning inference has become an essential task for embedded edge devices requiring the deployment of costly deep neural network accelerators onto extremely resource-constrained hardware. Although many optimization strategies have been proposed to improve the efficiency of standalone accelerators, the optimization for end-to-end performance of a computing device with heterogeneous cores is still challenging and often overlooked, especially for low power devices. In this paper, we propose a unified reconfigurable architecture, referred as Neural CPU (NCPU), for low-cost embedded systems. The proposed architecture is built on a binary neural network accelerator with the capability to emulate an in-order RISC-V CPU pipeline. The NCPU supports flexible programmability of RISC-V and maintains data locally to avoid costly core-to-core data transfer. A two-core NCPU SoC is designed and fabricated in a 65nm CMOS process. Compared with the conventional heterogeneous architecture, a single NCPU achieves 35% area reduction and 12% energy saving at 0.4V, which is suitable for low power and low-cost embedded edge devices. The NCPU design also features the capability of smooth switching between general-purpose CPU operation and a binary neural network inference to realize full utilization of the cores. The implemented two-core NCPU SoC achieves an end-to-end performance speed-up of 43% or an equivalent 74% energy saving based on use cases of real-time image classification and motion detection. Yuhao Ju, Russ Joseph, Jie Gu 0001 |
MICRO | 4 |
| 2019 | Digital Compatible Synthesis, Placement and Implementation of Mixed-Signal Time-Domain ComputingabstractMixed-signal time-domain computing (TC) has recently drawn significant attention due to its high efficiency in applications such as machine learning accelerators. However, due to the nature of analog and mixed-signal design, there is a lack of a systematic flow of synthesis and place & route for time-domain circuits. This paper proposed a comprehensive design flow for TC. In the front-end, a variation-aware digital compatible synthesis flow is proposed. In the back-end, a placement technique using graph-based optimization engine is proposed to deal with the especially stringent matching requirement in TC. Simulation results show significant improvement over the prior analog placement methods. A 55nm test chip is used to demonstrate that the proposed design flow can meet the stringent timing matching target for TC with significant performance boost over conventional digital design. Zhengyu Chen 0002, Hai Zhou 0001, Jie Gu 0001 |
DAC | 3 |
| 2019 | R-Accelerator: An RRAM-Based CGRA Accelerator With Logic ContractionabstractIn this paper, a novel RRAM-based reconfigurable accelerator (R-accelerator) design is proposed, which makes special use of existing RRAM device for high-efficient reconfigurable application-specific computing. The proposed R-accelerator design consists of RRAM-based arithmetic unit (AU) array, fully integrated into commercial EDA design tools. A significant area saving is achieved compared with conventional digital counterpart due to the proposed logic contraction technique, as well as saving of storage space and routing congestions. For enabling the optimization of the AU array, this paper also proposes a systematical method on the synthesis of the AU array under routing channel constraint for application-specific designs. Two automatic mapping algorithms, including simultaneous mapping and incremental mapping algorithms, are proposed and compared. The experiments using 45-nm CMOS technology on a case study of dynamic time warping example and general benchmark programs show up to 49% area reduction and 28% performance enhancement using the proposed R-accelerator technique compared with conventional application-specified integrated circuit (ASIC) design. Zhengyu Chen 0002, Hai Zhou 0001, Jie Gu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Compiler-guided instruction-level clock scheduling for timing speculative processorsabstractDespite the significant promise that circuit-level timing speculation has for enabling operation in marginal conditions, overheads associated with recovery prove to be a serious drawback. We show that fine-grained clock adjustment guided by the compiler can be used to stretch and shrink the clock to maximize benefits of timing speculation and reduce the overheads associated with recovery. We present a formulation for compiler-driven clock scheduling and explore the benefits in several scenarios. Our results show that there are significant opportunities to exploit timing slack when there are appropriate channels for the compiler to select clock period at cycle-level. Yuanbo Fan, Jie Gu 0001, Simone Campanoni, Russ Joseph |
DAC | 3 |
| 2018 | Cyclic locking and memristor-based obfuscation against CycSAT and inside foundry attacksabstractThe high cost of IC design has made chip protection one of the first priorities of the semiconductor industry. Although there is a common impression that combinational circuits must be designed without any cycles, circuits with cycles can be combinational as well. Such cyclic circuits can be used to reliably lock ICs. Moreover, since memristor is compatible with CMOS structure, it is possible to efficiently obfuscate cyclic circuits using polymorphic memristor-CMOS gates. In this case, the layouts of the circuits with different functionalities look exactly identical, making it impossible even for an inside foundry attacker to distinguish the defined functionality of an IC by looking at its layout. In this paper, we propose a comprehensive chip protection method based on cyclic locking and polymorphic memristor-CMOS obfuscation. The robustness against state-of-the-art key-pruning attacks is demonstrated and the overhead of the polymorphic gates is investigated. Amin Rezaei 0001, Yuanqi Shen, Shuyu Kong, Jie Gu 0001, Hai Zhou 0001 |
DATE | 4 |
| 2018 | Design and optimization of edge computing distributed neural processor for biomedical rehabilitation with sensor fusionabstractModern biomedical devices use sensor fusion techniques to improve the classification accuracy of motion intent of users for rehabilitation application. The design of motion classifier observes significant challenges due to the large number of channels and stringent communication latency requirement. This paper proposes an edge-computing distributed neural processor to effectively reduce the data traffic and physical wiring congestion. A special local and global networking architecture is introduced to significantly reduce traffic among multi-chips in edge computing. To optimize the design space of the features selected, a systematic design methodology is proposed. A novel mixed-signal feature extraction approach with assistance of neural network distortion recovery is also provided to significantly reduce the silicon area. A 12-channel 55nm CMOS test chip was implemented to demonstrate the proposed systematic design methodology. The measurement shows the test chip consumes only 20uW power, more than 10,000X less power than the current clinically used microprocessor and can perform edge-computing networking operation within 5ms time. Kofi Otseidu, Joshua Bryne, Levi J. Hargrove, Jie Gu 0001 |
ICCAD | 5 |
| 2018 | R-Accelerator: A Reconfigurable Accelerator with RRAM Based Logic Contraction and Resource Optimization for Application Specific ComputingabstractIn this paper, we introduce a novel reconfigurable accelerator (R-accelerator) design which embeds RRAM device into traditional logic circuits for high-performance application specific computing. To facilitate the synthesis of the proposed RRAM based logic cell, a special logic contraction technique is developed to maximize the area saving. In order to optimize the arithmetic unit array for instruction set mapping and interconnect routing, a new resource allocation algorithm is also proposed to achieve further saving in area and power. Using a fully integrated design flow with commercial design tools, our experimental results show that the proposed RRAM based R-accelerator architecture offers 45% area improvement, 33% power reduction and 32% performance enhancement in a 45nm CMOS process compared with conventional CMOS design. Zhengyu Chen 0002, Hai Zhou 0001, Jie Gu 0001 |
ICCD | 3 |
| 2018 | A Comprehensive Stochastic Design Methodology for Hold-Timing Resiliency in Voltage-Scalable Design
Zhengyu Chen 0002, Geng Xie, Jie Gu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Design and Synthesis of Self-Healing Memristive Circuits for Timing Resilient Processor DesignabstractModern microprocessors suffer from significant on-chip variation at the advanced technology nodes. The development of CMOS-compatible memristive devices has brought nonvolatile capability into silicon technology. This paper explores new applications for memristive devices to resolve performance degradations that result from process variation. Novel self-healing flip-flop and clock buffers are developed to automatically detect timing violation and to perform timing recovery by tuning the resistance values of memristor devices. To incorporate the circuit techniques into VLSI circuits design, novel device placement and tuning algorithms have been developed. The proposed design methodology is demonstrated in a 45-nm fast Fourier transform processor design. Our test results show that performance gains of up to 20% can be achieved using the proposed self-healing circuits, with only 1% area Shuyu Kong, Hai Zhou 0001, Jie Gu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Greybox Design Methodology: A Program Driven Hardware Co-optimization with Ultra-Dynamic Clock ManagementabstractIn this paper, a novel Greybox design methodology is proposed to establish a design and co-optimization flow across the boundary of conventional software and hardware design. The dynamic timing of each software instruction is simulated and associated with processor hardware design, which provides the basis of ultra-dynamic clock management. The proposed scheme effectively implements the instruction-based clock management and achieves 21.71% frequency speedup. Besides, a novel program-driven hardware optimization flow is proposed, in which software operations are mapped with hardware gate netlist and sorted by the usage frequency. The experiments on an ARM based pipeline design in commercial 65nm CMOS process show an extra 10% frequency speedup is obtained with high optimization efficiency. Overall, the proposed Greybox design method achieves frequency speedup by 31.56%, comparing with conventional design method. Russ Joseph, Jie Gu 0001 |
DAC | 3 |
| 2017 | Cell-to-array thermal-aware analysis of stacked RRAMabstractThe crossbar resistive random access memory (RRAM) has been studied extensively due to its low-power, low-cost, high density and nonvolatile characteristics. However, the dependence of RRAM performance parameters on temperature, from cell to array level, is less explored. Particularly, thin-film based RRAM that is integrated into 3D ICs is subject to severe thermal conditions. Hence, temperature dependence of RRAM behavior needs to be well understood in order to construct the most effective thermal management strategy for these systems. In this paper, a detailed RRAM device thermal model is proposed. Our experiments show that the temperature of the surrounding environment critically impacts the RRAM readout margin, which in turn, jeopardizes its promised advantages for high-density stacked integration. Our thermal model can be utilized at the architectural level to predict the latency variation in the RRAM memory as a function of the thermal environment and drive thermal management policies for the 3D IC. Yingyi Luo, Seda Ogrenci Memik, Jie Gu 0001 |
ISCAS | 3 |
| 2016 | Exploration of associative power management with instruction governed operation for ultra-low power designabstractThis paper explores a novel associative low power operation where instructions govern the operation of on-chip regulators in real time. Based on explicit association between long delay instruction patterns and hardware performance, an instruction based power management scheme is developed with energy models formulated for deriving the energy efficiency of the associative operation. The proposed system scheme is demonstrated using a low power microprocessor design with an integrated switched capacitor regulator in 45nm CMOS technology. Simulations on benchmark programs show a power saving of around 14% from the proposed scheme. A novel compiler optimization strategy is also proposed to further improve the energy efficiency. Yuanbo Fan, Russ Joseph, Jie Gu 0001 |
DAC | 4 |
| 2016 | Analysis and Design of Energy Efficient Time Domain Signal ProcessingabstractTime domain signal processing (TDSP) encodes information into time rather than voltage with higher efficiency than conventional digital design. This paper performs systematical analysis on the design principle and energy efficiency of TDSP. Variation impact, which poses significant challenges to TDSP, is evaluated and a variation driven design methodology is proposed to achieve an optimum tradeoff between energy efficiency and design robustness. Several novel circuit level design techniques such as dual encoding strategy and bit-scalable design are also proposed in this work to significantly improve the energy efficiency of TDSP. Design example on a critical building block of facial recognition application was used to demonstrate the potential of the technique. The result in a 45nm technology shows 3.3X energy-delay product reduction and 34% area saving can be achieved using TDSP compared with conventional digital design technique. Zhengyu Chen 0002, Jie Gu 0001 |
ISLPED | 2 |
| 2016 | Comprehensive Analysis, Modeling and Design for Hold-Timing Resiliency in Voltage Scalable DesignabstractResiliency to timing violation is a crucial requirement for low power electronics operating across a wide range of supply voltages. Although many existing solutions enhance setup timing tolerance for the higher performance, an accurate modeling and design strategy for hold resiliency dealing with conflicting requirement from both high voltages and low voltages has not been established. This paper proposes a novel voltage-scalable modeling technique that leverages conventional static timing analysis and efficient statistical analysis to achieve accurate stochastic hold timing analysis. Several highly nonlinear behaviors of circuit operation are also incorporated into the proposed model to achieve a model accuracy of within 10% of spice Monte-Carlos simulation. Leveraging the proposed modeling technique, a novel hold resilience design technique is proposed to eliminate the excessive hold fixing operation for low voltage operation and its associated performance degradation at high voltage while still being compatible with conventional design closure flow. The proposed design methodology is demonstrated in a 45nm DSP processor design enabling a voltage-scalable operation from 0.35V to 0.9V eliminating more than 20,000 hold buffers as well as 23% performance degradation at high voltages due to hold fixing. Geng Xie, Jie Gu 0001 |
ISLPED | 3 |