Yuhao Ju

dblp:263/0469 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-2509-400XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLA: Enhancing Security and Privacy for Generative Models with Logic-Locked Accelerators
abstract
We introduce LLA, an effective intellectual property (IP) protection scheme for generative AI models. LLA leverages the synergy between hardware and software to defend against various supply chain threats, including model theft, model corruption, and information leakage. On the software side, it embeds key bits into neurons that can trigger outliers to degrade performance and applies invariance transformations to obscure the key values. On the hardware side, it integrates a lightweight locking module into the AI accelerator while maintaining compatibility with various dataflow patterns and toolchains. An accelerator with a pre-stored secret key acts as a license to access the model services provided by the IP owner. The evaluation results show that LLA can withstand a broad range of oracle-guided key optimization attacks, while incurring a minimal computational overhead of less than 0.1% for 7,168 key bits.
You Li 0008, Guannan Zhao, Yuhao Ju, Yunqi He, Jie Gu 0001, Hai Zhou 0001
AAAI3
2026 A Physics-Informed Neural Network Surrogate for Runtime PDN and Dynamic Droop Prediction in 2.5-D Chiplet Integration
Xi Chen 0099, Yuhao Ju, Seda Ogrenci Memik, Jie Gu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 Modeling, Design and In-situ Demonstration of Bio-inspired Central Pattern Generator and Neuromorphic Computing Circuits for Complex Kinematic Control of Quadruped Robots
Qiankai Cao, Yuhao Ju, Zhengyu Chen 0002, Jie Gu 0001
ACM Great Lakes Symposium on VLSI2
2025 Development of a Physics-Informed Neural Network Model for Rapid Power Integrity Analysis in Die-Level and Die-Package Co-Design for 2.5-D Chiplet Solutions
abstract
This work presents a novel power distribution network (PDN) analysis using the emerging physics-informed neural network (PINNs) model. Different from conventional solver-based analysis, PINN allows rapid analysis and prediction while maintaining the physics compliance for high-fidelity analysis. An adaptive multi-objective training strategy is introduced, incorporating an interconnection matrix and layer labeling to accelerate convergence across complex PDN structures. An embedded workload vector and a transient modulator with linear superposition method extend the model’s applicability to a wide range of power scenarios. The developed method is applied to both chip-level PDN analysis and Chiplet 2.5-D chip-package co-analysis, showing high accuracy and fast runtime compared with conventional methods. The approach captures both steady-state IR droop and dynamic transient supply droop, including IR and L•di/dt noise from package and on-die PDN. Experiments on 2.5-D Chiplets with RISC-V processors and CNN accelerators show that the proposed PINN-based method achieves a 299x and 7x reduction in runtime compared to conventional EDA tools or prior work and saves up to 80% of training data than traditional neural networks models.
Xi Chen 0099, Yuhao Ju, Jie Gu 0001
ISLPED2
2024 LLM-MARK: A Computing Framework on Efficient Watermarking of Large Language Models for Authentic Use of Generative AI at Local Devices
abstract
As generative AI such as ChatGPT rapidly evolves, the increasing incidence of data misconduct such as the proliferation of counterfeit news or unauthorized use of Large Language Models (LLMs) presents a significant challenge for consumers to obtain authentic information. While new watermarking schemes are recently being proposed to protect the intellectual property (IP) of LLM, the computation cost is unfortunately too high for the targeted real-time execution on local devices. In this work, a specialized hardware-efficient watermarking computing framework is proposed enabling model authentication at local devices. By employing the proposed hardware hashing for fast lookup and pruned bitonic sorting network acceleration, the developed architecture framework enables fast and efficient watermarking of LLM on the small local devices. The proposed architecture is evaluated on Xilinx XCZU15EG FPGA, demonstrating 30x computing speed-up, making this architecture highly suitable for integration into local mobile devices. The proposed algorithm to architecture codesign framework offers a practical solution to the immediate challenges posed by LLM misuse, providing a feasible hardware solution for Intellectual Property protection in the era of generative AI.
Yuhao Ju, Xi Chen 0099, Jie Gu 0001
DAC2
2020 NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End Performance
abstract
Machine learning inference has become an essential task for embedded edge devices requiring the deployment of costly deep neural network accelerators onto extremely resource-constrained hardware. Although many optimization strategies have been proposed to improve the efficiency of standalone accelerators, the optimization for end-to-end performance of a computing device with heterogeneous cores is still challenging and often overlooked, especially for low power devices. In this paper, we propose a unified reconfigurable architecture, referred as Neural CPU (NCPU), for low-cost embedded systems. The proposed architecture is built on a binary neural network accelerator with the capability to emulate an in-order RISC-V CPU pipeline. The NCPU supports flexible programmability of RISC-V and maintains data locally to avoid costly core-to-core data transfer. A two-core NCPU SoC is designed and fabricated in a 65nm CMOS process. Compared with the conventional heterogeneous architecture, a single NCPU achieves 35% area reduction and 12% energy saving at 0.4V, which is suitable for low power and low-cost embedded edge devices. The NCPU design also features the capability of smooth switching between general-purpose CPU operation and a binary neural network inference to realize full utilization of the cores. The implemented two-core NCPU SoC achieves an end-to-end performance speed-up of 43% or an equivalent 74% energy saving based on use cases of real-time image classification and motion detection.
Yuhao Ju, Russ Joseph, Jie Gu 0001
MICRO2