Hongyang Jia

dblp:02/11188 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
abstract
Neural network (NN) accelerators with multi-chipmodule (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resource underutilization and off-chip communication overheads. Traditional parallelization schemes for NN inference on MCM architectures, such as intra-layer parallelism and inter-layer pipelining, show incompetency in breaking through both challenges, limiting the scalability of MCM architectures. We observed that existing works typically deploy layers separately rather than considering them jointly. This underexploited dimension leads to compromises between system computation and communication, thus hindering optimal utilization, especially as hardware/software scale. To address this limitation, we propose Scope, a merged pipeline framework incorporating this overlooked multi-layer dimension, thereby achieving improved throughput and scalability by relaxing tradeoffs between computation, communication and memory costs. This new dimension, however, adds to the complexity of design space exploration (DSE). To tackle this, we develop a series of search algorithms that achieves exponential-to-linear complexity reduction, while identifying solutions that rank in the top 0.05% of performance. Experiments show that Scope achieves up to $1.73 \times$ throughput improvement while maintaining similar energy consumption for ResNet-152 inference compared to state-of-the-art approaches.
Zongle Huang, Hongyang Jia, Kaiwei Zou, Yongpan Liu
ASP-DAC2
2026 SHyLA: 3D-Stacked NVM-DRAM Hybrid LLM-Inference Architecture Exploiting Data and Memory Heterogeneity
Fuyao Zhou, Shunan Dong, Huazhong Yang, Yongpan Liu, Hongyang Jia
ISCA9
2025 Presto: A Unified RISC-V-Compatible SoC for Multi-Scheme FHE Acceleration over Module Lattice
Luchang Lei, Gangfeng Du, Zhenyu Guan 0002, Huazhong Yang, Yongpan Liu, Song Bian 0001, Hongyang Jia
HCS11
2025 MCHEAS: Optimizing Large-Parameter NTT Over Multicluster In-Situ FHE Accelerating System
abstract
Fully Homomorphic encryption (FHE) enables high-level security but with a heavy computation workload, necessitating software-hardware co-design for aggressive acceleration. Recent works on specialized accelerators for HE evaluation have made significant progress in supporting lightweight RNS-CKKS applications, especially those with high-density in-memory computing techniques. To fulfill higher computational demands for more general applications, this article proposes multicluster HE accelerating system (MCHEAS), an accelerating system comprising multiple in-situ HE processing accelerators, each functioning as a cluster to perform large-parameter RNS-CKKS evaluation collaboratively. MCHEAS features optimization strategies including the synchronous, preemptive swap, square-diagonal, and odd-even index separation. Using these strategies to compile the computation and transmission of number theoretic transform (NTT) coefficients, the method optimizes the intercluster data swaps, a major bottleneck in NTT computations. Evaluations show that under 1 GHz, with different intercluster data transfer bandwidths, our approach accelerates NTT computations by 26.40% to 51.75%. MCHEAS also improves computing unit utilization by 10.30% to 33.97%, with a maximum peak utilization rate of up to 99.62%. MCHEAS achieves 17.63% to 34.67% speedups for HE operations involving NTT, and 15.12% to 30.62% speedups for demonstrated applications, while enhancing the computing units’ utilization by 5.18% to 21.87% during application execution. Furthermore, we compare MCHEAS with SOTA designs under a specific intercluster data transfer bandwidth, achieving up to$81.45\times $their area efficiencies in applications.
Zhenyu Guan 0002, Luchang Lei, Hongyang Jia, Yi Chen 0012, Bo Zhang 0142, Changrui Ren, Jin Dong 0004, Song Bian 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 SASDenSebLE: A Compact Vision Transformer Inference Architecture With Saturation-Approximate Softmax Dataflow Enabling Sequence-Parallelism Boosted Layer-Fusion Execution
Zongle Huang, Shupei Fan, Luchang Lei, Huazhong Yang, Yongpan Liu, Hongyang Jia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2025 Guest Editorial TCAS-I Special Issue on the ESSERC 2024 Conference
abstract
status: Published
Georges Gielen, Jan Craninckx, Hongyang Jia
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Invited: Automatic Hardware/Software Design for High-Speed Autonomous Unmanned Aerial Vehicles Guided by a Flight Model
abstract
Autonomous Unmanned Aerial Vehicles (UAVs) are on the rise in the industrial and academic communities. Since most UAVs are severely size, weight, and power (SWaP) constrained, building computing system for high-speed UAVs is challenging. Current domain-specific hardware-software (HW-SW) designs for UAVs are mainly bottom-up, focusing on optimizing a single module in the whole system, such as visual-inertial odometry (VIO), depth estimation, or planning. But this leads to underdesign for flight speed as agile navigation depends on a tight combination of multiple modules. To find the optimal HW-SW design for systematic flight performance, we propose a top-down automatic design framework. A flight model is introduced to guide the inter-module and the intra-module HW-SW optimization towards the system-level goal. For the perception algorithms, we define the representative design space. And some critical non-AI operators are accelerated and profiled on embedded GPU to achieve better hardware performance. The design framework is evaluated on a micro UAV equipped with a Nvidia Jetson Orin NX. In a specific navigation scenario, the design found by our framework achieve 40% and 65% increase on flight speed than two manual design methods respectively.
Yuanfan Xu, Suquan Zhang, Yunfei Xiang, Hongyang Jia, Yu Wang 0002
DAC5
2024 ESC-NTT: An Elastic, Seamless and Compact Architecture for Multi-Parameter NTT Acceleration
abstract
Fully homomorphic encryption (FHE) and post-quantum cryptography (PQC) heavily rely on number theoretic transform (NTT) to accelerate polynomial multiplication, However, most existing NTT accelerators lack flexibility when the underlying modulus and polynomial lengths change. Current designs often store twiddle factors in on-chip storage, facing a noticeable drawback when frequent parameter changes occur, leading to a potential 50% decrease in computation speed due to the input bandwidth limitations. To address this challenge, we propose ESC-NTT, a fully-pipelined and flexible architecture for handling NTTs with varying parameters. ESC-NTT, a complete custom architecture, continuously performs$N$-point (inverse) NTT, negacyclic NTT (NCN), and inverse NCN (INCN) without introducing bubbles during modulus and NTT length switches. Additionally, we introduce a twiddle factor generator (TFG) module to replace on-chip factor storage and save 68.7% twiddle factors' bandwidth compared to inputting every factor. In the experiment, ESC-NTT is implemented on a Xilinx Alveo U280 FPGA and synthesized in a 28 nm CMOS technology. In the case of frequent modulus switching and same on-chip storage, the calculation speed of ESC-NTT is 1.05× to 241.39× that of existing FHE accelerators when performing 4096-point NTT.
Zhenyu Guan 0002, Luchang Lei, Hongyang Jia, Yi Chen 0012, Bo Zhang 0142, Jin Dong 0004, Song Bian 0001
DATE6
2023 Data-Driven Security and Stability Rule in High Renewable Penetrated Power System Operation
abstract
Power systems around the world are experiencing an energy revolution that substitutes fossil fuels with renewable energy. Such a transition poses two significant challenges: highly variable generators that add short-term and long-term difficulties for supply–demand balance, and a high proportion of convertor-based devices that may jeopardize power system security and stability. At the same time, machine learning techniques provide more opportunities to study the complex power system security and stability problems. This article summarizes the machine learning framework to embed security rules into power system operation optimization under high renewable energy penetration. First, we explore how high penetration renewable energy impacts power system security and stability. Then, we review how the complex security and stability boundary of power systems is modeled using various machine learning techniques. Finally, we show how the machine learning model is transformed into optimization constraints that can be embedded into the power system operation model. The framework is substantiated through case studies of practical power systems.
Ning Zhang 0008, Hongyang Jia, Qingchun Hou
Proc. IEEE2
2019 A Programmable Embedded Microprocessor for Bit-scalable In-memory Computing
abstract
This article consists of a collection of slides from the author's conference presentation.
Hongyang Jia, Hossein Valavi, Yinqi Tang, Naveen Verma
Hot Chips Symposium1
2018 Genetic Programming for Energy-Efficient and Energy-Scalable Approximate Feature Computation in Embedded Inference Systems
abstract
With the increasing interest in deploying embedded sensors in a range of applications, there is also interest in deploying embedded inference capabilities. Doing so under the strict and often variable energy constraints of the embedded platforms requires algorithmic, in addition to circuit and architectural, approaches to reducing energy. A broad approach that has recently received considerable attention in the context of inference systems is approximate computing. This stems from the observation that many inference systems exhibit various forms of tolerance to data noise. While some systems have demonstrated significant approximation-versus-energy knobs to exploit this, they have been applicable to specific kernels and architectures; the more generally available knobs have been relatively weak, resulting in large data noise for relatively modest energy savings (e.g., voltage overscaling, bit-precision scaling). In this work, we explore the use of genetic programming (GP) to compute approximate features. Further, we leverage a method that enhances tolerance to feature-data noise through directed retraining of the inference stage. Previous work in GP has shown that it generalizes well to enable approximation of a broad range of computations, raising the potential for broad applicability of the proposed approach. The focus on feature extraction is deliberate because they involve diverse, often highly nonlinear, operations, challenging general applicability of energy-reducing approaches. We evaluate the proposed methodologies through two case studies, based on energy modeling of a custom low-power microprocessor with a classification accelerator. The first case study is on electroencephalogram-based seizure detection. We find that the choice of two primitive functions (square root, subtraction) out of seven possible primitive functions (addition, subtraction, multiplication, logarithm, exponential, square root, and square) enables us to approximate features in 0.41$mJ$per feature vector (FV), as compared to 4.79$mJ$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 11.68$\times$. The important system-level performance metrics for seizure detection are sensitivity, latency, and number of false alarms per hour. Our set of GP models achieves 100 percent sensitivity, 4.37 second latency, and 0.15 false alarms per hour. The baseline performance is 100 percent sensitivity, 3.84 second latency, and 0.06 false alarms per hour. The second case study is on electrocardiogram-based arrhythmia detection. In this case, just one primitive function (multiplication) suffices to approximate features in 1.13$\mu J$per FV, as compared to 11.69$\mu J$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 10.35$\times$. The important system-level metrics in this case are sensitivity, specificity, and accuracy. Our set of GP models achieves 81.17 percent sensitivity, 80.63 percent specificity, and 81.86 percent accuracy, whereas the baseline achieves 82.05 percent sensitivity, 88.12 percent specificity, and 87.92 percent accuracy. These case studies demonstrate the possibility of a significant reduction in feature extraction energy at the expense of a slight degradation in system performance.
Hongyang Jia, Naveen Verma, Niraj K. Jha
IEEE Trans. Computers2
2014 Personalizing Statistical Models for Asthma Prognosis and Therapeutics
Hongyang Jia, Patricia Flatley Brennan, Junbo Son, Yu-Ting Hung
AMIA1
2014 Register allocation for hybrid register architecture in nonvolatile processors
abstract
Nonvolatile processors (NVP) have been an emerging topic in recent years due to its zero standby power, data retention and instant-on features. The conventional full replacement architecture in NVP has drawbacks of large area overhead and high backup energy. This paper provides a partial replacement based hybrid register architecture to significantly abate above problems. However, the hybrid register architecture can induce potential critical data loss and backup errors. In this paper, we propose a critical-data overflow aware register allocation (CORA). Different from other register allocation methods, CORA efficiently reduces the possibility of critical data spilling and backup errors. The experiment results show that CORA reduces the critical data overflow rate by up to 52%. The hybrid register architecture reduces the chip area by 45.1% and backup energy by 82.8% when using CORA.
Hongyang Jia, Yongpan Liu, Qing'an Li, Chun Jason Xue, Huazhong Yang
ISCAS2
2012 An energy harvesting nonvolatile sensor node and its application to distributed moving object detection
abstract
Energy harvesting sensor nodes based on real nonvolatile processors are demonstrated to show the desirable characteristics of those systems, such as no battery, zero stand-by power, microsecond-scale sleep and wake-up time, high resilience to random power failures and fine-grained power management. Furthermore, we show its applications to a distributed moving object detection system, one of novel nonvolatile computing systems.
Yongpan Liu, Hongyang Jia, Shan Su, Jinghuan Wen, Wenzhu Zhang, Lin Zhang 0001, Huazhong Yang
IPSN3