Zunkai Huang

dblp:174/6588 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-7501-8959ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 A 50 μW/Gbps/Lane Power-Efficient MIPI D-PHY Receiver With Architecture-Level Adaptive and Structural Optimizations for Micro-Displays
abstract
Achieving high power efficiency in Mobile Industry Processor Interface (MIPI) D-PHY receivers is crucial for micro-display chips in AR/VR systems, where stringent power constraints exist. However, existing designs often sacrifice power efficiency for higher data rates due to architectural limitations, neglecting optimization for low-power applications. To address this issue, we propose a receiver architecture that substantially enhances power efficiency through three key techniques. First, we improve the gain-bandwidth product (GBW) by employing an autonomous gain scheduling analog front-end (AFE) that dynamically tunes the gain while reducing drive current. Second, we reduce clocking overhead by introducing a self-monitoring interferometric deserializer that enables clock-free pre-scaling and halves the DDR sampling frequency. Third, we increase transition speed and minimize short-circuit power by utilizing a chaotic topological flow actuator (CTFA) with multi-path current feedthrough. Compared to prior state-of-the-art designs, the proposed receiver achieves a power efficiency of$50~\mu $W/Gbps/lane ($42~\mu $A/Gbps/lane), reducing power and current consumption by 46% and 45%, respectively, using a standard 180-nm process.
Haoran Zeng, Yingqi Feng, Tianai Li, Hang Ye 0007, Zunkai Huang, Hui Wang 0036, Yongxin Zhu 0001, Qiliang Li, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 A Robust and Efficient Multiscenario Object Detection Network for Edge Devices
abstract
Deep-learning-based object detection has been increasingly attractive in various intelligent edge applications, including remote sensing and autonomous driving. However, achieving an optimal tradeoff between computing efficiency and detection accuracy is challenging. YOLO-series networks provide fast and lightweight detection but often sacrifice accuracy. In this work, we propose the multibranch cascading aggregation YOLO (MCA-YOLO) model, designed for multiscenario object detection tasks on edge devices. MCA-YOLO enhances detection accuracy by learning comprehensive feature information while maintaining computational efficiency through three key components: multibranch spatial pyramid pooling (MSPP), Ghost-convolution (GC)-based efficient layer aggregation network (G-ELAN), and hierarchical aggregation neck (HAN) that integrates MSPP and G-ELAN with GC. We train and validate MCA-YOLO using four benchmark datasets: VOC, COCO, SIMD, and VisDrone. The experimental results demonstrate significant improvements in detection accuracy and inference speed. Besides, we deploy the MCA-YOLO on a Jetson Xavier NX edge device embedded in an unmanned aerial vehicle (UAV), creating a remote sensing system capable of real-time, high-accuracy object detection. Our code is publicly available athttps://github.com/lawlawCodes/MCA-YOLO.
Zhihuan Chen, Aiwen Luo, Jialu Zheng, Zunkai Huang
IEEE Geosci. Remote. Sens. Lett.5
2025 Bit-Sparsity Aware Acceleration With Compact CSD Code on Generic Matrix Multiplication
abstract
The ever-increasing demand for matrix multiplication in artificial intelligence (AI) and generic computing emphasizes the necessity of efficient computing power accommodating both floating-point (FP) and quantized integer (QINT). While state-of-the-art bit-sparsity-aware acceleration techniques have demonstrated impressive performance and efficiency in neural networks through software-driven methods such as pruning and quantization, these approaches are not always feasible in typical generic computing scenarios. In this paper, we propose Bit-Cigma, a hardware-centric architecture that leverages bit-sparsity to accelerate generic matrix multiplication. Bit-Cigma features (1) CCSD encoding, an optimized on-chip sparsification technique based on canonical signed digit (CSD) representation; (2) segmented dot product, a multi-stage exponent matching technique for long FP vectors; and (3) the versatility to efficiently process both FP and QINT data types. CCSD encoding halves the cost of CSD encoding while achieving optimal bit-sparsity, and segmented dot product improves both accuracy and throughput. Bit-Cigma cores are implemented using 65 nm technology at 1 GHz, demonstrating substantial gains in performance and efficiency for both FP and QINT configurations. Compared to state-of-the-art Bitlet, Bit-Cigma achieves 3.2$\boldsymbol{\times}$performance, 6.1$\boldsymbol{\times}$area efficiency, and 15.3$\boldsymbol{\times}$energy efficiency when processing FP32 data while ensuring zero computing error.
Zixuan Zhu 0001, Chundong Wang 0001, Zunkai Huang, Yongxin Zhu 0001
IEEE Trans. Computers5
2023 ELANet: Effective Lightweight Attention-Guided Network for Real-Time Semantic Segmentation
Qingming Yi, Guoshuai Dai, Zunkai Huang, Aiwen Luo
Neural Process. Lett.4
2023 LMFFNet: A Well-Balanced Lightweight Network for Fast and Accurate Semantic Segmentation
abstract
Real-time semantic segmentation is widely used in autonomous driving and robotics. Most previous networks achieved great accuracy based on a complicated model involving mass computing. The existing lightweight networks generally reduce the parameter sizes by sacrificing the segmentation accuracy. It is critical to balance the parameters and accuracy for real-time semantic segmentation. In this article, we propose a lightweight multiscale-feature-fusion network (LMFFNet) mainly composed of three types of components: split-extract-merge bottleneck (SEM-B) block, feature fusion module (FFM), and multiscale attention decoder (MAD), where the SEM-B block extracts sufficient features with fewer parameters. FFMs fuse multiscale semantic features to effectively improve the segmentation accuracy and the MAD well recovers the details of the input images through the attention mechanism. Without pretraining, LMFFNet-3-8 achieves 75.1% mean intersection over union (mIoU) with 1.4 M parameters at 118.9 frames/s using RTX 3090 GPU. More experiments are investigated extensively on various resolutions on other three datasets of CamVid, KITTI, and WildDash2. The experiments verify that the proposed LMFFNet model makes a decent tradeoff between segmentation accuracy and inference speed for real-time tasks. The source code is publicly available at https://github.com/Greak-1124/LMFFNet.
Jialin Shen, Qingming Yi, Jian Weng 0001, Zunkai Huang, Aiwen Luo, Yicong Zhou
IEEE Trans. Neural Networks Learn. Syst.5
2021 SparkNoC: An energy-efficiency FPGA-based accelerator using optimized lightweight CNN for edge computing
Zunkai Huang, Hui Wang 0036, Victor Chang 0001, Yongxin Zhu 0001, Songlin Feng
J. Syst. Archit.2
2020 Anomaly Detection Based on RBM-LSTM Neural Network for CPS in Advanced Driver Assistance System
abstract
Advanced Driver Assistance System (ADAS) is a typical Cyber Physical System (CPS) application for human–computer interaction. In the process of vehicle driving, we use the information from CPS on ADAS to not only help us understand the driving condition of the car but also help us change the driving strategies to drive in a better and safer way. After getting the information, the driver can evaluate the feedback information of the vehicle, so as to enhance the ability to assist in driving of the ADAS system. This completes a complete human–computer interaction process. However, the data obtained during the interaction usually form a large dimension, and irrelevant features sometimes hide the occurrence of anomalies, which poses a significant challenge to us to better understand the driving states of the car. To solve this problem, we propose an anomaly detection framework based on RBM-LSTM. In this hybrid framework, RBM is trained to extract general underlying features from data collected by CPS, and LSTM is trained from the features learned by RBM. This framework can effectively improve the prediction speed and present a good prediction accuracy to show vehicle driving condition. Besides, drivers are allowed to evaluate the prediction results, so as to improve the accuracy of prediction. Through the experimental results, we can find that the proposed framework not only simplifies the training of the entire neural network and increases the training speed but also greatly improves the accuracy of the interaction-driven data analysis. It is a valid method to analyze the data generated during the human interaction.
Hanlin Zhu, Yongxin Zhu 0001, Victor Chang 0001, Cong He, Ching-Hsien Hsu, Hui Wang 0036, Songlin Feng, Zunkai Huang
ACM Trans. Cyber Phys. Syst.10
2019 Designing efficient accelerator of depthwise separable convolutional neural network on FPGA
Zunkai Huang, Hui Wang 0036, Songlin Feng
J. Syst. Archit.3