Yucong Huang

dblp:167/9059 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Model-Based Imaginative Planning for Embodied Agents
abstract
Junru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang, Weien Zhou, Feifei Wang, Yang Yang, Ying Wen, Wen Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang, Weien Zhou
ACL (1)3
2026 TranCAD: Transforming tabular data into color images for deep semi-supervised anomaly detection
Yucong Huang, Feng Xu 0008, Xin Lyu 0001, Zhennan Xu
Expert Syst. Appl.1
2026 RV-WINO: A RISC-V Neural Network Accelerator Based on Winograd Algorithm Fabricated in 55-nm CMOS Process
abstract
The rapid evolution of artificial intelligence (AI) in IoT applications necessitates the execution of inference tasks on edge devices. However, the deployment of computation-intensive neural networks on resource-constrained edge systems presents a significant challenge. This brief presents the RV-WINO processor, the first silicon implementation of a RISC-V processor based on the Winograd algorithm for convolution and general matrix multiplication (GEMM) acceleration. The processor incorporates a Winograd module, which significantly reduces multiplication operations during convolutions, leading to a substantial decrease in energy consumption. In addition, the processor includes a matrix multiplication module that reuses the multipliers of the Winograd module, accelerating fully connected and dot product operations in neural networks. The RV-WINO processor fabricated in a 55-nm CMOS process achieves the peak computational performance of 0.95 and 2.39 GOPS in INT32 and INT8 modes, with its peak energy efficiency reaching 112 and 237 GOPS/W. In convolutional neural network (CNN) inference tasks, the execution time is reduced by over 80% compared with the baseline processor.
Yucong Huang, Qu Lu, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye
IEEE Trans. Very Large Scale Integr. Syst.2
2025 Logic Gate Network Inference Acceleration with RISC-V Custom Instruction Set
abstract
Logic Gate Networks (LGNs) exploit the similarity between neural networks and logic circuit networks and replace the neurons with logic gates.Consequently, the computation inside the neurons can be replaced by Boolean operations (16 operations for two-input logic).LGNs can be implemented by logic-based instructions in processors and significantly reduce the computation overhead during inference.However, the encoding and decoding processes at the input and output stages of LGNs face efficiency challenges when using traditional RISC-V instruction sets.This limitation arises because these processes rely on one-bit operations, which cannot fully utilize the 32-bit bandwidth of standard instructions.In this work, we proposed four custom RISC-V-based instructions to accelerate the encoding and decoding processes of LGNs.An applicationspecific RISC-V processor, called RV-LGN, has been implemented on FPGA and synthesized using Synopsys® Design Compiler with the CMOS 55nm process.The custom instructions can be called via in-line assembly in C code, making RV-LGN highly promising for implementation in edge devices.Benchmark tests on MIT-BIH, MNIST, and CIFAR-10 classification tasks demonstrate that RV-LGN achieves a runtime reduction of over 87% compared to a generic RISC-V RV32IM ISA processor.Additionally, power consumption during LGN inference is significantly reduced.For the MIT-BIH dataset, the energy consumption is 0.098 µJ/Beat, while MNIST and CIFAR-10 tasks require 0.18 µJ/Image and 0.51 µJ/Image, respectively.These results highlight the superior efficiency of RV-LGN compared to other processors.
Chenxi Feng, Xinyu Kang, Yuru Li, Yucong Huang, Terry Tao Ye
CF5
2025 Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
abstract
For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy for Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.
Jiesong Lian, Yucong Huang, Chengdong Ma, Ying Wen 0001, Long Hu, Yixue Hao
ECAI2
2025 A Low-Noise, High-Input-Impedance Pre-Amplifier for Piezoelectric MEMS Microphone
abstract
This paper proposes a low-noise, high-input-impedance pre-amplifier for piezoelectric MEMS microphones. It connects directly to the piezoelectric audio sensor without the need for auxiliary biases and operates under a low supply voltage, with a low power consumption, and a small silicon footprint. To recognize a wide range of sound pressure levels with low THD, the pre-amplifier can switch between two modes, i.e., a 20dB high-gain mode and a 10dB low-gain mode. It also incorporates an impedance-enhanced ESD design for the IO pads and a source follower for impedance conversion and input noise suppression. The design is implemented in 0.18-µm CMOS process, and the performance estimation is based on post-layout simulation. The input referred noise of the proposed pre-amplifier is 8.77µVrms in low-gain mode and 3.65µVrms in high-gain mode. it also features a DC input impedance of 190 GΩ, 74.56 dB PSRR, 10.4/20.9 dB system gain, 0.136%/0.044% THD with 94 dBSPL input. It consumes 86.5 µA of current under a 1.6-3.6 V power supply and occupies an active silicon area of 0.23 mm2.
Weiye Song, Yucong Huang, Terry Tao Ye
ISCAS2
2025 NNia-8: An 8-Core RISC-V Neural Network Inference Accelerator with Efficient Processing Elements and Memory Utilization
Yucong Huang, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye
NPC (2)2
2025 MSADNet: Multi-Scale Adaptive Dual Attention Network for Multivariate Time Series Anomaly Detection
abstract
Multivariate time series anomaly detection requires consideration of the features across different variables. It involves extracting and aggregating the temporal features and dependencies at multiple scales. In order to capture local temporal patterns and long-term dependencies in multivariate time series, most existing methods segment time series into patches with different sizes for multi-scale modeling. However, these methods still face the following challenges: 1) They typically use the fixed-size patches instead of resizing the patches according to different time series patterns; 2) After segmenting time series into multiple patches, they usually struggle to model the local and global dependencies. To address these issues, we propose a multi-scale adaptive dual attention network (MSADNet) for multivariate time series anomaly detection. Specifically, an adaptive weight allocation mechanism is proposed to dynamically adjust the patch size of each multivariate time series and further enable adaptive multi-scale feature extraction and aggregation. Meanwhile, a dual patchwise attention module is proposed to efficiently model the global and local dependencies of time series from two views: interpatch and intra-patch. Extensive experiments on five real-world datasets show that MSADNet achieves an average F1 score of 0.9622 and outperforms existing state-of-the-art methods.
Yuyan Yang, Yucong Huang
SMC4
2025 RV-SCNN: A RISC-V Processor With Customized Instruction Set for SNN and CNN Inference Acceleration on Edge Platforms
abstract
The rapid advancement of artificial intelligence (AI) applications has driven an increasing demand for conducting inference tasks on edge devices. However, implementing computation-intensive neural networks on resource-constrained edge systems remains a significant challenge. In this article, we propose a novel processor architecture called RV-SCNN to address this challenge. The architecture is based on the RISC-V generic instruction set and incorporates various single instruction multiple data (SIMD) custom instruction extensions to accelerate the computation of spike neural networks (SNNs) and convolutional neural networks (CNNs), enabling efficient execution of complex neural network models. The core operators of the processor are shared by both SNN and CNN operations, thus supporting both computation modes. Other acceleration implementations include an internal hardware loop control unit that reduces the instruction overhead, an address calculation unit and an interlayer fusion unit that minimize the memory access overhead, as well as an image to column (IM2COL) unit that improves the computational efficiency of the$3 \times 3$convolutions in SNNs and CNNs. The custom instructions are called through inline assembly in the C program, providing higher flexibility compared to traditional ASICs and supporting custom complex SNN/CNN network structures. Compared to traditional instruction sets, the RV-SCNN processor reduces the execution time of CNNs and SNNs by over 90%. We validate the processor on FPGA platform and evaluate its performance under CMOS 55-nm process. The processor achieves an operational efficiency of 9.88 pJ/SOP in SNN network inference tasks, while the peak energy efficiency reaches 679 GOPS/W in CNN network inference.
Chenxi Feng, Xinyu Kang, Qi Wang 0051, Yucong Huang, Terry Tao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 IGFNet: An Interactive-Guided Fusion Network for Hyperspectral Pansharpening
abstract
Hyperspectral pansharpening is an efficient approach to obtaining high-resolution hyperspectral images (HR-HSIs) by fusing low-resolution hyperspectral images (LR-HSIs) with high-resolution panchromatic images (HR-PANs). However, the spatial and spectral distortions in reconstructed HR-HSIs are almost inevitable due to the modal gap between LR-HSIs and HR-PANs. Therefore, the performance of multi-source features fusion largely hinges on the ability to extract and align heterogeneous features across modalities. Most of the existing methods focus on integrating decoupled spatial and spectral information from different sources directly, which poses a dual challenge in aligning both spatial and spectral features effectively. To address the issues above-mentioned, a novel method named Interactive-Guided Fusion Network (IGFNet) is proposed, which is built upon a multi-stage progressive fusion framework. A high-resolution branch (HR) is introduced to interactively guide the alignment between cross-modal features, by fusing up-sampled HSI and PAN as a joint spatial-spectral guidance signal. Furthermore, the alignment is progressively conducted across stages, narrowing the modality gap and enhancing the representation of HR spatial-spectral feature. Additionally, we designed parameter-free spatial, spectral, and spatial-spectral attention mechanisms to extract global and local features effectively. Extensive experiments on reduced-resolution and full-resolution datasets demonstrate that IGFNet outperforms state-of-the-art across various metrics. Specifically, with a scaling factor of 4 on the Pavia University dataset, our method achieves a 2.09% relative improvement in PSNR, a 0.99% relative increase in SSIM, a 2.6% relative reduction in SAM, while reducing the parameter count by compared to the baseline MDA-Net.
Zhennan Xu, Xin Lyu 0001, Feng Xu 0008, Xin Li 0090, Yucong Huang, Caifeng Wu, Yiwei Fang
IEEE Trans. Geosci. Remote. Sens.5
2024 RV-GEMM: Neural Network Inference Acceleration with Near-Memory GEMM Instructions on RISC-V
abstract
General Matrix Multiply (GEMM), as a fundamental operation in neural network, plays an important role in artificial intelligence and signal processing applications. In this paper, we proposed three SMID RISC-V custom instructions to accelerate GEMM computations, supporting multiple precisions including 32-bit, 16-bit and 8-bit fixed. Furthermore, we implemented address calculation and loop control units along with the GEMM acceleration module to reduce the memory access overhead. These three GEMM custom instructions, along with the near-memory optimization units, were incorporated in the RV-GEMM processor and implemented on the FPGA platform for speedup evaluation. It was also compiled in Synopsys Design Compiler with CMOS 55nm process for hardware overhead estimation. Compared to the baseline RISC-V processor, for GEMM computations under precisions of 32-bit, 16-bit and 8-bit fixed, the RV-GEMM processor achieved speedup ratios of 15.8×, 28.7× and 42.5×. The peak energy efficiency also reached 260 GOPS/W, 420 GOPS/W and 609 GOPS/W, respectively.
Chenxi Feng, Bingzhen Chen, Qi Wang 0051, Yucong Huang, Terry Tao Ye
CF5
2024 RWriC: A Dynamic Writing Scheme for Variation Compensation for RRAM-based In-Memory Computing
abstract
RRAM-based compute-in-memory (CIM) suffers from programming variation issues, specifically device-to-device variation (DDV) and cycle-to-cycle variation (CCV), which can have a detrimental impact on inference accuracy. To address these variation issues, we propose RWriC, a dynamic Writing scheme for variation Compensation for RRAM-based CIM. RWriC sequentially programs the weights, implemented by multiple RRAM cells, starting from the high significance cell (HSC) and moving towards the low significance cell (LSC). This approach leverages the knowledge of current cumulative errors and the programming targets (PTs) of other RRAM cells to dynamically adjust the PT of the RRAM currently under programming. By shifting the PT of HSC, RWriC enables the LSC to compensate for the programming errors of the HSC. Moreover, when the variation is substantial, RWriC allows the magnitude of LSC to be scaled up, providing an even wider compensation range. Through the combined application of the shifting and scaling techniques, experimental results show that the inference accuracy for ResNet50 on the CIFAR-10 dataset only drops by 0.9% under 18% device variation. In comparison to the conventional writing scheme, our RWriC approach achieves a 5-11x improvement in variation robustness for ResNet50 and Yolov8 across different tasks.
Yucong Huang, Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui, Terry Tao Ye
DAC1
2023 RVComp: Analog Variation Compensation for RRAM-Based in-Memory Computing
abstract
Resistive Random Access Memory (RRAM) has shown great potential in accelerating memory-intensive computation in neural network applications. However, RRAM-based computing suffers from significant accuracy degradation due to the inevitable device variations. In this paper, we propose RVComp, a fine-grained analog Compensation approach to mitigate the accuracy loss of in-memory computing incurred by the Variations of the RRAM devices. Specifically, weights in the RRAM crossbar are accompanied by dedicated compensation RRAM cells to offset their programming errors with a scaling factor. A programming target shifting mechanism is further designed with the objectives of reducing the hardware overhead and minimizing the compensation errors under large device variations. Based on these two key concepts, we propose double and dynamic compensation schemes and the corresponding support architecture. Since the RRAM cells only account for a small fraction of the overall area of the computing macro due to the dominance of the peripheral circuitry, the overall area overhead of RVComp is low and manageable. Simulation results show RVComp achieves a negligible 1.80% inference accuracy drop for ResNet18 on the CIFAR-10 dataset under 30% device variation with only 7.12% area and 5.02% power overhead and no extra latency.
Jingyu He, Yucong Huang, Miguel Angel Lastras-Montaño, Terry Tao Ye, Chi-Ying Tsui, Kwang-Ting Cheng
ASP-DAC2
2021 55nm CMOS Analog Circuit Implementation of LIF and STDP Functions for Low-Power SNNs
abstract
Spiking neural networks (SNNs) demonstrate great potentials to achieve low-power computation for AI applications. SNN uses spike trains, instead of binary bit-steams to encode input and output information, therefore, analog implementation of SNN will have more advantages than digital implementation in terms of power consumption and hardware overheads. Leaky Integrate-and-Fire (LIF) and Spike Timing Dependent Plasticity (STDP) models are the two fundamental mechanisms of SNN operation. In this paper, we propose a 55nm analog CMOS implementation of the LIF and STDP functions. Testing results demonstrate that the circuit can closely imitate the behavior of the LIF and STDP mechanisms, while demanding a much lower power consumption (around 1nJ per spike with the pulse width of 0.5ms). The proposed LIF and STDP circuits can be used as building blocks to construct a complete SNN architecture.
Zhitao Yang, Zhujiang Han, Yucong Huang, Terry Tao Ye
ISLPED3
2020 Analog Circuit Implementation of Neurons with Multiply-Accumulate and ReLU Functions
abstract
Although Artificial Neural Networks (ANNs) are inspired by biological neural systems, most of ANNs today are implemented with digital circuitry and use binary values in computation. In recent years, analog-based neuromorphic system has gained lots of attention as it provides a natural interface for brain-machine interaction. In this paper, we present analog designs of a complete neuron system, where the Multiply-Accumulate (MAC) and Rectified Linear Unit (ReLU) functions are all implemented in analog circuits. The design uses SMIC 55nm standard LP CMOS process node and operates at low supply voltage (1.2 V). The simulation results in SPECTRE demonstrate that the MAC's linear error is no more than 0.5% and total harmonic distortion (THD) is less than 1.6% when the inputs vary from peak (-10 µA) to peak (10 µA) at 10 MHz, the -3dB bandwidth is 288 MHz, the maximum power consumption is 540 µW and the static power consumption is 493 µW under 100MHz input signal frequency. More specifically, our design is resilient to the fluctuation of power supply, which helps to achieve high precision of computation.
Yucong Huang, Zhitao Yang, Jianghan Zhu, Terry Tao Ye
ACM Great Lakes Symposium on VLSI1
2020 Analog Circuit Implementation of LIF and STDP Models for Spiking Neural Networks
abstract
Spiking Neural Networks (SNN) is one special implementation of Artificial Neural Networks (ANN), where the input signals are encoded in the temporal relationship between consecutive spikes (spike trains) instead of real-numbered values. Nevertheless, SNN is believed to be a closer representation of the biological neural system, because it imitates the current spikes that are transmitted between neurons in real biological systems. Practical and simplified SNN models include the Leaky Integrate-and-Fire (LIF) function of the neurons and the Spiking Timing Dependent Plasticity (STDP) function of the synapses. While most ANN architectures can be implemented with digital logic gates, SNN is more suitable to be implemented in analog circuits. In this paper, we propose revised analog circuit implementations of SNN neurons with LIF and STDP functions. Compared with previous works by other researchers, our proposed analog designs use fewer components and can be cascaded to form a complete neural system. The circuits are designed and simulated with SMIC 55nm CMOS LP process. The simulated results demonstrate that the analog neural system can work under a very small current (less than 10 µA) and voltage supply (1.0 Volts), and consumes less power consumption than digital implementations.
Zhitao Yang, Yucong Huang, Jianghan Zhu, Terry Tao Ye
ACM Great Lakes Symposium on VLSI2
2015 Enhanced DCTCP to explicitly inform of packet loss
abstract
DCTCP is designed to address the TCP incast congestion in data center networks. But it may cause TCP timeout due to the lack of prompt packet loss detection, thus restricts the maximum number of concurrent senders. In this paper, an enhanced DCTCP called E-DCTCP is designed based on the notion of ECN. E-DCTCP composes of two phases, i.e. loss detection and backoff retransmission. In the loss detection phase, if any switch detects a packet (e.g. A) to be dropped, it stamps a dedicated ECN flag on A, and immediately routes A back to its sender. In the backoff retransmission phase, if a sender receives any backtracked packet A, it retransmits A based on a simplified binary exponential backoff algorithm. We show by simulations that E-DCTCP provides higher goodput and shorter flow completion time than DCTCP when the number of concurrent senders increases.
Yucong Huang
ICC1