Guangyao Wang

dblp:229/6222 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Noise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-Memory
abstract
Diffusion models achieve state-of-the-art image generation but impose heavy computational burdens on digital computers. Compute-in-memory (CIM) architectures offer promising acceleration, but inherent noise causes severe performance degradation through weight perturbations. We find that reducing sampling steps improves robustness but limits generation versatility, and that noise at earlier steps causes more severe degradation due to error accumulation. Based on these insights, we propose EtaMix, a novel noise-aware sampling strategy that interpolates between stochastic and deterministic sampling without requiring training or hardware modifications. EtaMix applies more stochastic sampling initially to offset weight perturbations, then gradually transitions to deterministic sampling. Experimental results show EtaMix achieves up to 2.01× and 5.12× FID improvements under different noise conditions for DDPM and DDIM, respectively.
Yuannuo Feng, Wenyong Zhou, Yuexi Lv, Guangyao Wang, Zhengwu Liu, Ngai Wong 0001, Wang Kang 0001
DATE5
2026 NoiseGuard: A Comprehensive Framework With Noise Modeling, Noise-Aware Training, and Noise Compensation for In-Memory Computing SoC
Guangyao Wang, Yizhe Chen, Yuexi Lv, Yuannuo Feng, Jenny Ma, Saiya Wang, Guilin Zhao, Yong Pei, Minghua Tang, Wang Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 HyIMC: Analog-Digital Hybrid In-Memory Computing SoC for High-Quality Low-Latency Speech Enhancement
abstract
In-memory computing (IMC) holds significant promise for accelerating deep learning-based speech enhancement (DL-SE). However, existing IMC architectures face challenges in simultaneously achieving high precision, energy efficiency, and the necessary parallelism for DL-SE's inherent temporal dependencies. This paper introduces HyIMC, a novel hybrid analog-digital IMC architecture designed to address these limitations. HyIMC features: 1) a hybrid analog-digital design optimized for DL-SE algorithms; 2) a schedule controller that efficiently manages recurrent dataflow within skip connections; and 3) non-key dimension shrinkage, a model compression technique that preserves accuracy. Implemented on a 40nm eFlash-based IMC SoC prototype, HyIMC achieves 160 TOPS/W energy efficiency, compresses the DL-SE model size by ~600%, improves the feature of merit by ~1200%, and enhances perceptual evaluation of speech quality by ~120%.
Wanru Mao, Guangyao Wang, Tianshuo Bai, Jingcheng Gu, Xitong Yang, Aifei Zhang, Xiaohang Wei, Wang Kang 0001
DATE3
2025 Towards Accurate Characterization of In-Memory Computing Non-Idealities: A Physics & Data Co-Driven Generative Framework
abstract
Analog in-memory computing (IMC) promises unprecedented energy efficiency for deep learning acceleration, but suffers from non-idealities that severely degrade inference accuracy in fabricated chips. Therefore, accurate modeling of these non-idealities becomes significant. In this work, we present PDGM-IMC, the first physics and data co-driven generative framework for IMC non-idealities characterizing. Unlike traditional physical models that fail to model complex non-ideality behaviors, or black-box neural networks that lack interpretability and generalization, PDGM-IMC leverages normalizing flows with custom transformations directly derived from device physics principles. This novel approach enables explicit modeling of complex probability distributions, spatial correlations, and die-to-die variations that previous methods could not capture. Validated on multiple dies of a commercial eFlash-based IMC SoC, PDGM-IMC improves modeling accuracy by 4.6× for the input circuit and IMC array and by 2.0× for the output circuit, significantly outperforming existing approaches. By extracting the statistical signature of fabricated chips, PDGM-IMC enables accurate pre-silicon prediction of post-silicon behavior, fundamentally transforming hardware-aware neural network optimization for analog accelerators. The source code and the pre-trained models are publicly available at https://github.com/BUAA-BASIC-Lab/PDGM-IMC.
Jing Kou, Guangyao Wang, Saiya Wang, Yuexi Lv, Xinghao Cui, Wei. W. Xing, Wang Kang 0001
ICCAD2
2025 Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
abstract
Recent advancements in large language models (LLMs) have significantly improved text-to-SQL systems. However, most datasets and LLM-based methods tend to focus narrowly on SQL generation, often neglecting the complexities inherent in real-world conversational queries. This oversight can result in unreliable responses, particularly for ambiguous questions that cannot be directly addressed with SQL. To address this gap, we propose MMSQL, a comprehensive test suite designed to evaluate LLMs’ question classification and SQL generation capabilities by simulating real-world scenarios with diverse question types and multi-turn Q&A interactions. Utilizing MMSQL, we assessed the performance of popular LLMs, including both open-source and closed-source models, and identified key factors influencing their performance in these contexts. Furthermore, we introduce an LLM-based multi-agent framework that employs specialized agents to identify question types and determine appropriate answering strategies. Experimental results demonstrate that this method effectively enhances baseline models’ ability to handle diverse question types in conversational scenarios. Our approach simultaneously considers multiple question types and multi-turn interactions, providing a new, realistic perspective, and offering a valuable advancement toward more reliable and versatile text-to-SQL systems. Our dataset and code are publicly available at https://mcxiaoxiao.github.io/MMSQL.
Ziming Guo, Yinggang Sun, Guangyao Wang
IJCNN5
2025 Inter-layer explainable variational autoencoder model for multivariate time series anomaly detection
Guangyao Wang, Wenzhi Yang, Guoyin Wang 0001
Eng. Appl. Artif. Intell.2
2024 An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design Framework
abstract
Despite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection.
Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Weighted graph convolution over dependency trees for nontaxonomic relation extraction on public opinion information
Guangyao Wang, Shengquan Liu, Fuyuan Wei
Appl. Intell.1
2022 Automated universal fractures detection in X-ray images based on deep learning approach
Shuzhen Lu, Sheng-Sheng Wang 0001, Guangyao Wang
Multim. Tools Appl.3