EDBT 2026 Demo / reviewers in the wild / expert
Yikang Ouyang
dblp:357/2639
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0007-0714-9501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space ReductionabstractMultiplication is a fundamental operation in many applications, and multipliers are widely adopted in various circuits. However, optimizing multipliers is challenging due to the extensive design space. In this article, we propose a multiplier design optimization framework based on reinforcement learning. We utilize matrix and tensor representations for the compressor tree of a multiplier, enabling seamless integration of convolutional neural networks as the agent network. The agent optimizes the multiplier structure using a Pareto-driven reward customized to balance area and delay. Furthermore, we enhance the original framework with parallel reinforcement learning and design space pruning techniques and extend its capability to optimize fused multiply-accumulate designs. Experiments conducted on different bit widths of multipliers demonstrate that multipliers produced by our approach outperform all baseline designs in terms of area, power, and delay. The performance gain is further validated by comparing the area, power, and delay of processing element arrays using multipliers from our approach and baseline approaches. Dongsheng Zuo, Jiadong Zhu, Yikang Ouyang, Yuzhe Ma |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | SMART-GPO: Gate-Level Sensitivity Measurement with Accurate Estimation for Glitch Power OptimizationabstractDynamic power consumption is a significant concern in modern integrated circuits. This issue is primarily caused by signal toggling, including unwanted toggles known as glitches. With the number of operations increasing in circuits, glitches can lead to significant additional dynamic power. This paper presents SMART-GPO, a novel framework that efficiently and accurately estimates and reduces glitch power. Our approach samples cycles for accurate glitch estimation, followed by gate-sizing and Vth assignment to optimize glitch power based on sensitivity measurements. We validated SMART-GPO on the Berkeley Out-of-Order Machine (BOOM) and Rocket SoCs with TSMC N28 technology. It achieves a mean absolute percentage error (MAPE) of 2% on glitch power estimation when running power analysis on only 1% simulation cycles. The optimization results demonstrate that our framework reduces glitch power by more than 9%, which outperforms previous approaches substantially. Yikang Ouyang, Yuchao Wu, Dongsheng Zuo, Subhendu Roy, Tinghuan Chen, Zhiyao Xie, Yuzhe Ma |
ASP-DAC | 1 |
| 2025 | Efficient Continuous Logic Optimization with Diffusion ModelabstractThe logic synthesis optimization flow is crucial to the quality of results (QoR), which applies a sequence of transformations to a design. Recently, there has been a growing focus on the automatic optimization of synthesis flows to improve QoR, utilizing techniques such as Bayesian optimization and reinforcement learning, which may fall short in efficiency due to the exponentially large search space. In contrast, continuous optimization offers notable efficiency advantages by leveraging the explicit gradient. However, despite its potential, several significant concerns remain to be addressed. On one hand, it is essential to obtain a reliable gradient. On the other hand, a major challenge arises from the fact that searching within a continuous space can yield solutions that deviate from feasible ones. In this paper, we propose an efficient approach to optimize synthesis sequences within a continuous latent space. Specifically, the gradient information is derived from a QoR surrogate model, while the discrepancies between solutions and feasible transformations are minimized by a diffusion model. Experimental results on extensive benchmarks demonstrate that the proposed method not only achieves lower area and delay but also improves efficiency by 5 X to 130 X, compared with previous methods. Yikang Ouyang, Jiadong Zhu, Tinghuan Chen, Yuzhe Ma |
DAC | 1 |
| 2024 | RL-OPC: Mask Optimization With Deep Reinforcement LearningabstractMask optimization is a vital step in the VLSI manufacturing process in advanced technology nodes. As one of the most representative techniques, optical proximity correction (OPC) is widely applied to enhance printability. Since conventional OPC methods consume prohibitive computational overhead, recent research has applied machine learning techniques for efficient mask optimization. However, existing discriminative learning models rely on a given dataset for supervised training, and generative learning models usually leverage a proxy optimization objective for end-to-end learning, which may limit the feasibility. In this article, we pioneer introducing the reinforcement learning (RL) model for mask optimization, which directly optimizes the preferred objective without leveraging a differentiable proxy. Intensive experiments show that our method outperforms state-of-the-art solutions, including academic approaches and commercial toolkits. Xiaoxiao Liang, Yikang Ouyang, Bei Yu 0001, Yuzhe Ma |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | RL-MUL: Multiplier Design Optimization with Deep Reinforcement LearningabstractMultiplication is a fundamental operation in many applications, and multipliers are widely adopted in various circuits. However, optimizing multipliers is challenging and non-trivial due to the huge design space. In this paper, we propose RL-MUL, a multiplier design optimization framework based on reinforcement learning. Specifically, we utilize matrix and tensor representations for the compressor tree of a multiplier, based on which the convolutional neural networks can be seamlessly incorporated as the agent network. The agent can learn to adjust the multiplier structure based on a Pareto-driven reward which is customized to accommodate the trade-off between area and delay. Experiments are conducted on different bit widths of multipliers. The results demonstrate that the multipliers produced by RL-MUL dominate all baseline designs in terms of both area and delay. The performance gain of RL-MUL is further validated by comparing the area and delay of processing element arrays using multipliers from RL-MUL and baseline approaches. Dongsheng Zuo, Yikang Ouyang, Yuzhe Ma |
DAC | 2 |