VLDB 2026 Research / reviewers in the wild / expert
Haisheng Zheng
dblp:142/6344
· DBLP profile ↗
21ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-8555-4544ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CombRewriter: Enabling Combinational Logic Simplification in MLIR-Based Hardware CompilerabstractModern Hardware Description Languages (HDLs) play a pivotal role in enabling swift and adaptable hardware development. A hardware compiler translates high-level designer intents into a concrete hardware implementation, the quality of which directly determines ultimate circuit performance. However, current hardware compilers may overlook opportunities for combinational logic simplification, leading to RTL code that contains redundant logic and degrades the Quality of Results (QoR) of the synthesized netlist. This paper presents CombRewriter, a novel approach that incorporates compilation-level optimization techniques into combinational logic simplification. Experimental results demonstrate that the proposed method effectively reduces netlist area. Haisheng Zheng, Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
ASP-DAC | 1 |
| 2025 | Efficient OpAmp Adaptation for Zoom Attention to Golden ContextsabstractLarge language models (LLMs) have shown significant promise in question-answering (QA) tasks, particularly in retrieval-augmented generation (RAG) scenarios and long-context applications. However, their performance is hindered by noisy reference documents, which often distract from essential information. Despite fine-tuning efforts, Transformer-based architectures struggle to prioritize relevant content. This is evidenced by their tendency to allocate disproportionate attention to irrelevant or later-positioned documents. Recent work proposes the differential attention mechanism to address this issue, but this mechanism is limited by an unsuitable common-mode rejection ratio (CMRR) and high computational costs. Inspired by the operational amplifier (OpAmp), we propose the OpAmp adaptation to address these challenges, which is implemented with adapters efficiently. By integrating the adapter into pre-trained Transformer blocks, our approach enhances focus on the golden context without costly training from scratch. Empirical evaluations on noisy-context benchmarks reveal that our Qwen2.5-OpAmp-72B model, trained with our OpAmp adaptation, surpasses the performance of state-of-the-art LLMs, including DeepSeek-V3 and GPT-4o.Our code is available at https://github.com/wuhy68/OpampAdapter. Haoyuan Wu, Rui Ming, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
ACL (1) | 3 |
| 2025 | ChatLS: Multimodal Retrieval-Augmented Generation and Chain-of-Thought for Logic Synthesis Script CustomizationabstractLarge Language Models (LLMs) have demonstrated significant potential in automating the Electronic Design Automation (EDA) process through effective integration with EDA tools. This paper targets the customization of logic synthesis scripts, which is crucial for accommodating the unique characteristics of each design in the EDA workflow. The proposed framework, called ChatLS, integrates multimodal retrieval-augmented generation (RAG) and chain-of-thought (CoT) reasoning, enabling LLMs to collaboratively analyze design features and precisely customize synthesis scripts. Experimental results demonstrate that ChatLS has achieved superior performance in customizing synthesis scripts with a commercial logic synthesis tool. Haisheng Zheng, Haoyuan Wu, Zhuolun He |
DAC | 1 |
| 2025 | iRw: An Intelligent RewritingabstractThis paper proposes a novel machine learning-driven rewriting algorithm to optimize And-Inverter Graphs (AIGs) for refining combinational logic prior to technology mapping. The algorithm, called i Rw, iteratively extracts subcircuits in AIGs and replaces them with more streamlined implementations. These subcircuits are identified using an original extraction algorithm, while the compact implementations are produced through rewriting techniques guided by a machine learning model. This approach efficiently enables the generation of logically equivalent subcircuits with minimal overhead. Experiments on benchmark circuits indicate that the proposed methodology outperforms state-of-the-art AIG rewriting techniques in both quality and runtime. Haisheng Zheng, Haoyuan Wu, Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
DATE | 1 |
| 2025 | MM-GRADE: A Multi-Modal EDA Tool Documentation QA Framework Leveraging Retrieval Augmented GenerationabstractThe complexity of EDA tools necessitates the development of advanced documentation query answering systems to enhance user efficiency and reduce the associated learning curve. Recent innovations in the use of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) for EDA tool documentation have demonstrated significant progress; however, these approaches typically lack the multi-modal capabilities required to effectively handle visual data, such as circuit layout and GUI screenshots provided through user input. To address the concern, we introduce a multi-modal RAG system that incorporates two domain-customized modules: a multi-modal retriever model finetuned by the customized bilevel hard negative mining (BHNM) strategy, and a vision large language model (VLLM) finetuned using a tailored extract-score-answer pipeline. Moreover, we have manually curated ORD-MMBench, a multi-modal QA benchmark comprising 120 high-quality question-document-answer triplets based on OpenROAD documentation. Experimental results demonstrate that our customized RAG framework outperforms state-of-the-art multi-modal RAG flows and models on ORD-MMBench. Yuan Pu 0001, Zhuolun He, Shutong Lin, Jiajun Qin, Xinyun Zhang 0001, Hairuo Han, Haisheng Zheng, Cheng Zhuo, Qi Sun 0002, David Z. Pan, Bei Yu 0001 |
ICCAD | 7 |
| 2025 | Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG AlignmentabstractUnderstanding the structure and function of circuits is crucial for electronic design automation (EDA).
Circuits can be formulated as And-Inverter graphs (AIGs), enabling efficient implementation of representation learning through graph neural networks (GNNs).
Masked modeling paradigms have been proven effective in graph representation learning.
However, masking augmentation to original circuits will destroy their logical equivalence, which is unsuitable for circuit representation learning.
Moreover, existing masked modeling paradigms often prioritize structural information at the expense of abstract information such as circuit function.
To address these limitations, we introduce MGVGA, a novel constrained masked modeling paradigm incorporating masked gate modeling (MGM) and Verilog-AIG alignment (VGA).
Specifically, MGM preserves logical equivalence by masking gates in the latent space rather than in the original circuits, subsequently reconstructing the attributes of these masked gates.
Meanwhile, large language models (LLMs) have demonstrated an excellent understanding of the Verilog code functionality.
Building upon this capability, VGA performs masking operations on original circuits and reconstructs masked gates under the constraints of equivalent Verilog codes, enabling GNNs to learn circuit functions from LLMs.
We evaluate MGVGA on various logic synthesis tasks for EDA and show the superior performance of MGVGA compared to previous state-of-the-art methods.
Our code is available at https://github.com/wuhy68/MGVGA. Haoyuan Wu, Haisheng Zheng, Yuan Pu 0001, Bei Yu 0001 |
ICLR | 2 |
| 2025 | Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design AutomationabstractHaoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
NAACL (Long Papers) | 2 |
| 2025 | On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language UnderstandingabstractLarge language models (LLMs) achieve remarkable performance in code generation tasks.
However, a significant performance disparity persists between popular programming languages (e.g., Python, C++) and others.
To address this capability gap, we leverage the code translation task to train LLMs, thereby facilitating the transfer of coding proficiency across diverse programming languages.
Moreover, we introduce OORL for training, a novel reinforcement learning (RL) framework that integrates on-policy and off-policy strategies.
Within OORL, on-policy RL is applied during code translation, guided by a rule-based reward signal derived from unit tests.
Complementing this coarse-grained rule-based reward, we propose Group Equivalent Preference Optimization (GEPO), a novel preference optimization method.
Specifically, GEPO trains the LLM using intermediate representations (IRs) groups.
LLMs can be guided to discern IRs equivalent to the source code from inequivalent ones, while also utilizing signals about the mutual equivalence between IRs within the group.
This process allows LLMs to capture nuanced aspects of code functionality.
By employing OORL for training with code translation tasks, LLMs improve their recognition of code functionality and their understanding of the relationships between code implemented in different languages.
Extensive experiments demonstrate that our OORL for LLMs training with code translation tasks achieves significant performance improvements on code benchmarks across multiple programming languages. Haoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao, Xueyi Chen, Yikai Yang, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
NeurIPS | 7 |
| 2025 | Energy performance prediction of centrifugal pumps based on adaptive support vector regression
Huican Luo, Peijian Zhou, Jiayi Cui, Haisheng Zheng, Yantian Wang |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | IncreMacro: Incremental Macro Placement RefinementabstractThis article proposes$\textsf {IncreMacro}$, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this,$\textsf {IncreMacro}$utilizes several methods, including kd-tree-based macro diagnosis, gradient-based macro shifting, constraint-graph-based LP for macro legalization, and diffusion-based cell migration. By employing these techniques iteratively,$\textsf {IncreMacro}$meets two critical solution requirements of macro placement: 1) pushing macros toward the chip boundary and 2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into$\textsf {AutoDMP}$and$\textsf {DREAMPlace}~4.0$, and is evaluated on seven RISC-V benchmark circuits and four TILOS macro placement circuit designs at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by$\textsf {AutoDMP}~(\textsf {DREAMPlace}~4.0$), our approach reduces routed wirelength by 15.1% (14.9%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 99.9 (82.6%) and 99.9% (81.3%), and reduces the total power consumption by 4.4% (4.3%). Meanwhile, compared with$\textsf {IncreMacro}$[1], our approach augmented with the cell migration algorithm improves the routed WNS and TNS by 24.7% and 23.1%, and remains the average routed wirelength and total power consumption almost unchanged. Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Jiajun Qin, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | LSTP: A Logic Synthesis Timing PredictorabstractThe ever-growing complexity of modern VLSI circuits brings about a substantial increase in the design cycle. As for logic synthesis, how to efficiently obtain physical characteristics of a design for subsequent design space exploration emerges as a critical issue. In this paper, we propose ${\mathsf{LSTP}}$, an ML-based logic synthesis predictor, which can rapidly predict the post-synthesis timing of a broad range of circuit designs. Specifically, we explicitly take optimization sequences into consideration so that we can comprehend the synergy between optimization passes and their effects on netlists. Experimental results demonstrate that we outperform state-of-the-art remarkably. Haisheng Zheng, Zhuolun He, Fangzhou Liu 0005, Zehua Pei, Bei Yu 0001 |
ASPDAC | 1 |
| 2024 | CBTune: Contextual Bandit Tuning for Logic SynthesisabstractLogic synthesis pre-optimization involves applying a sequence of transformations called synthesis flow to reduce the circuit's Boolean logic graph, like AIG. However, the challenge lies in selecting and arranging these transformations due to the exponentially expanding solution space. In this work, we propose CBTune, a novel online learning framework that utilizes a contextual bandit algorithm to explore the solution space and generate synthesis flows efficiently. We develop the Syn-LinUCB algorithm as the agent, which incorporates circuit characteristics and leverages long-term payoffs to guide decision-making, thus ef-fectively preventing getting trapped in local optima. Experimental results show that our framework achieves the optimal synthesis flow with a lower time cost, substantially reducing the number of AIG nodes and 6-LUTs compared to SOTA approaches. Fangzhou Liu 0005, Zehua Pei, Ziyang Yu 0001, Haisheng Zheng, Zhuolun He, Tinghuan Chen, Bei Yu 0001 |
DATE | 4 |
| 2024 | Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General TasksabstractLarge language models (LLMs) have demonstrated considerable proficiency in general natural language processing (NLP) tasks.Instruction tuning, a successful paradigm, enhances the ability of LLMs to follow natural language instructions and exhibit robust generalization across general tasks.However, these models often encounter performance limitations across multiple tasks due to constrained model capacity.Expanding this capacity during the instruction tuning phase poses significant challenges.To address this issue, we introduce parameter-efficient sparsity crafting (PESC), which crafts dense models into sparse models using the mixture-of-experts (MoE) architecture.PESC integrates adapters into the MoE layers of sparse models, differentiating experts without altering the individual weights within these layers.This method significantly reduces computational costs and GPU memory requirements, facilitating model capacity expansion through a minimal parameter increase when guaranteeing the quality of approximation in function space compared to original sparse upcycling.Our empirical evaluation demonstrates the effectiveness of the PESC method.Using PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.5.Our code is available at https://github.com/wuhy68/ Parameter-Efficient-MoE. Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
EMNLP | 2 |
| 2024 | IncreMacro: Incremental Macro Placement RefinementabstractThis paper proposes IncreMacro, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this, IncreMacro utilizes several methods including kd-tree-based macro diagnosis, gradient-based macro shifting and constraint-graph-based LP for macro legalization. By employing these techniques iteratively, IncreMacro meets two critical solution requirements of macro placement: (1) pushing macros to the chip boundary; and (2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into DREAMPlace and AutoDMP, and is evaluated on several RISC-V benchmark circuits at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by DREAMPlace (AutoDMP), IncreMacro reduces routed wirelength by 6.5% (16.8%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 59.9% (99.6%) and 63.9% (99.9%), and reduces the total power consumption by 3.3% (4.9%). Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
ISPD | 5 |
| 2024 | ChatEDA: A Large Language Model Powered Autonomous Agent for EDAabstractThe integration of a complex set of Electronic Design Automation (EDA) tools to enhance interoperability is a critical concern for circuit designers. Recent advancements in large language models (LLMs) have showcased their exceptional capabilities in natural language processing and comprehension, offering a novel approach to interfacing with EDA tools. This research paper introduces ChatEDA, an autonomous agent for EDA empowered by a large language model, AutoMage, complemented by EDA tools serving as executors. ChatEDA streamlines the design flow from the Register-Transfer Level (RTL) to the Graphic Data System Version II (GDSII) by effectively managing task decomposition, script generation, and task execution. Through comprehensive experimental evaluations, ChatEDA has demonstrated its proficiency in handling diverse requirements, and our fine-tuned AutoMage model has exhibited superior performance compared to GPT-4 and other similar LLMs. Haoyuan Wu, Zhuolun He, Xinyun Zhang 0001, Xufeng Yao, Su Zheng, Haisheng Zheng, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | OpenDRC: An Efficient Open-Source Design Rule Checking Engine with Hierarchical GPU AccelerationabstractDesign rule checking (DRC) is an essential procedure in physical verification, yet few open-source DRC tools are accessible in academia. To fill in the gap, we present OpenDRC, an open-source DRC engine that aims for extremely high efficiency. OpenDRC maintains hierarchical layouts with layer-wise bounding volume hierarchies and performs adaptive row-based partition to identify independent regions for check pruning and/or parallel processing. For common design rules, OpenDRC provides a sequential mode that runs cell-level sweeplines, and a parallel mode that launches edge-based GPU check kernels. Experiments demonstrate that OpenDRC outperforms state-of-the-art multi-threading and GPU-accelerated design rule checkers. The source code is available at https://github.com/opendrc/opendrc. Zhuolun He, Yihang Zuo, Jiaxi Jiang, Haisheng Zheng, Yuzhe Ma, Bei Yu 0001 |
DAC | 4 |
| 2023 | AlphaSyn: Logic Synthesis Optimization with Efficient Monte Carlo Tree SearchabstractRecent years have seen rising research in logic synthesis recipe generation to improve the Quality-of-Result (QoR). However, existing approaches typically have low efficiency and are stuck at local optima. In this work, we propose a logic synthesis optimization framework, AlphaSyn, that incorporates a domain-specific Monte Carlo tree search (MCTS) algorithm. AlphaSyn enables exploration across the entire search space while optimizing sampling points utilization. We further develop a synthesis-specific upper confidence bound for trees (SynUCT) algorithm for the selection phase and a well-designed learning strategy to enhance the stability of the MCTS algorithm. The AlphaSyn algorithm is fully parallelized for efficiency with asynchronous MCTS exploration and significance-base resource allocation. For standard-cell technology mapping on the ASAP 7nm library among other tasks, experimental results show that AlphaSyn outperforms SOTA FlowTune with an average 8.74% area reduction and$\boldsymbol{1.24}\times$runtime speedup. Zehua Pei, Fangzhou Liu 0005, Zhuolun He, Guojin Chen, Haisheng Zheng, Keren Zhu 0001, Bei Yu 0001 |
ICCAD | 5 |
| 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPUabstractOver the past few years, super-resolution (SR) processing has achieved astonishing progress along with the development of deep learning. Nevertheless, the rigorous requirement for real-time inference, especially for video tasks, leaves a harsh challenge for both the model architecture design and the hardware-level implementation. In this article, we propose a hardware-aware acceleration on embedded GPU devices as a full-stack SR deployment framework. The most critical stage with dictionary learning applied in SR flow was analyzed in details and optimized with a tailored dictionary slimming strategy. Moreover, we also delve into the programming architecture of hardware while analyzing the model structure to optimize the computation kernels to reduce inference latency and maximize the throughput given restricted computing power. In addition, we further accelerate the model with 8-bit integer inference by quantizing the weights in the compressed model. An adaptive 8-bit quantization flow for SR task enables the quantized model to achieve a comparable result with the full-precision baselines. With the help of our approaches, the computation and communication bottlenecks in the deep dictionary learning-based SR models can be overcome effectively. The experiments on both edge embedded device NVIDIA NX and 2080Ti prove that our framework exceeds the performance of state-of-the-art NVIDIA TensorRT significantly and can achieve real-time performance. Wenqian Zhao 0002, Qi Sun 0002, Wenbo Li 0002, Haisheng Zheng, Nianjuan Jiang, Jiangbo Lu, Bei Yu 0001, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | GTuner: tuning DNN computations on GPU via graph attention networkabstractIt is an open problem to compile DNN models on GPU and improve the performance. A novel framework, GTuner, is proposed to jointly learn from the structures of computational graphs and the statistical features of codes to find the optimal code implementations. A Graph ATtention network (GAT) is designed as the performance estimator in GTuner. In GAT, graph neural layers are used to propagate the information in the graph and a multi-head self-attention module is designed to learn the complicated relationships between the features. Under the guidance of GAT, the GPU codes are generated through auto-tuning. Experimental results demonstrate that our method outperforms the previous arts remarkably. Qi Sun 0002, Xinyun Zhang 0001, Hao Geng, Yuxuan Zhao 0001, Haisheng Zheng, Bei Yu 0001 |
DAC | 6 |
| 2021 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPUabstractRecent years have witnessed impressive progress in super-resolution (SR) processing. However, its real-time inference requirement sets a challenge not only for the model design but also for the on-chip implementation. In this paper, we implement a full-stack SR acceleration framework on embedded GPU devices. The special dictionary learning algorithm used in SR models was analyzed in detail and accelerated via a novel dictionary selective strategy. Besides, the hardware programming architecture together with the model structure is analyzed to guide the optimal design of computation kernels to minimize the inference latency under the resource constraints. With these novel techniques, the communication and computation bottlenecks in the deep dictionary learning-based SR models are tackled perfectly. The experiments on the edge embedded NVIDIA NX and 2080Ti show that our method outperforms the state-of-the-art NVIDIA TensorRT significantly and can achieve real-time performance. Wenqian Zhao 0002, Qi Sun 0002, Wenbo Li 0002, Haisheng Zheng, Bei Yu 0001, Martin D. F. Wong |
ICCAD | 5 |
| 2013 | A novel micro-deformation monitoring system for large scale structureabstractThis paper describes a novel remote sensing technique for deformation monitoring. This system consists of two parts: a set of radio beacons and a telemetering receiver. The radio beacons are installed apart on the large structure. These beacons have the same carrier modulated by the same PN codes with different delays that exceed one chip period. The distance between any two PN codes is unique so as to distinguish the beacons from each other. The telemetering receiver receives the mixed signal from all the beacons, and uses de-correlation to recover every beacon's carrier. Phase detectors measure the phase difference between all other carriers and the reference carrier. The micro-deformation value Δd is proportional to the phase difference Δθ, and then the deformation curve can be drawn out. The simulation result has proved that the technique proposed by this paper is correct. Andrew G. Dempster, Haiyong Wu, Haisheng Zheng |
IGARSS | 6 |