EDBT 2026 Demo / reviewers in the wild / expert
Yanqing Zhang 0002
dblp:342/9534-2
· DBLP profile ↗
22ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-0836-6766ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 5 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GEM: GPU-Accelerated Emulator-Inspired RTL SimulationabstractIn this paper, we present a GPU-accelerated RTL simulator addressing critical challenges in high-speed circuit verification. Traditional CPU-based RTL simulators struggle with scalability and performance, and while FPGA-based emulators offer acceleration, they are costly and less accessible. Previous GPU-based attempts have failed to speed up RTL simulation due to the heterogeneous nature of circuit partitions, which conflicts with the SIMT (Single Instruction, Multiple Thread) paradigm of GPUs. Inspired by the design of emulators, our approach introduces a novel virtual Very Long Instruction Word (VLIW) architecture, designed for efficient CUDA execution. We also design a flow that maps circuit logic to the architecture in a process analogous to the FPGA CAD flow. This architecture mitigates issues of irregular memory access and thread divergence, unlocking GPU potential for RTL simulation. Our solution achieves up to $64 \times$ speed-up over the best CPU simulators, democratizing high-speed RTL simulation with accessible hardware and establishing a new frontier for GPUaccelerated circuit verification. Zizheng Guo 0001, Yanqing Zhang 0002, Runsheng Wang, Yibo Lin, Haoxing Ren |
DAC | 2 |
| 2025 | ChipVQA: Benchmarking Visual Language Models for Chip DesignabstractLarge-language models (LLMs) have exhibited great potential to assist chip designs and analysis. Recent research and efforts are mainly focusing on text-based tasks including general QA, debugging, design tool scripting, and so on. However, chip design and implementation workflow usually require a visual understanding of diagrams, flow charts, graphs, schematics, waveforms, etc, which demands the development of multimodality foundation models. In this paper, we propose ChipVQA, a benchmark designed to evaluate the capability of visual language models for chip design. ChipVQA includes 142 carefully designed and collected VQA questions covering five chip design disciplines: Digital Design, Analog Design, Architecture, Physical Design and Semiconductor Manufacturing. Unlike existing VQA benchmarks, ChipVQA questions are carefully designed by chip design experts and require indepth domain knowledge and reasoning to solve. We conduct comprehensive evaluations on both open-source and proprietary multimodal models that are greatly challenged by the benchmark suit. ChipVQA is available at https://github.com/phdyang007/chipvqa. Qijing Huang 0001, Nathaniel Ross Pinckney, Walker J. Turner, Wenfei Zhou, Yanqing Zhang 0002, Chia-Tung Ho, Chen-Chia Chang, Haoxing Ren |
DATE | 6 |
| 2025 | SimPart: A Simple Yet Effective Replication-Aided Partitioning Algorithm for Logic Simulation on GPU
Yi-Hua Chung, Shui Jiang, Wan-Luan Lee, Yanqing Zhang 0002, Haoxing Ren, Tsung-Yi Ho, Tsung-Wei Huang |
Euro-Par (3) | 4 |
| 2024 | BoolGebra: Attributed Graph-Learning for Boolean Algebraic ManipulationabstractLogic optimization is an essential stage in the design automation flow for digital systems as the performance of the system at logic level can have significant impacts on the final chip area, timing closure, and the power efficiency of the system. Logic optimization is a technology-independent circuit optimization at the logic level conducted on multi-level technology-independent representations such as And-Inverter-Graphs (AIGs) [1] and Majority-Inverter-Graphs (MIGs) [2] of the digital logic. Existing state-of-the-art (SOTA) Directed-Acyclic-Graphs (DAGs) aware Boolean optimization algorithms, such as structural rewriting (rw) [1], resubstitution (rs) [3], and refactoring (rf) [1] in ABC [4], are conducted on the AIG data structure with a graph-level single optimization concept, i.e., all nodes in the graph have one same fixed optimization opportunity, while overlooking other potential optimization opportunities. [5] proposes orchestrated logic optimization, which is a fine-grained node-level logic optimization method incorporating multiple optimization techniques within a single AIG traversal. However, the enlarged search space pose a significant challenge in searching optimal solutions without domain knowledge. Anthony Agnesina, Yanqing Zhang 0002, Haoxing Ren, Cunxi Yu |
DATE | 3 |
| 2024 | GL0AM: GPU Logic Simulation Using 0-Delay and Re-simulation Acceleration MethodabstractIn this work, we present GL0AM, a novel GPU accelerated logic simulator that performs delay annotated gate-level simulation, supporting a wide range of sequential components, including SRAMs, and simulation scenarios. We propose a methodology to perform simulation in 2 phases in order to increase the parallelism exposed to the GPU, where the first phase performs 0-delay cycle simulation that records sequential gate (flops, latches, and clock-gates) waveform results, and the second phase performs re-simulation to attain the remaining combinational gate waveforms. We propose to use netlist graph partitioning to minimize synchronization overheads and increase parallelism during the 0-delay simulation phase. GL0AM achieves simulation execution time speedup of 19-537X on an NVIDIA H100 GPU when compared to a commercial simulator across a diverse set of benchmarks. GL0AM provides a 2-5X speedup, iso- GPU platform and benchmark, over a recent state-of-the-art GPU-accelerated logic simulator.1 Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 1 |
| 2023 | GenFuzz: GPU-accelerated Hardware Fuzzing using Genetic Algorithm with Multiple InputsabstractHardware fuzzing has emerged as a promising automatic verification technique to efficiently discover and verify hardware vulnerabilities. However, hardware fuzzing can be extremely time-consuming due to compute-intensive iterative simulations. While recent research has explored several approaches to accelerate hardware fuzzing, nearly all of them are limited to single-input fuzzing using one thread of a CPU-based simulator. As a result, we propose Gen-Fuzz, a GPU-accelerated hardware fuzzer using a genetic algorithm with multiple inputs. Measuring experimental results on a real industrial design, we show that GenFuzz running on a single A6000 GPU and eight CPU cores achieves 80× runtime speed-up when compared to state-of-the-art hardware fuzzers. Dian-Lun Lin, Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany, Shih-Hsin Wang, Tsung-Wei Huang |
DAC | 2 |
| 2022 | GATSPI: GPU accelerated gate-level simulation for power improvementabstractIn this paper, we present GATSPI, a novel GPU accelerated logic gate simulator that enables ultra-fast power estimation for industry-sized ASIC designs with millions of gates. GATSPI is written in PyTorch with custom CUDA kernels for ease of coding and maintainability. It achieves simulation kernel speedup of up to 1668X on a single-GPU system and up to 7412X on a multiple-GPU system when compared to a commercial gate-level simulator running on a single CPU core. GATSPI supports a range of simple to complex cell types from an industry standard cell library and SDF conditional delay statements without requiring prior calibration runs and produces industry-standard SAIF files from delay-aware gate-level simulation. Finally, we deploy GATSPI in a glitch-optimization flow, achieving a 1.4% power saving with a 449X speedup in turnaround time compared to a similar flow using a commercial simulator. Yanqing Zhang 0002, Haoxing Ren, Akshay Sridharan, Brucek Khailany |
DAC | 1 |
| 2022 | Why are Graph Neural Networks Effective for EDA Problems?: (Invited Paper)abstractIn this paper, we discuss the source of effectiveness of Graph Neural Networks (GNNs) in EDA, particularly in the VLSI design automation domain. We argue that the effectiveness comes from the fact that GNNs implicitly embed the prior knowledge and inductive biases associated with given VLSI tasks, which is one of the three approaches to make a learning algorithm physics-informed. These inductive biases are different to those common used in GNNs designed for other structured data, such as social networks and citation networks. We will illustrate this principle with several recent GNN examples in the VLSI domain, including predictive tasks such as switching activity prediction, timing prediction, parasitics prediction, layout symmetry prediction, as well as optimization tasks such as gate sizing and macro and cell transistor placement. We will also discuss the challenges of applications of GNN and the opportunity of applying self-supervised learning techniques with GNN for VLSI optimization. Haoxing Ren, Siddhartha Nath, Yanqing Zhang 0002, Hao Chen 0059 |
ICCAD | 3 |
| 2022 | From RTL to CUDA: A GPU Acceleration Flow for RTL Simulation with Batch StimulusabstractHigh-throughput RTL simulation is critical for verifying today’s highly complex SoCs. Recent research has explored accelerating RTL simulation by leveraging event-driven approaches or partitioning heuristics to speed up simulation on a single stimulus. To further accelerate throughput performance, industry-quality functional verification signoff must explore running multiple stimulus (i.e., batch stimulus) simultaneously, either with directed tests or random inputs. In this paper, we propose RTLFlow, a GPU-accelerated RTL simulation flow with batch stimulus. RTLflow first transpiles RTL into CUDA kernels that each simulates a partition of the RTL simultaneously across multiple stimulus. It also leverages CUDA Graph and pipeline scheduling for efficient runtime execution. Measuring experimental results on a large industrial design (NVDLA) with 65536 stimulus, we show that RTLflow running on a single A6000 GPU can achieve a 40 × runtime speed-up when compared to an 80-thread multi-core CPU baseline. Dian-Lun Lin, Haoxing Ren, Yanqing Zhang 0002, Brucek Khailany, Tsung-Wei Huang |
ICPP | 3 |
| 2021 | MAVIREC: ML-Aided Vectored IR-Drop Estimation and ClassificationabstractVectored IR drop analysis is a critical step in chip signoff that checks the power integrity of an on-chip power delivery network. Due to the prohibitive runtimes of dynamic IR drop analysis, the large number of test patterns must be whittled down to a small subset of worst-case IR vectors. Unlike the traditional slow heuristic method that select a few vectors with incomplete coverage, MAVIREC uses machine learning techniques -- 3D convolutions and regression-like layers -- for accurately recommending a larger subset of test patterns that exercise worst-case scenarios. In under 30 minutes, MAVIREC profiles 100K-cycle vectors and provides better coverage than a state-of-the-art industrial flow. Further, MAVIREC's IR drop predictor shows 10x speedup with under 4mV RMSE relative to an industrial flow. Vidya A. Chhabria, Yanqing Zhang 0002, Haoxing Ren, Ben Keller, Brucek Khailany, Sachin S. Sapatnekar |
DATE | 2 |
| 2021 | 2021 ICCAD CAD Contest Problem C: GPU Accelerated Logic RewritingabstractLogic rewriting is an important optimization function that can improve Quality of Results (QoR) in modern VLSI circuits. This optimization function usually has a greedy approach and involves steps such as graph traversal, cut computation and ranking, and functional matching. For logic rewriting to be effective in improving the QoR, there should be many local rewriting iterations which can be very slow for industrial level benchmark circuits. One effective solution to speed up the logic rewriting operation is to upload its time consuming steps to Graphics Processing Units (GPUs) to benefit from massively parallel computations that is available there. In this regard, the present contest problem studies the possibility of using GPUs in accelerating a classical logic rewriting function. State-of-the-art large-scale open-source benchmark circuits as well as industrial-level designs will be used to test the GPU accelerated logic rewriting function. Ghasem Pasandi, Sreedhar Pratty, Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 4 |
| 2020 | FIST: A Feature-Importance Sampling and Tree-Based Method for Automatic Design Flow Parameter TuningabstractDesign flow parameters are of utmost importance to chip design quality and require a painfully long time to evaluate their effects. In reality, flow parameter tuning is usually performed manually based on designers' experience in an ad hoc manner. In this work, we introduce a machine learning-based automatic parameter tuning methodology that aims to find the best design quality with a limited number of trials. Instead of merely plugging in machine learning engines, we develop clustering and approximate sampling techniques for improving tuning efficiency. The feature extraction in this method can reuse knowledge from prior designs. Furthermore, we leverage a state-of-the-art XGBoost model and propose a novel dynamic tree technique to overcome overfitting. Experimental results on benchmark circuits show that our approach achieves 25% improvement in design quality or 37% reduction in sampling cost compared to random forest method, which is the kernel of a highly cited previous work. Our approach is further validated on two industrial designs. By sampling less than 0.02% of possible parameter sets, it reduces area by 1.83% and 1.43% compared to the best solutions hand-tuned by experienced designers. Zhiyao Xie, Guanqi Fang, Yu-Hung Huang, Haoxing Ren, Yanqing Zhang 0002, Brucek Khailany, Shao-Yun Fang, Jiang Hu 0001, Yiran Chen 0001, Erick Carvajal Barboza |
ASP-DAC | 5 |
| 2020 | GRANNITE: Graph Neural Network Inference for Transferable Power EstimationabstractThis paper introduces GRANNITE, a GPU-accelerated novel graph neural network (GNN) model for fast, accurate, and transferable vector-based average power estimation. During training, GRANNITE learns how to propagate average toggle rates through combinational logic: a netlist is represented as a graph, register states and unit inputs from RTL simulation are used as features, and combinational gate toggle rates are used as labels. A trained GNN model can then infer average toggle rates on a new workload of interest or new netlists from RTL simulation results in a few seconds. Compared to traditional power analysis using gate-level simulations, GRANNITE achieves >18.7X speedup with an error of only <; 5.5% across a diverse set of benchmark circuits. Compared to a GPU-accelerated conventional probabilistic switching activity estimation approach, GRANNITE achieves much better accuracy (on average 25.9% lower error) at similar runtimes. Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
DAC | 1 |
| 2020 | Opportunities for RTL and Gate Level Simulation using GPUs (Invited Talk)abstractThis paper summarizes the opportunities in accelerating simulation on parallel processing hardware platforms such as GPUs. First, we give a summary of prior art. Then, we propose the idea that coding frameworks usually used for popular machine learning (ML) topics, such as PyTorch/DGL.ai, can also be used for exploring simulation purposes. We demo a crude oblivious two-value cycle gate-level simulator using the higher level ML framework APIs that exhibits >20X speedup, despite its simplistic construction. Next, we summarize recent advances in GPU features that may provide additional opportunities to further state-of-the-art results. Finally, we conclude and touch upon some potential areas for furthering research into the topic of GPU accelerated simulation. Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 1 |
| 2020 | Problem C: GPU Accelerated Logic Re-simulation : (Invited Talk)abstractLogic "re"-simulation can be defined as gate level simulation where the input waveforms at every primary input and pseudo-primary input (such as register/RAM outputs) are known. Such waveforms could come from the unit's RTL simulation trace or Automatic Test Pattern Generation (ATPG) vectors. This type of simulation is useful in doing functional verification on gate level netlists and power analysis, since we can take the known trace on all primary and pseudo-primary inputs, re-simulate the trace using propagation of signals through timing-aware gate-level combinational logic, and verify that results at the primary and pseudo-primary outputs match the reference RTL simulation results. However, gate level simulation is usually much slower than RTL simulation. Thus, there is motivation for faster solutions. In this contest, we ask contestants to use Graphic Processing Units (GPUs) to speedup the re-simulation task. Yanqing Zhang 0002, Haoxing Ren, Ben Keller, Brucek Khailany |
ICCAD | 1 |
| 2019 | PRIMAL: Power Inference using Machine LearningabstractThis paper introduces PRIMAL, a novel learning-based framework that enables fast and accurate power estimation for ASIC designs. PRIMAL trains machine learning (ML) models with design verification testbenches for characterizing the power of reusable circuit building blocks. The trained models can then be used to generate detailed power profiles of the same blocks under different workloads. We evaluate the performance of several established ML models on this task, including ridge regression, gradient tree boosting, multi-layer perceptron, and convolutional neural network (CNN). For average power estimation, ML-based techniques can achieve an average error of less than 1% across a diverse set of realistic benchmarks, outperforming a commercial RTL power estimation tool in both accuracy and speed (15x faster). For cycle-by-cycle power estimation, PRIMAL is on average 50x faster than a commercial gate-level power analysis tool, with an average error less than 5%. In particular, our CNN-based method achieves a 35x speed-up and an error of 5.2% for cycle-by-cycle power estimation of a RISC-V processor core. Furthermore, our case study on a NoC router shows that PRIMAL can achieve a small estimation error of 4.5% using cycle-approximate traces from SystemC simulation. Haoxing Ren, Yanqing Zhang 0002, Ben Keller, Brucek Khailany, Zhiru Zhang |
DAC | 3 |
| 2019 | A 0.11 PJ/OP, 0.32-128 Tops, Scalable Multi-Chip-Module-Based Deep Neural Network Accelerator Designed with A High-Productivity vlsi MethodologyabstractThis article consists of a collection of slides from the author's conference presentation. Rangharajan Venkatesan, Sophia Shao, Brian Zimmer, Jason Clemons, Matthew Fojtik, Nan Jiang 0009, Ben Keller, Alicia Klinefelter, Nathaniel Ross Pinckney, Priyanka Raina, Stephen G. Tell, Yanqing Zhang 0002, William J. Dally, Joel S. Emer, C. Thomas Gray, Stephen W. Keckler, Brucek Khailany |
Hot Chips Symposium | 12 |
| 2019 | MAGNet: A Modular Accelerator Generator for Neural NetworksabstractDeep neural networks have been adopted in a wide range of application domains, leading to high demand for inference accelerators. However, the high cost associated with ASIC hardware design makes it challenging to build custom accelerators for different targets. To lower design cost, we propose MAGNet, a modular accelerator generator for neural networks. MAGNet takes a target application consisting of one or more neural networks along with hardware constraints as input and produces synthesizable RTL for a neural network accelerator ASIC as well as valid mappings for running the target networks on the generated hardware. MAGNet consists of three key components: (i) MAGNet Designer, a highly configurable architectural template designed in C++ and synthesizable by high-level synthesis tools. MAGNet Designer supports a wide range of design-time parameters such as different data formats, diverse memory hierarchies, and dataflows. (ii) MAGNet Mapper, an automated framework for exploring different software mappings for executing a neural network on the generated hardware. (iii) MAGNet Tuner, a design space exploration framework encompassing the designer, the mapper, and a deep learning framework to enable fast design space exploration and co-optimization of architecture and application. We demonstrate the utility of MAGNet by designing an inference accelerator optimized for image classification application using three different neural networks-AlexNet, ResNet, and DriveNet. MAGNet-generated hardware is highly efficient and leverages a novel multi-level dataflow to achieve 40 fJ/op and 2.8 TOPS/mm2in a 16nm technology node for the ResNet-50 benchmark with <; 1% accuracy loss on the ImageNet dataset. Rangharajan Venkatesan, Sophia Shao, Miaorong Wang, Jason Clemons, Steve Dai, Matthew Fojtik, Ben Keller, Alicia Klinefelter, Nathaniel Ross Pinckney, Priyanka Raina, Yanqing Zhang 0002, Brian Zimmer, William J. Dally, Joel S. Emer, Stephen W. Keckler, Brucek Khailany |
ICCAD | 11 |
| 2019 | Simba: Scaling Deep-Learning Inference with Multi-Chip-Module-Based ArchitectureabstractPackage-level integration using multi-chip-modules (MCMs) is a promising approach for building large-scale systems. Compared to a large monolithic die, an MCM combines many smaller chiplets into a larger system, substantially reducing fabrication and design costs. Current MCMs typically only contain a handful of coarse-grained large chiplets due to the high area, performance, and energy overheads associated with inter-chiplet communication. This work investigates and quantifies the costs and benefits of using MCMs with fine-grained chiplets for deep learning inference, an application area with large compute and on-chip storage requirements. To evaluate the approach, we architected, implemented, fabricated, and tested Simba, a 36-chiplet prototype MCM system for deep-learning inference. Each chiplet achieves 4 TOPS peak performance, and the 36-chiplet MCM package achieves up to 128 TOPS and up to 6.1 TOPS/W. The MCM is configurable to support a flexible mapping of DNN layers to the distributed compute and storage units. To mitigate inter-chiplet communication overheads, we introduce three tiling optimizations that improve data locality. These optimizations achieve up to 16% speedup compared to the baseline layer mapping. Our evaluation shows that Simba can process 1988 images/s running ResNet-50 with batch size of one, delivering inference latency of 0.50 ms. Sophia Shao, Jason Clemons, Rangharajan Venkatesan, Brian Zimmer, Matthew Fojtik, Nan Jiang 0009, Ben Keller, Alicia Klinefelter, Nathaniel Ross Pinckney, Priyanka Raina, Stephen G. Tell, Yanqing Zhang 0002, William J. Dally, Joel S. Emer, C. Thomas Gray, Brucek Khailany, Stephen W. Keckler |
MICRO | 12 |
| 2018 | A modular digital VLSI flow for high-productivity SoC designabstractA high-productivity digital VLSI flow for designing complex SoCs is presented. The flow includes high-level synthesis tools, an object-oriented library of synthesizable SystemC and C++ components, and a modular VLSI physical design approach based on fine-grained globally asynchronous locally synchronous (GALS) clocking. The flow was demonstrated on a 16nm FinFET testchip targeting machine learning and computer vision. Brucek Khailany, Evgeni Khmer, Rangharajan Venkatesan, Jason Clemons, Joel S. Emer, Matthew Fojtik, Alicia Klinefelter, Michael Pellauer, Nathaniel Ross Pinckney, Sophia Shao, Shreesha Srinath, Christopher Torng, Sam Likun Xi, Yanqing Zhang 0002, Brian Zimmer |
DAC | 14 |
| 2012 | Body Sensor Networks: A Holistic Approach From Silicon to UsersabstractBody sensor networks (BSNs) are emerging cyber–physical systems that promise to improve quality of life through improved healthcare, augmented sensing and actuation for the disabled, independent living for the elderly, and reduced healthcare costs. However, the physical nature of BSNs introduces new challenges. The human body is a highly dynamic physical environment that creates constantly changing demands on sensing, actuation, and quality of service (QoS). Movement between indoor and outdoor environments and physical movements constantly change the wireless channel characteristics. These dynamic application contexts can also have a dramatic impact on data and resource prioritization. Thus, BSNs must simultaneously deal with rapid changes to both top–down application requirements and bottom–up resource availability. This is made all the more challenging by the wearable nature of BSN devices, which necessitates a vanishingly small size and, therefore, extremely limited hardware resources and power budget. Current research is being performed to develop new principles and techniques for adaptive operation in highly dynamic physical environments, using miniaturized, energy-constrained devices. This paper describes a holistic cross-layer approach that addresses all aspects of the system, from low-level hardware design to higher level communication and data fusion algorithms, to top-level applications. Benton H. Calhoun, John C. Lach, John A. Stankovic, David D. Wentzloff, Kamin Whitehouse, Adam T. Barth, Jonathan K. Brown, Qiang Li 0025, Nathan E. Roberts, Yanqing Zhang 0002 |
Proc. IEEE | 11 |
| 2010 | System design principles combining sub-threshold circuit and architectures with energy scavenging mechanismsabstractUltra low power (ULP) circuits and energy scavenging mechanisms, though conceptually appealing, have been mainly studied in isolation to date. In this paper, we observe energy harvesting prototypes to derive system-level models and to reveal practical issues specific to various types of energy harvesting systems. Our models predict how design decisions affect overall lifetime. We use our model to derive system driven principles for optimizing architecture, voltage selection, and sub-threshold circuit designs across different types of power harvesting systems. Benton H. Calhoun, Sudhanshu Khanna, Yanqing Zhang 0002, Joseph F. Ryan 0002, Brian P. Otis |
ISCAS | 3 |