Qijun Zhang

dblp:78/4722 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 7 first-author · 19 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ReadyPower: A Reliable, Interpretable, and Handy Architectural Power Model Based on Analytical Framework
abstract
Power is a primary objective in modern processor design, requiring accurate yet efficient power modeling techniques. Architecturelevel power models are necessary for early power optimization and design space exploration. However, classical analytical architecture-level power models (e.g., McPAT) suffer from significant inaccuracies. Emerging machine learning (ML)-based power models, despite their superior accuracy in research papers, are not widely adopted in the industry. In this work, we point out three inherent limitations of ML-based power models: unreliability, limited interpretability, and difficulty in usage. This work proposes a new analytical power modeling framework named ReadyPower, which is ready-for-use by being reliable, interpretable, and handy. We observe that the root cause of the low accuracy of classical analytical power models is the discrepancies between the real processor implementation and the processor’s analytical model. To bridge the discrepancies, we introduce architecture-level, implementationlevel, and technology-level parameters into the widely adopted McPAT analytical model to build ReadyPower. The parameters at three different levels are decided in different ways. In our experiment, averaged across different training scenarios, ReadyPower achieves $\gt20 \%$ lower mean absolute percentage error (MAPE) and $\gt0.2$ higher correlation coefficient R compared with the ML-based baselines, on both BOOM and XiangShan CPU architectures.
Qijun Zhang, Shang Liu 0006, Yao Lu 0031, Mengming Li, Zhiyao Xie
ASP-DAC1
2026 Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
abstract
Tensor parallelism (TP) in large-scale LLM inference and training introduces frequent collective operations that dominate inter-GPU communication. While in-switch computing, exemplified by NVLink SHARP (NVLS), accelerates collective operations by reducing redundant data transfer, its communication-centric design philosophy introduces the mismatch between its communication mode and the memory semantic requirement of LLM's computation kernel. Such a mismatch isolates the compute and communication phases, resulting in underutilized resources and limited overlap in multi-GPU systems. To address the limitation, we propose CAIS, the first ComputeAware In-Switch computing framework that aligns communication modes with computation's memory semantics requirement. CAIS consists of three integral techniques: (1) compute-aware ISA and microarchitecture extension to enable compute-aware in-switch computing. (2) merging-aware TB (Thread Block) coordination to improve the temporal alignment for efficient request merging. (3) graph-level dataflow optimizer to achieve a tight cross-kernel overlap. Evaluations on LLM workloads show that CAIS achieves$1.38 \times$average end-to-end training speedup over the SOTA NVLS-enabled solution, and$1.61 \times$over T3, the SOTA compute-communicate overlap solutions but do not leverage NVLS, demonstrating its effectiveness in accelerating TP on multi-GPU systems.
Chen Zhang 0001, Qijun Zhang, Zhuoshan Zhou, Yijia Diao, Zhipeng Tu, Zhuoran Song, Zhigang Ji, Jingwen Leng, Minyi Guo
HPCA2
2026 ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses
abstract
Irregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns, their effectiveness fundamentally depends on temporal address recurrence and large metadata storage. When memory addresses exhibit weak or no recurrence, as in indirect memory accesses, temporal prefetchers achieve limited performance gains while incurring substantial storage overhead. This paper proposes Instruction-Correlation Prefetching (ICP), a new hardware prefetching mechanism that exploits instruction-level correlations rather than memory-address correlations to handle irregular memory accesses. ICP observes that although memory addresses may not repeat, the instructions generating them often recur with stable data-dependency relationships. By learning these persistent instruction correlations, ICP speculatively computes and prefetches future irregular accesses using the execution results of their correlated predecessors. Across irregular SPEC CPU and GAP benchmarks, ICP outperforms the state-of-the-art temporal prefetcher Triangel by 14.0% and the indirect prefetcher DMP by 6.0%, while requiring only 2.1 KB of hardware storage, over three orders of magnitude smaller than temporal prefetchers.
Mengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang, Xiangfeng Sun, Ceyu Xu, Yuan Xie 0001, Shang Liu 0006, Zhiyao Xie
ISCA4
2026 Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
abstract
Mixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP) becomes a performance challenge. We observe substantial redundant inter-GPU data transfers in MoE that can be potentially addressed by in-switch computing. Unfortunately, the existing solution, NVLink SHARP (NVLS), can only support static collectives with regular patterns, incapable of dynamic communication with irregular patterns in MoE. To bridge the functionality gap, we propose DySHARP, an integral dynamic in-switch computing solution to accelerate MoE, encompassing both communication primitives and communication-aware scheduling: 1) Dynamic multimem addressing co-designs ISA, architecture, and runtime, as a dynamic extension to NVLS, reducing redundant traffic. However, the resulting traffic reduction is inherently asymmetric between two directions, preventing it from directly translating into speedup. 2) Token-centric kernel fusion deeply fuses the dispatch-computation-combine pipeline, resolving this asymmetry to translate traffic reduction into actual speedup. Compared with the state-of-the-art solution, DySHARP achieves up to 1.79× speedup.
Qijun Zhang, Chen Zhang 0001, Zhuoshan Zhou, Zhipeng Tu, Guangyu Sun 0003, Zhiyao Xie, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He 0002, Minyi Guo
ISCA1
2026 MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
Zhuoshan Zhou, Chen Zhang 0001, Qijun Zhang, Zhe Zhou 0002, Zhipeng Tu, Guangyu Sun 0003, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He 0002, Minyi Guo
ISCA4
2026 How Assisted Driving-Based Intelligent Transportation Systems Promote Sustainability: Insights From Penetration-Aware Experiments
abstract
In the assisted driving-based Intelligent Transportation System (ADITS) featuring speed-guided functionality, intelligent vehicles are crucial for enhancing transportation sustainability. However, the impact of ADITS on transportation sustainability at different penetration rates remains unclear. This study combines road testing with simulation to address this gap. By integrating road test data, traffic survey data, ADITS algorithms, and Monte Carlo uncertainty analysis, the study validates the simulation results’ reliability. Correction coefficients are provided to refine energy consumption and carbon emissions simulations. The findings indicate that the differences between simulation data and on-road test data are minimal. However, under varying penetration rate environments, the differences in most variables are statistically significant. At 100% penetration, the adjusted energy use in congested traffic stands at 1.05 MJ/km, with carbon emissions of 81.6 g/km. In uncongested scenarios, these values drop to 0.65 MJ/km and 55.4 g/km, respectively. With the increasing penetration rate, energy efficiency and decarbonization efficiency gradually improve, achieving an optimization of 23%−27%. Intelligent vehicles optimized for uncongested conditions exhibit smoother driving patterns, mitigating aggressive maneuvers and contributing to green transport. In congested scenarios, intelligent vehicles face constraints from leading vehicles, limiting speed optimization. At lower penetration rates (0.25), ADITS can worsen traffic conditions. However, with higher penetration rates, speed fluctuations decrease, leading to more uniform road speeds, reduced high-speed accelerations, and lower energy consumption and carbon emissions. These findings provide theoretical support for the implementation and development of intelligent transportation systems.
Qijun Zhang, Zhenyu Jia, Hongjun Mao
IEEE Trans. Intell. Transp. Syst.2
2025 Towards Big Data in AI for EDA Research: Generation of New Pseudo Circuits at RTL Stage
abstract
Machine learning (ML) techniques have demonstrated remarkable effectiveness in electronic design automation (EDA). ML models need to be trained on diverse circuit datasets for better accuracy and generalization capabilities. However, the availability of circuit data remains a long-standing severe issue. The strong data privacy concern in the semiconductor industry makes direct sharing of circuit IPs almost impossible. To address the data availability problem, open-source datasets like CircuitNet have been proposed, but they mostly focus on collecting labels of several existing open-source designs, instead of generating any new designs. In this work, we make an innovative exploration to directly generate new pseudo-circuits without human effort. We believe that generating pseudo-circuits is the most promising, if not the only, approach to achieving "big data" in the semiconductor industry in the foreseeable future. We demonstrate that pseudo-circuits can significantly boost the performance of ML models in early design quality predictions, as early as the pre-synthesis RTL stage.
Shang Liu 0006, Wenji Fang, Yao Lu 0031, Qijun Zhang, Zhiyao Xie
ASP-DAC4
2025 FirePower: Towards a Foundation with Generalizable Knowledge for Architecture-Level Power Modeling
abstract
Power efficiency is a critical design objective in modern processor design. A high-fidelity architecture-level power modeling method is greatly needed by CPU architects for guiding early optimizations. However, traditional architecture-level power models can not meet the accuracy requirement, largely due to the discrepancy between the power model and actual design implementation. While some machine learning (ML)-based architecture-level power modeling methods have been proposed in recent years, the data-hungry ML model training process requires sufficient similar known designs, which are unrealistic in many development scenarios.
Qijun Zhang, Mengming Li, Yao Lu 0031, Zhiyao Xie
ASP-DAC1
2025 Pointer: An Energy-Efficient ReRAM-based Point Cloud Recognition Accelerator with Inter-layer and Intra-layer Optimizations
abstract
Point cloud is an important data structure for a wide range of applications, including robotics, AR/VR, and autonomous driving. To process the point cloud, many deep-learning-based point cloud recognition algorithms have been proposed. However, to meet the requirement of applications like autonomous driving, the algorithm must be fast enough, rendering accelerators necessary at the inference stage. But existing point cloud accelerators are still inefficient due to two challenges. First, the multi-layer perceptron (MLP) during feature computation is the performance bottleneck. Second, the feature vector fetching operation incurs heavy DRAM access.
Qijun Zhang, Zhiyao Xie
ASP-DAC1
2025 ATLAS: A Self-Supervised and Cross-Stage Netlist Power Model for Fine-Grained Time-Based Layout Power Analysis
abstract
Accurate power prediction in VLSI design is crucial for effective power optimization, especially as designs get transformed from gate-level netlist to layout stages. However, traditional accurate power simulation requires time-consuming back-end processing and simulation steps, which significantly impede design optimization. To address this, we propose ATLAS, which can predict the ultimate time-based layout power for any new design in the gate-level netlist. To the best of our knowledge, ATLAS is the first work that supports both time-based power simulation and general cross-design power modeling. It achieves such general timebased power modeling by proposing a new pre-training and fine-tuning paradigm customized for circuit power. Targeting golden per-cycle layout power from commercial tools, our ATLAS achieves the mean absolute percentage error (MAPE) of only ${0. 5 8 \%, ~} {0. 4 5 \%}$, and ${5. 1 2 \%}$ for the clock tree, register, and combinational power groups, respectively, without any layout information. Overall, the MAPE for the total power of the entire design is $\lt1 \%$, and the inference speed of a workload is significantly faster than the standard flow of commercial tools.
Yao Lu 0031, Wenji Fang, Jing Wang 0171, Qijun Zhang, Zhiyao Xie
DAC5
2025 AutoPower: Automated Few-Shot Architecture-Level Power Modeling by Power Group Decoupling
abstract
Power efficiency is a critical design objective in modern CPU design. Architects need a fast yet accurate architecture-level power evaluation tool to perform early-stage power estimation. However, traditional analytical architecture-level power models are inaccurate. The recently proposed machine learning (ML)-based architecture-level power model requires sufficient data from known configurations for training, making it unrealistic. In this work, we propose AutoPower targeting fully automated architecture-level power modeling with limited known design configurations. We have two key observations: (1) The clock and SRAM dominate the power consumption of the processor, and (2) The clock and SRAM power correlate with structural information available at the architecture level. Based on these two observations, we propose the power group decoupling in AutoPower. First, AutoPower decouples across power groups to build individual power models for each group. Second, AutoPower designs power models by further decoupling the model into multiple sub-models within each power group. In our experiments, AutoPower can achieve a low mean absolute percentage error (MAPE) of $4.36 \%$ and a high $R^{2}$ of 0.96 even with only two known configurations for training. This is $5 \%$ lower in MAPE and 0.09 higher in $R^{2}$ compared with McPAT-Calib, the representative ML-based power model.
Qijun Zhang, Yao Lu 0031, Mengming Li, Zhiyao Xie
DAC1
2025 Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency
abstract
Hardware prefetching plays a critical role in hiding the off-chip DRAM latency. The complexity of applications results in a wide variety of memory access patterns, prompting the development of numerous cache-prefetching algorithms. Consequently, commercial processors often employ a hybrid of these algorithms to enhance the overall prefetching performance. Nonetheless, since these prefetchers share hardware resources, conflicts arising from competing prefetching requests can negate the benefits of hardware prefetching. Under such circumstances, several prefetcher selection algorithms have been proposed to mitigate conflicts between prefetchers. However, these prior solutions suffer from two limitations. First, the input demand request allocation is inaccurate. Second, the prefetcher selection criteria are coarse-grained. In this paper, we address both limitations by introducing an efficient and widely applicable prefetcher selection algorithm—Alecto1, which tailors the demand requests for each prefetcher. Every demand request is first sent to Alecto to identify suitable prefetchers before being routed to prefetchers for training and prefetching. Our analysis shows that Alecto is adept at not only harmonizing prefetching accuracy, coverage, and timeliness but also significantly enhancing the utilization of the prefetcher table, which is vital for temporal prefetching. Alecto outperforms the state-of-the-art RL-based prefetcher selection algorithm—Bandit by $2.76 \%$ in single-core, and $\mathbf{7. 5 6 \%}$ in eight-core. For memory-intensive benchmarks, Alecto outperforms Bandit by $\mathbf{5. 2 5 \%}$. Alecto consistently delivers state-of-the-art performance in scheduling various types of cache prefetchers. In addition to the performance improvement, Alecto can reduce the energy consumption associated with accessing the prefetchers’ table by $48 \%$ ($7 \%$ energy reduction on the entire memory hierarchy), while only adding less than 1 KB of storage overhead.1The name Alecto stands for the combination of selection and allocation.
Mengming Li, Qijun Zhang, Yongqing Ren, Zhiyao Xie
HPCA2
2025 Profile-Guided Temporal Prefetching
abstract
Temporal prefetching shows promise for handling irregular memory access patterns, which are common in data-dependent and pointer-based data structures.Recent studies introduced on-chip metadata storage to reduce the memory traffic caused by accessing metadata from off-chip DRAM.However, existing prefetching schemes struggle to efficiently utilize the limited on-chip storage.An alternative solution, software indirect access prefetching, remains ineffective for optimizing temporal prefetching.In this work, we propose Prophet-a hardware-software codesigned framework that leverages profile-guided methods to optimize metadata storage management.Prophet profiles programs using counters instead of traces, injects hints into programs to guide metadata storage management, and dynamically tunes these hints to enable the optimized binary to adapt to different program inputs.Prophet is designed to coexist with existing hardware temporal prefetchers, delivering efficient, high-performance solutions for frequently executed workloads while preserving the original runtime scheme for less frequently executed workloads.Prophet outperforms the state-of-the-art temporal prefetcher, Triangel, by 14.23%, effectively addressing complex temporal patterns where prior profile-guided solutions fall short (only achieving 0.1% performance gain).Prophet delivers superior performance across all evaluated workload inputs, introducing negligible profiling, analysis, and instruction overhead.
Mengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang, Yao Lu 0031, Yongqing Ren, Zhiyao Xie
ISCA2
2025 ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
abstract
Power is the primary design objective of large-scale integrated circuits (ICs), especially for complex modern processors (i.e., CPUs). Accurate CPU power evaluation requires designers to go through the whole time-consuming IC implementation process, easily taking months. At the early design stage (e.g., architecture-level), classical power models are notoriously inaccurate. Recently, ML-based architecture-level power models have been proposed to boost accuracy, but the data availability is a severe challenge. Currently, there is no open-source dataset for this important ML application. A typical dataset generation process involves correct CPU design implementation and repetitive execution of power simulation flows, requiring significant design expertise, engineering effort, and execution time. Even private in-house datasets often fail to reflect realistic CPU design scenarios. In this work, we propose ArchPower, the first open-source dataset for architecture-level processor power modeling. We go through complex and realistic design flows to collect the CPU architectural information as features and the ground-truth simulated power as labels. Our dataset includes 200 CPU data samples, collected from 25 different CPU configurations when executing 8 different workloads. There are more than 100 architectural features in each data sample, including both hardware and event parameters. The label of each sample provides fine-grained power information, including the total design power and the power for each of the 11 components. Each power value is further decomposed into four fine-grained power groups: combinational logic power, sequential logic power, memory power, and clock power. ArchPower is available at https://github.com/hkust-zhiyao/ArchPower.
Qijun Zhang, Yao Lu 0031, Mengming Li, Shang Liu 0006, Zhiyao Xie
NeurIPS1
2025 Transferable Presynthesis PPA Estimation for RTL Designs With Data Augmentation Techniques
abstract
In modern VLSI design flow, evaluating the quality of register-transfer level (RTL) designs involves time-consuming logic synthesis using electronic design automation tools, a process that often slows down early optimization. While recent machine learning (ML) solutions offer some advancements, they typically struggle with maintaining high accuracy across any given RTL design. In this work, we propose an innovative transferable presynthesis power, performance, and area (PPA) estimation framework named MasterRTL. It first converts the hardware description language code to a new bit-level design representation named the simple operator graph (SOG). By only adopting single-bit simple operators, this SOG proves to be a general representation that unifies different design types and styles. The SOG is also more similar to the target gate-level netlist, reducing the gap between the RTL representation and netlist. In addition to the new SOG representation, MasterRTL proposes new ML methods for the RTL-stage modeling of timing, power, and area separately. Compared with the state-of-the-art solutions, the experiment on a comprehensive dataset with 90 different designs shows accuracy improvement by 0.33, 0.22, and 0.15 in correlation for total negative slack (TNS), worst negative slack (WNS), and power, respectively. Besides the prediction of the synthesis results, MasterRTL also excels in accurately predicting layout-stage PPA based on the RTL designs and in adapting across different technology nodes and process corners. Furthermore, we investigate two effective data augmentation techniques: 1) a graph generation method and 2) a large language model (LLM)-based approach. Our results validate the effectiveness of the generated RTL designs in mitigating the data shortage challenges.
Wenji Fang, Yao Lu 0031, Shang Liu 0006, Qijun Zhang, Ceyu Xu, Lisa Wu Wills, Hongce Zhang, Zhiyao Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 RTLCoder: Fully Open-Source and Efficient LLM-Assisted RTL Code Generation Technique
abstract
The automatic generation of RTL code (e.g., Verilog) using natural language instructions and large language models (LLMs) has attracted significant research interest recently. However, most existing approaches heavily rely on commercial LLMs, such as ChatGPT, while open-source LLMs tailored for this specific design generation task exhibit notably inferior performance. The absence of high-quality open-source solutions restricts the flexibility and data privacy of this emerging technique. In this study, we present a new customized LLM solution with a modest parameter count of only 7B, achieving better performance than GPT-3.5 on all representative benchmarks for RTL code generation. Especially, it outperforms GPT-4 in VerilogEval Machine benchmark. This remarkable balance between accuracy and efficiency is made possible by leveraging our new RTL code dataset and a customized LLM algorithm, both of which have been made fully open-source. Furthermore, we have successfully quantized our LLM to 4-bit with a total size of 4 GB, enabling it to function on a single laptop with only slight performance degradation. This efficiency allows the RTL generator to serve as a local assistant for engineers, ensuring all design privacy concerns are addressed.
Shang Liu 0006, Wenji Fang, Yao Lu 0031, Jing Wang 0171, Qijun Zhang, Hongce Zhang, Zhiyao Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 An Architecture-Level CPU Modeling Framework for Power and Other Design Qualities
abstract
Power efficiency is a critical design objective in modern microprocessor design. To evaluate the impact of architectural-level design decisions, an accurate yet efficient architecture-level power model is desired. However, widely adopted analytical power models like McPAT and Wattch have been criticized for their unreliable accuracy, while machine learning (ML) methods like McPAT-Calib rely on sufficient known designs for training and perform poorly when available designs are limited, which is the case in realistic scenarios. In this work, we propose PANDA, an innovative architecture-level solution that combines the advantages of analytical and ML power models. It achieves unprecedented high accuracy on unknown new designs even when there are very limited designs for training. Besides being an excellent average power model, we also extend PANDA to support the time-based power trace prediction, which can enable the analysis of peak power, power fluctuations, and voltage fluctuation. This is highly challenging at the architecture level. Other qualities, such as area, performance, and energy accurately, can also be supported. In addition to single design quality, PANDA can model the tradeoffs among different design qualities, such as the tradeoff between power and timing, by predicting the Pareto-optimal curve. Finally, PANDA can further support power prediction for unknown new technology nodes. Our experiment shows that, for average power prediction, our method can achieve high accuracy with a correlation coefficient R of 0.99 and mean absolute percentage error (MAPE) of 7.91% even when only one configuration is known, outperforming McPAT-Calib which has R of -0.24 and MAPE of 35.96%. For time-based power trace prediction, our method can achieve a low MAPE of 4.34%, outperforming the state-of-the-art method Powertrain which has a MAPE of 53.8%.
Qijun Zhang, Mengming Li, Andrea Mondelli, Zhiyao Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model
abstract
Inspired by the recent success of large language models (LLMs) like ChatGPT, researchers start to explore the adoption of LLMs for agile hardware design, such as generating design RTL based on natural-language instructions. However, in existing works, their target designs are all relatively simple and in a small scale, and proposed by the authors themselves, making a fair comparison among different LLM solutions challenging. In addition, many prior works only focus on the design correctness, without evaluating the design qualities of generated design RTL. In this work, we propose an open-source benchmark named RTLLM, for generating design RTL with natural language instructions. To systematically evaluate the auto-generated design RTL, we summarized three progressive goals, named syntax goal, functionality goal, and design quality goal. This benchmark can automatically provide a quantitative evaluation of any given LLM-based solution. Furthermore, we propose an easy-to-use yet surprisingly effective prompt engineering technique named self-planning, which proves to significantly boost the performance of GPT-3.5 in our proposed benchmark.
Yao Lu 0031, Shang Liu 0006, Qijun Zhang, Zhiyao Xie
ASPDAC3
2024 Unleashing Flexibility of ML-based Power Estimators Through Efficient Development Strategies
abstract
Power is a primary design objective in modern VLSI design. Efficient and accurate power evaluation tools are in high demand to provide prompt power feedback for early design optimization. However, it is time-consuming to simulate long, fine-grained (e.g., per-cycle) power traces in complex designs with commercial power simulators. In recent years, machine learning (ML)-based power models have emerged as a potential solution to make fast predictions on per-cycle power based on signal toggling activities. Despite their immense potential, these models are currently underutilized in realistic development scenarios, largely due to the challenges associated with their development and updates. To overcome the barriers, this work optimizes the often-neglected power model development process by an in-depth examination of power data's impact on model accuracy. We propose efficient strategies to minimize the overhead involved in model development. Furthermore, we enhance model flexibility by enabling the transfer of existing models to updated design RTLs with negligible additional costs.
Yao Lu 0031, Qijun Zhang, Zhiyao Xie
ISLPED2
2023 MasterRTL: A Pre-Synthesis PPA Estimation Framework for Any RTL Design
abstract
In modern VLSI design flow, the register-transfer level (RTL) stage is a critical point, where designers define precise design behavior with hardware description languages (HDLs) like Verilog. Since the RTL design is in the format of HDL code, the standard way to evaluate its quality requires time-consuming subsequent synthesis steps with EDA tools. This time-consuming process significantly impedes design optimization at the early RTL stage. Despite the emergence of some recent ML-based solutions, they fail to maintain high accuracy for any given RTL design. In this work, we propose an innovative pre-synthesis PPA estimation framework named MasterRTL. It first converts the HDL code to a new bit-level design representation named the simple operator graph (SOG). By only adopting single-bit simple operators, this SOG proves to be a general representation that unifies different design types and styles. The SOG is also more similar to the target gate-level netlist, reducing the gap between RTL representation and netlist. In addition to the new SOG representation, MasterRTL proposes new ML methods for the RTL-stage modeling of timing, power, and area separately. Compared with state-of-the-art solutions, the experiment on a comprehensive dataset with 90 different designs shows accuracy improvement by 0.33, 0.22, and 0.15 in correlation for total negative slack (TNS), worst negative slack (WNS), and power, respectively.
Wenji Fang, Yao Lu 0031, Shang Liu 0006, Qijun Zhang, Ceyu Xu, Lisa Wu Wills, Hongce Zhang, Zhiyao Xie
ICCAD4
2023 PANDA: Architecture-Level Power Evaluation by Unifying Analytical and Machine Learning Solutions
abstract
Power efficiency is a critical design objective in modern microprocessor design. To evaluate the impact of architectural-level design decisions, an accurate yet efficient architecture-level power model is desired. However, widely adopted data-independent analytical power models like McPAT and Wattch have been criticized for their unreliable accuracy. While some machine learning (ML) methods have been proposed for architecture-level power modeling, they rely on sufficient known designs for training and perform poorly when the number of available designs is limited, which is typically the case in realistic scenarios. In this work, we derive a general formulation that unifies existing architecture-level power models. Based on the formulation, we propose PANDA, an innovative architecture-level solution that combines the advantages of analytical and ML power models. It achieves unprecedented high accuracy on unknown new designs even when there are very limited designs for training, which is a common challenge in practice. Besides being an excellent power model, it can predict area, performance, and energy accurately. PANDA further supports power prediction for unknown new technology nodes. In our experiments, besides validating the superior performance and the wide range of functionalities of PANDA, we also propose an application scenario, where PANDA proves to identify high-performance design configurations given a power constraint.
Qijun Zhang, Shiyu Li 0001, Guanglei Zhou, Jingyu Pan, Chen-Chia Chang, Yiran Chen 0001, Zhiyao Xie
ICCAD1
2018 Fast Neural Network Training on FPGA Using Quasi-Newton Optimization Method
Qiang Liu 0011, Ruoyu Sang, Tao Zhang 0025, Qijun Zhang
IEEE Trans. Very Large Scale Integr. Syst.6
2017 Fast and Automated Electromigration Analysis for CMOS RF PA Design
Junjie Gu, Haipeng Fu, Weicong Na, Qijun Zhang, Jianguo Ma
J. Electron. Test.4
2016 Knowledge-Based Neural Network Model for FPGA Logical Architecture Development
abstract
This paper proposes a knowledge-based neural network (KBNN) modeling approach for field-programmable gate array (FPGA) logical architecture design. The KBNN embeds the existing FPGA analytical models (AMs) into an NN. The NN can complement the AMs according to their needs to provide further increased model accuracy, while maintaining the meaningful trends successfully captured in the AMs. The obtained KBNN predicts the routing channel width required by circuit implementations on various FPGA architectures, which can be used by architects to quickly and accurately evaluate various FPGA architectures in early development stages. Experimental results show that the KBNN-based approach achieves an average error of 2%, which shows 75% accuracy enhancement over the existing AMs for routing channel width estimation of a set of benchmark circuits and FPGA architectures. The KBNN model has been applied to three FPGA architecture development scenarios to demonstrate its practical application and effectiveness.
Qiang Liu 0011, Qijun Zhang
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Accuracy Improvement of Energy Prediction for Solar-Energy-Powered Embedded Systems
abstract
Solar energy prediction is a key to the power management in the electronic embedded system that operates using the harvested solar energy. This paper proposes accuracy improvement approaches for the solar energy prediction based on artificial neural networks, in order to increase the robustness of solar-energy-powered systems. Two complementary neural network models, multilayer perceptron (MLP) network and knowledge-based neural network (KBNN), are exploited to predict the future solar energy, through offline and online training. MLP is constructed under the guidance of the proposed input parameter selection approach and is used when the training data are sufficient. KBNN is employed to take advantage of the existing prediction models and is especially valuable when the training data are insufficient. Built on top of the existing prediction approaches, our work results in a synergy that can overcome the accuracy limitation of the existing prediction approaches. The experimental results show the prediction accuracy improvements by up to 65.4%, compared with the existing approaches. The results also demonstrate the capability of KBNN in providing a reliable model, especially when fewer training data are available.
Qiang Liu 0011, Qijun Zhang
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Neural network based pre-placement wirelength estimation
abstract
We propose a neural network based approach for estimating the total wirelength of a digital circuit, mapped onto an FPGA, before circuit placement and routing. A 3-layer MLP neural network is trained to learn the behavior of a placement tool and then quickly predicts the wirelength of a circuit design with the accuracy similar to one obtained after placement. A priori knowledge about the wirelength of circuit designs can be used to effectively guide the design exploration processes at the early design stages. This breaks the repetitive CAD design flow and reduces the design cycle. In this work, five circuit parameters and two FPGA architecture parameters are considered in the wirelength estimation. The proposed approach is evaluated by comparing the wirelength given by the trained neural networks and the placement tool VPR for the IWLS2005 circuit benchmark. Results show that the neural network's estimation has an average error below 0.6% compared to VPR. The neural network model is also compared to a linear model for the wirelength estimation, showing 7.39 times improvement in the estimation accuracy.
Qiang Liu 0011, Jianguo Ma, Qijun Zhang
FPT3
2012 An enhanced Neuro-Space mapping method for nonlinear microwave device modeling
abstract
In this article, a new Neuro-Space mapping method is presented aimed at using neural networks to automatically enhance nonlinear device models, such as FET models. Compared with previously published space mapping methods, our proposed method produces better modeling accuracy and provides more effective combinations of mapping structure with existing coarse model. In our proposed models, separate mappings for voltage and current at gate and drain are used as the mapping structure. Training methods for mapping neural networks are also proposed. Application examples on modeling MESFET devices and the use of new models in DC, S-parameter and combined DC and S-parameter simulation demonstrate that our proposed Neuro-Space mapping model matches more closely with the device data than that by the traditional Neuro-Space mapping method for modeling nonlinear microwave devices.
Lin Zhu 0005, Yongtao Ma, Qijun Zhang
ISCAS3
2006 Evolving musical performance profiles using genetic algorithms with structural fitness
abstract
This paper presents a system that uses Genetic Algorithm (GA) to evolve hierarchical pulse sets (i.e., hierarchical duration vs. amplitude matrices) for expressive music performance by machines. The performance profile for a piece of music is represented using pulse sets and the fitness (for the GA) is derived from the structure of the piece to be performed; hence the term "structural fitness". Randomly initiated pulse sets are selected and evolved using GA. The fitness value is calculated by measuring the pulse set's ability of highlighting musical structures. This measurement is based upon generative rules for expressive music performance. This is the first stage of a project, which is aimed at the design of a dynamic model for the evolution of expressive performance profiles by interacting agents in an artificial society of musicians and listeners.
Qijun Zhang, Eduardo Reck Miranda
GECCO1