Yuntao Lu

dblp:205/8761 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 LLM-Assisted Circuit Verification: A Comprehensive Survey
Hongduo Liu, Yuntao Lu, Xufeng Yao, Bei Yu 0001
ASP-DAC2
2026 Comfort temperature assessment for honeybee colonies based on long-term monitoring
Yuntao Lu, Shijuan Li, Shengping Liu
Expert Syst. Appl.1
2026 DeepVerifier: Learning to Update Test Sequences for Coverage-Guided Verification
abstract
Verification is critical in ensuring the reliable operation of modern, complex computing systems. However, as processor designs become increasingly sophisticated, conventional static verification techniques struggle to generate high-quality test sequences that achieve comprehensive coverage. Dynamic simulation-based approaches, which leverage coverage-driven objectives, can increase confidence in correct processor functionality but often suffer from low verification efficiency due to the generation of redundant test sequences and significant computational overhead. To address these challenges, this paper presents DeepVerifier, a novel coverage-guided test generation framework that leverages data-driven learning of existing test sequences and their associated coverage feedback. DeepVerifier uses a language model to learn the semantic representations of test sequences, ensure adherence to syntax constraints, and estimate the relationship between test sequences and coverage scores. By updating test sequences with higher coverage, DeepVerifier can significantly improve the efficiency and effectiveness of the verification process. Experimental results of verifying an out-of-order RISC-V microprocessor demonstrate that the framework accurately estimates the coverage scores of test sequences and updates high-quality sequences that contribute to higher coverage. This coverage-guided test generation technique holds promise for enhancing the reliability of modern processor designs.
Yuntao Lu, Yuxuan Zhao 0001, Ziyue Zheng, Yangdi Lyu, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.1
2025 DiffCCD: Differentiable Concurrent Clock and Data Optimization
abstract
Timing optimization following clock tree synthesis (post-CTS) is a crucial step in very large scale integration (VLSI) physical design for achieving timing closure. During this stage, clock skew significantly impacts circuit timing performance, making useful skew optimization essential for enhancing design quality. However, traditional skew optimization methods face challenges due to their insufficient consideration of physical implementation constraints. To overcome these limitations, we propose a GPU-accelerated differentiable concurrent clock and data (CCD) optimization framework, which simultaneously optimizes clock skew and logic delays to enhance overall timing performance with the consideration of physical constraints. We implement the CCD optimization method as a step involving buffer sizing in the clock network and refining placement results. The key innovation of our approach lies in formulating a smooth and differentiable process for CCD optimization with a calibration mechanism to ensure accurate gradient computations. Additionally, we employ an alternating direction method of multipliers (ADMM)-based strategy to decompose the entire optimization problem into several manageable subproblems, effectively balancing timing optimization with physical implementation constraints. Experimental results on open-source industrial benchmarks demonstrate that our CCD optimization framework achieves superior timing closure compared to a baseline approach within an open-source physical design tool. Our method yields an average improvement of 22.4% in worst negative slack (WNS) and 45.0% in total negative slack (TNS), along with a 9.434× runtime speedup. To our knowledge, this is the first work to incorporate clock skew effects into gradient-based timing optimization.
Yuhao Ji, Yuntao Lu, Zuodong Zhang, Zizheng Guo 0001, Yibo Lin, Bei Yu 0001
ICCAD2
2025 VIRTUAL: Vector-based Dynamic Power Estimation via Decoupled Multi-Modality Learning
abstract
Dynamic power analysis in digital integrated circuits (ICs) conventionally relies on gate-level synthesis and simulation, creating a critical bottleneck in iterative design flows. We propose VIRTUAL, a multi-modality learning framework for rapid post-synthesis dynamic power estimation directly from Register-Transfer Level (RTL) implementations and input waveform vectors, eliminating the need for gate-level synthesis and extensive simulation. By decoupling features from input port waveforms and RTL implementations, VIRTUAL employs a transformer-based encoder to extract temporal patterns from input port waveforms and a graph neural network (GNN) to capture structural and functional dependencies within RTL implementations. Through self-supervised contrastive learning across sequential and graph modalities, the framework learns robust power-relevant representations with minimal labeled data. Subsequently, VIRTUAL refines the multi-modality embeddings using a lightweight fusion module and a power prediction head, enabling dynamic power estimation for fixed clock periods within seconds or minutes. Experimental evaluations demonstrate approximately a Pearson correlation coefficient (PCC) of 0.842 and a mean absolute percentage error (MAPE) of 23.43%, while achieving 14.27×speedup compared to traditional gate-level power analysis workflows. Experimental results across diverse RTL designs and input port waveforms validate that our proposed learning-based approach maintains the accuracy while significantly reducing design iteration time, transforming hours of synthesis and simulation into minutes of direct prediction.
Yuntao Lu, Yihan Wen, Jianan Mu, Huawei Li 0001, Bei Yu 0001
ICCAD1
2017 A high-performance FPGA accelerator for sparse neural networks: work-in-progress
abstract
Neural networks have been widely used in a large range of domains, researchers tune numbers of layrs, neurons and synapses to adapt various applications. As a consequence, computations and memory of neural networks models are both intensive. As large requirements of memory and computing resources, it is difficult to deploy neural networks on resource-limited platforms. Sparse neural networks, which prune redundant neurons and synapses, alleviate computation and memory pressure. However, conventional accelerators cannot benefit from the sparse feature.
Yuntao Lu, Lei Gong 0003, Chongchong Xu, Yiwei Zhang 0001, Chao Wang 0003, Xuehai Zhou
CASES1
2017 A Power-Efficient Accelerator for Convolutional Neural Networks
abstract
Convolutional neural networks(CNNs) have been widely applied in various applications. However, the computation-intensive convolutional layers and memory-intensive fully connected layers have brought many challenges to the implementation of CNN on embedded platforms. To overcome this problem, this work proposes a power-efficient accelerator for CNNs, and different methods are applied to optimize the convolutional layers and fully connected layers. For the convolutional layer, the accelerator first rearranges the input features into matrix on-the-fly when storing them to the on-chip buffers. Thus the computation of convolutional layer can be completed through matrix multiplication. For the fully connected layer, the batch-based method is used to reduce the required memory bandwidth, which also can be completed through matrix multiplication. Then a two-layer pipelined computation method for matrix multiplication is proposed to increase the throughput. As a case study, we implement a widely used CNN model, LeNet-5, on an embedded device. It can achieve a peak performance of 34.48 GOP/s and the power efficiency with the value of 19.45 GOP/s/W under 100MHz clock frequency which outperforms previous approaches.
Chao Wang 0003, Lei Gong 0003, Chongchong Xu, Yiwei Zhang 0001, Yuntao Lu, Xi Li 0003, Xuehai Zhou
CLUSTER6
2017 OmniGraph: A Scalable Hardware Accelerator for Graph Processing
abstract
Large-scale graphs processing attracts more and more attentions, and it has been widely applied in many application domains. FPGA is a promising platform to implement graph processing algorithms with high power-efficiency and parallelism. In this paper, we propose OmniGraph, a scalable hardware accelerator for graph processing. OmniGraph can process graphs with different sizes adaptively and is adaptable to various graph algorithms. OmniGraph improves the preprocessing methodology based on Interval-Shard and consists of three computation engines, vertices on-chip && edges on-chip engine, vertices on-chip && edges off-chip engine, and vertices off-chip && edges off-chip engine. Experimental results on the state-of-the-art Xilinx Virtex-7 board demonstrate that case studies in OmniGraph achieve 1.03x-8.13x average speedup comparing to GraphChi on Intel core2 processors.
Chongchong Xu, Chao Wang 0003, Lei Gong 0003, Yuntao Lu, Yiwei Zhang 0001, Xi Li 0003, Xuehai Zhou
CLUSTER4
2017 A Power-Efficient Accelerator Based on FPGAs for LSTM Network
abstract
Today, artificial neural networks (ANNs) are widely used in a variety of applications, including speech recognition, face detection, disease diagnosis, etc. And as the emerging field of ANNs, Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) which contains complex computational logic. To achieve high accuracy, researchers always build large-scale LSTM networks which are time-consuming and power-consuming. In this paper, we present a hardware accelerator for the LSTM neural network layer based on FPGA Zedboard and use pipeline methods to parallelize the forward computing process. We also implement a sparse LSTM hidden layer, which consumes fewer storage resources than the dense network. Our accelerator is power-efficient and has a higher speed than ARM Cortex-A9 processor.
Yiwei Zhang 0001, Chao Wang 0003, Lei Gong 0003, Yuntao Lu, Chongchong Xu, Xi Li 0003, Xuehai Zhou
CLUSTER4
2017 Evaluation and Trade-offs of Graph Processing for Cloud Services
abstract
Large-scale data is often represented as graphs in the field of modern cloud computing. Graph processing attracts more and more attentions when utilizing the cloud computing service. With the increasing attentions to process massive graphs (e.g., social networks, web graphs, transport networks, and bioinformatics), many state-of-the-art open source graph computing systems on a single node have been proposed, including GraphChi, X-Stream, and GridGraph. GraphChi adopts a vertex-centric model while the latter two adopt an edge-centric model. However, there is a lack of evaluations and analyses to the performance of these systems, which makes it difficult for users to choose the best system for their applications. In this paper, to make the graph processing provide excellent cloud services to users, we propose an evaluation framework, conduct a series of extensive experiments to evaluate the performance and analyze the bottlenecks of these systems on graphs with different characteristics and different kinds of algorithms. The metrics we adopt in this paper are principles to design graph computing systems on a single node, such as RunTime, CPU Utilization, and Data Locality. The results demonstrate the trade-offs among different graph frameworks and X-Stream is more suitable to process transport networks on WCC and BFS, compared to GridGraph. Besides, we present several discussions on GridGraph. The results of our work are concluded as a reference for users, researchers, and developers.
Chongchong Xu, Jinhong Zhou, Yuntao Lu, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou
ICWS3