Deyuan Guo

dblp:78/10161 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0004-1756-1784ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Characterizing Digital DRAM PIM through Modeling and Benchmarking
abstract
The disparity between processor speed and memory bandwidth has become a growing performance bottleneck, particularly for memory-intensive workloads. Processing-in-Memory (PIM) mitigates this bottleneck by integrating computation directly within DRAM. However, the effectiveness of PIM varies significantly across workloads, architectures, and DRAM technologies, yet it is often assessed using tightly coupled simulators and benchmarks that lack portability and generality. This article extends PIMbench and PIMeval —a generalizable benchmark suite and an extensible PIM simulator to support a broader range of workloads, PIM architectures, and DRAM technologies. This evaluation incorporates roofline analysis and a breakdown of intra-memory execution stages to identify PIM-specific bottlenecks and performance scaling limits. The evaluation spans three classes of digital PIM architectures: subarray-level bit-serial, subarray-level bit-parallel, and bank-level bit-parallel. It further demonstrates how internal DRAM parameters such as subarray count and GDL width impact PIM performance. The code is publicly available at: https://github.com/UVA-LavaLab/PIMeval-PIMbench .
Farzana Siddique, Deyuan Guo, Hugo Abbot, Kyle Durrer, MohammadHosein Gholamrezaei, Morteza Baradaran, Ethan Ermovick, Alif Ahmed, Zhenxing Fan, Beenish Gul, Ashish Venkat, Kevin Skadron
ACM Trans. Archit. Code Optim.2
2025 InsightAlign: A Transferable Physical Design Recipe Recommender Based on Design Insights
abstract
Physical design tools have complex workflows with many different ways of optimizing power, performance, and area (PPA) out of a large number of options and hyperparameters in different engines and functionalities. Black-box optimization techniques are widely adopted to automate quality-of-result (QoR) exploration. Such exploration often proves impractical in real-world customer environments due to high computational demands, lengthy exploration cycles, and the need for large parallel jobs. To reduce the exploration space for viable compute resource requirements, we propose a novel design methodology to enable transferable learning by incorporating design insights crafted on top of physical design experts’ experience and streamlining QoR exploration as a sequence generation task for best recipe selection. We apply language model-inspired alignment techniques to learn the ranking of different recipe sets, enabling our model to generalize beyond known-good manually tuned expert design recipes. Extensive evaluations demonstrate our method’s superior QoRs and runtime performance on unseen industrial designs and rigorous benchmarks.
Hao-Hsiang Hsiao, Sudipto Kundu, Wei-Ting Chan, Deyuan Guo, Sung Kyu Lim
DAC5
2023 RL-CCD: Concurrent Clock and Data Optimization using Attention-Based Self-Supervised Reinforcement Learning
abstract
Concurrent Clock and Data (CCD) optimization is a well-adopted approach in modern commercial tools that resolves timing violations using a mixture of clock skewing and delay fixing strategies. However, existing CCD algorithms are flawed. Particularly, they fail to prioritize violating endpoints for different optimization strategies correctly, leading to flow-wise globally sub-optimal results. In this paper, we overcome this issue by presenting RL-CCD, a Reinforcement Learning (RL) agent that selects endpoints for useful skew prioritization using the proposed EP-GNN, an endpoint-oriented Graph Neural Network (GNN) model, and a Transformer-based self-supervised attention mechanism. Experimental results on 19 industrial designs in 5 − 12nm technologies demonstrate that RL-CCD achieves up to 64% Total Negative Slack (TNS) reduction and 66.5% number of violating endpoints (NVE) improvement over the native implementation of a commercial tool.
Yi-Chen Lu, Wei-Ting Chan, Deyuan Guo, Sudipto Kundu, Vishal Khandelwal, Sung Kyu Lim
DAC3
2018 A Scalable Solution for Rule-Based Part-of-Speech Tagging on Novel Hardware Accelerators
abstract
Part-of-speech (POS) tagging is the foundation of many natural language processing applications. Rule-based POS tagging is a wellknown solution, which assigns tags to the words using a set of predefined rules. Many researchers favor statistical-based approaches over rule-based methods for better empirical accuracy. However, until now, the computational cost of rule-based POS tagging has made it difficult to study whether more complex rules or larger rulesets could lead to accuracy competitive with statistical approaches. In this paper, we leverage two hardware accelerators, the Automata Processor (AP) and Field Programmable Gate Arrays (FPGA), to accelerate rule-based POS tagging by converting rules to regular expressions and exploiting the highly-parallel regular-expressionmatching ability of these accelerators. We study the relationship between rule set size and accuracy, and observe that adding more rules only poses minimal overhead on the AP and FPGA. This allows a substantial increase in the number and complexity of rules, leading to accuracy improvement. Our experiments on Treebank and Brown corpora achieve up to 2,600X and 1,914X speedups on the AP and on the FPGA respectively over rule-based methods on the CPU in the rule-matching stage, up to 58× speedup over the Perceptron POS tagger on the CPU in total testing time, and up to 253× speedup over the LSTM tagger on the GPU in total testing time, while showing a competitive accuracy compared to neural-network and statistical solutions.
Elaheh Sadredini, Deyuan Guo, Chunkun Bo, Reza Rahimi, Kevin Skadron, Hongning Wang
KDD2
2014 An Implementation of Message-Passing Interface over VxWorks for Real-Time Embedded Multi-Core Systems
abstract
Message-passing interface (MPI) has proved to be very successful in the high performance computing domain. However, suitability of MPI for embedded real-time system design is still under investigation. In this work, we have provided our methods and experiences of implementing MPI parallel environment for a real-time embedded multi-core system. Our main contributions were to: (1) enable hyper transport bus communication mechanism to establish MPI parallel environment; (2) support VxWorks operating system for establishing MPI parallel environment and (3) enhance the real-time property of MPI mechanism. The digital signal processor (DSP)-MPI presented in this work can also be used on other platforms supporting VxWorks operating system. The results indicate that the real-time property of DSP-MPI has been improved significantly compared with MPICH2. The test on realistic applications also shows that DSP-MPI can fulfill the requirement of our target multi-core platform.
Xu Yang 0003, Deyuan Guo, Hu He 0001, Haijing Tang
Comput. J.2
2011 An Efficient Shared Memory Based Virtual Communication System for Embedded SMP Cluster
abstract
With the prevalence of multi-core processors, it is a trend that the embedded cluster deploys SMP nodes to gain more computing power. As a crucial issue, the MPI inter-process communication has been suffering the contradiction between high performance and embedded constraints. Moreover, there is a big performance gap between intra- and inter-node communication for different infrastructures. In this paper, we design a virtual communication system called SMVN, which extends the shared memory mechanism typically used in intra-node case into the inter-node case. The SMVN utilizes the HT inter-chip interconnect interface in Godson-3A SMP nodes to build a mesh topology. It is Ethernet compatible by simulating bottom layers of TCP/IP protocol. With the design, the node interconnection can get rid of NICs, cables and switches. Furthermore, we exploit the zero-copy scheme and other optimizations to improve the performance. We port the MPICH2 library by socket channel and formulate its process allocation. The MPI latency and bandwidth tests show that the performance difference between two levels is small. The inter-node bandwidth is 27.3 MB/s, which is more than twice the theoretical peak value of 100 Mb Ethernet and reaches 84% of the intra-node performance.
Wenxuan Yin, Deyuan Guo
NAS4