EDBT 2026 Demo / reviewers in the wild / expert
Jinyang Guo 0001
dblp:224/6845-1
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-4405-0424ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingabstractGenerative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by$4.7-8.8 \%$while maintaining high-performance AU applications by reducing SLO violations by$\mathbf{7 - 1 1 \%}$compared with state-of-the-art resource managers. Xinkai Wang 0003, Chao Li 0009, Yiming Zhuansun, Jinyang Guo 0001, Xiaofeng Hou, Jing Wang 0055, Weigao Chen, Liping Zhang 0013, Minyi Guo |
HPCA | 4 |
| 2025 | TriCooling-Sim: Efficient Thermal Simulation for High-Density Micro AI Data Centers
Jinyang Guo 0001, Xinkai Wang 0003, Jing Wang 0055, Xiaofeng Hou, Chao Li 0009, Minyi Guo |
NPC (2) | 1 |
| 2025 | Improving Energy Efficiency of Graph Processing on Shared-Memory SystemsabstractWith the number of cores increasing in shared memory systems, the energy consumption of parallel computing on them is becoming increasingly prominent. Currently, researchers concern with the performance optimization, while ignoring the energy efficiency of graph processing. Meanwhile, existing works that optimize energy efficiency involve mainly the general benchmarks by using dynamic voltage and frequency scaling and thread throttling methods. However, these methods cannot be directly transplanted to graph processing, because most graph algorithms converge in fewer iterations and traditional energy efficiency optimization methods are not applicable to them and will produce much overhead, resulting in the fact that the loss outweighs the gain. And some energy-saving methods estimate the subsequent CPU frequency based on the run-time system state, which leads to an inaccurate prediction of the optimal energy-saving CPU frequency. In view of the above issues, we propose a pre-allocated thread throttling method and a static frequency scaling method. The former achieves thread throttling by establishing a pre-allocated scheduling method, which calculates the optimal energy-saving number of threads promptly when the graph is loaded; On this basis, in order to reduce the cost of dynamic frequency setting at runtime and improve the energy efficiency further, the latter introduces the static frequency scaling method to reduce the execution speed of some tasks by relaxing thread execution time. The experimental results show that the pre-allocated thread throttling method improves the energy efficiency by about 10% compared to the original framework, and the static frequency scaling method further improves it by about 20% with trivial performance loss. Le Luo 0002, Chao Li 0009, Jinyang Guo 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2023 | FPGA sharing in the cloud: a comprehensive analysis
Jinyang Guo 0001, Lu Zhang 0049, José Romero Hung, Chao Li 0009, Jieru Zhao, Minyi Guo |
Frontiers Comput. Sci. | 1 |
| 2023 | DRAGON: Dynamic Recurrent Accelerator for Graph Online ConvolutionabstractDespite the extraordinary applicative potentiality that dynamic graph inference may entail, its practical-physical implementation has been a topic seldom explored in literature. Although graph inference through neural networks has received plenty of algorithmic innovation, its transfer to the physical world has not found similar development. This is understandable since the most preeminent Euclidean acceleration techniques from CNN have little implication in the non-Euclidean nature of relational graphs. Instead of coping with the challenges arising from forcing naturally sparse structures into more inflexible stochastic arrangements, in DRAGON, we embrace this characteristic in order to promote acceleration. Inspired by high-performance computing approaches like Parallel Multi-moth Flame Optimization for Link Prediction (PMFO-LP), we propose and implement a novel efficient architecture, capable of producing similar speed-up and performance than baseline but at a fraction of its hardware requirements and power consumption. We leverage the hidden parallelistic capacity of our previously developed static graph convolutional processor ACE-GCN and expanded it with RNN structures, allowing the deployment of a multi-processing network referenced around a common pool of proximity-based centroids. Experimental results demonstrate outstanding acceleration. In comparison with the fastest CPU-based software implementation available in the literature, DRAGON has achieved roughly 191× speed-up. Under the largest configuration and dataset, DRAGON was also able to overtake a more power-hungry PMFO-LP by almost 1.59× in speed, and at around 89.59% in power efficiency. More importantly than raw acceleration, we demonstrate the unique functional qualities of our approach as a flexible and fault-tolerant solution that makes it an interesting alternative for an anthology of applicative scenarios. José Romero Hung, Chao Li 0009, Taolei Wang, Jinyang Guo 0001, Pengyu Wang 0003, Chuanming Shao, Jing Wang 0055, Guoyong Shi, Xiangwen Liu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2022 | Oversubscribing GPU Unified Virtual Memory: Implications and SuggestionsabstractRecent GPU architectures support unified virtual memory (UVM), which offers great opportunities to solve larger problems by memory oversubscription. Although some studies are concerned over the performance degradation under UVM oversubscription, the reasons behind workloads' diverse sensitivities to oversubscription is still unclear. In this work, we take the first step to select various benchmark applications and conduct rigorous experiments on their performance under different oversubscription ratios. Specifically,we take into account the variety of memory access patterns and explain applications' diverse sensitivities to oversubscription. We also consider prefetching and UVM hints, and discover their complex impact under different oversubscription ratios. Moreover, the strengths and pitfalls of UVM's multi-GPU support are discussed. We expect that this paper will provide useful experiences and insights for UVM system design. Chuanming Shao, Jinyang Guo 0001, Pengyu Wang 0003, Jing Wang 0055, Chao Li 0009, Minyi Guo |
ICPE | 2 |
| 2021 | ACE-GCN: A Fast Data-driven FPGA Accelerator for GCN EmbeddingabstractACE-GCN is a fast and resource/energy-efficient FPGA accelerator for graph convolutional embedding under data-driven and in-place processing conditions. Our accelerator exploits the inherent power law distribution and high sparsity commonly exhibited by real-world graphs datasets. Contrary to other hardware implementations of GCN, on which traditional optimization techniques are employed to bypass the problem of dataset sparsity, our architecture is designed to take advantage of this very same situation. We propose and implement an innovative acceleration approach supported by our “implicit-processing-by-association” concept, in conjunction with a dataset-customized convolutional operator. The computational relief and consequential acceleration effect arise from the possibility of replacing rather complex convolutional operations for a faster embedding result estimation. Based on a computationally inexpensive and super-expedited similarity calculation, our accelerator is able to decide from the automatic embedding estimation or the unavoidable direct convolution operation. Evaluations demonstrate that our approach presents excellent applicability and competitive acceleration value. Depending on the dataset and efficiency level at the target, between 23× and 4,930× PyG baseline, coming close to AWB-GCN by 46% to 81% on smaller datasets and noticeable surpassing AWB-GCN for larger datasets and with controllable accuracy loss levels. We further demonstrate the unique hardware optimization characteristics of our approach and discuss its multi-processing potentiality. José Romero Hung, Chao Li 0009, Pengyu Wang 0003, Chuanming Shao, Jinyang Guo 0001, Jing Wang 0055, Guoyong Shi |
ACM Trans. Reconfigurable Technol. Syst. | 5 |