Hui Guo 0004

dblp:70/2221-4 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-5131-0437ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 8 since 2021
YearPublicationVenuePosition
2025 SONet: Towards Practical Online Neural Network for Enhancing Hard-to-Predict Branches
Zhenxuan Xiong, Libo Huang 0002, Ling Yang 0008, Hui Guo 0004, Songwen Pei, Gang Chen 0023, Yongwen Wang
Euro-Par (2)4
2025 PolyPE: An Efficient Multi-Precision Multi-Mode Floating-Point Processing Element for HPC and AI
abstract
In this paper, an efficient multi-precision multimode floating-point Processing Element is designed for HPCenabled AI workloads, called PolyPE, in which Poly means multiprecision multi-mode. It supports both conventional and mixedprecision FMA operations, including single-FMA, dual-FMA, and quad-FMA modes, as well as quad-FMA-add for enhanced throughput. The supported precisions include double precision, single precision, half precision, TF32, and BF16. At each clock cycle, the processing element can perform one double-precision, two single-precision, or four half-precision operations. Compared to existing designs, it offers broader precision support, including TF32 and BF16, with higher throughput and lower hardware overhead, achieving up to 5× improvement over standard FMA. We integrated the design into an open-source GPGPU and extended its instruction set. Experimental results show up to 2.17× performance gain, with 27.2% and 41.2% reductions in LUT and FF usage, respectively, while preserving functional equivalence.
Zhenzhen Jia, Hongbing Tan, Ling Yang 0008, Hui Guo 0004, Junsheng Chang, Yongwen Wang, Libo Huang 0002
ICCD4
2025 Brief Announcement: LCTree: A Fast Hardware BVH Constructor for Real-Time Ray Tracing
abstract
Unlike traditional rasterization rendering, ray tracing is a groundbreaking technology that has revolutionized the realistic rendering of images, marking a significant leap forward. However, achieving real-time ray tracing in dynamic scene applications remains a challenging task. This difficulty arises primarily from the substantial technical bottlenecks related to the frequent need for reconstructing or incrementally updating acceleration structures essential for efficient ray calculations.
Run Yan, Su Yin, Hui Guo 0004, Yongwen Wang, Gang Chen 0023, Nong Xiao 0001, Libo Huang 0002
SPAA3
2024 QuickTree: A Fast Hardware BVH Construction Engine
abstract
Ray tracing has emerged as a powerful technique for generating visually stunning and realistic images compared to rasterization. With the continuous advancements in computer hardware, modern GPUs have integrated specialized ray tracing acceleration units to enhance rendering capabilities further. However, achieving realtime ray tracing presents a challenge in dynamic scenes, where spatial data structures used for accelerated rendering must be reconstructed or updated when there are changes in the scene primitives. This paper introduces QuickTree, a novel Bounding Volume Hierarchy (BVH) construction engine based on the linear BVH (LBVH) optimization algorithm. QuickTree addresses the challenge of dynamic scenes support by employing a highly parallel and pipelined system design. This innovative approach ensures fast construction speed. QuickTree demonstrates significant performance improvements. Compared to the currently fastest MergeTree, it has increased construction speed by 10% and reduced area by 45% compared to RayCore, which has the smallest chip area.
Yin Su, Hui Guo 0004, Run Yan, Yongwen Wang, Nong Xiao 0001, Gang Chen 0023, Libo Huang 0002
CF2
2024 A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AI
abstract
The dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs.
Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 MPRTA: An Efficient Multilevel Parallel Mobile Accelerator for High-Performance Ray Tracing
abstract
Ray tracing has been regarded as the future of graphics rendering technology for a long time. However, interactive ray tracing still faces challenges, especially in mobile devices, such as high computational intensity and multiple branches. In this brief, we aim to maximize overall efficiency by leveraging all forms of potential parallelism, including task, basic block, loop, and pipeline levels. We present multilevel parallel ray tracing accelerator (MPRTA), an innovative mobile accelerator that offers high performance and optimal efficiency for ray tracing. Experimental results indicate that MPRTA is$1.67\times $more efficient than the currently best-reported mobile accelerator.
Run Yan, Yin Su, Hui Guo 0004, Yashuai Lü, Nong Xiao 0001, Li Shen 0007, Yongwen Wang, Libo Huang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2023 A Scalable BFloat16 Dot-Product Architecture for Deep Learning
abstract
BFloat16(BF16) format has recently driven the development of deep learning due to its higher energy efficiency and less memory consumption than the traditional format. This paper presents a scalable BF16 dot-product(DoP) architecture for high-performance deep-learning computing. A novel 4-term DoP unit is proposed as a fundamental module in the architecture, which performs 4-term DoP operation in three cycles. More-term DoP units are constructed through the extension of the fundamental unit, in which early exponent comparison is performed to hide latency, and intermediate normalization and rounding are omitted to improve accuracy and further reduce latency. Compared with the discrete design, the proposed architecture reduces latency by 22.8% for 4-term DoP, and a larger proportion of latency is reduced as the size of the DoP operation increases. Compared with existing designs for BF16, the proposed architecture at 64-term exhibits better-normalized energy efficiency and higher throughput with at least 1.88× and 20.3× improvement, respectively.
Libo Huang 0002, Hongbing Tan, Hui Guo 0004
ACM Great Lakes Symposium on VLSI5
2021 GraphPEG: Accelerating Graph Processing on GPUs
abstract
Due to massive thread-level parallelism, GPUs have become an attractive platform for accelerating large-scale data parallel computations, such as graph processing. However, achieving high performance for graph processing with GPUs is non-trivial. Processing graphs on GPUs introduces several problems, such as load imbalance, low utilization of hardware unit, and memory divergence. Although previous work has proposed several software strategies to optimize graph processing on GPUs, there are several issues beyond the capability of software techniques to address. In this article, we present GraphPEG, a graph processing engine for efficient graph processing on GPUs. Inspired by the observation that many graph algorithms have a common pattern on graph traversal, GraphPEG improves the performance of graph processing by coupling automatic edge gathering with fine-grain work distribution. GraphPEG can also adapt to various input graph datasets and simplify the software design of graph processing with hardware-assisted graph traversal. Simulation results show that, in comparison with two representative highly efficient GPU graph processing software framework Gunrock and SEP-Graph, GraphPEG improves graph processing throughput by 2.8× and 2.5× on average, and up to 7.3× and 7.0× for six graph algorithm benchmarks on six graph datasets, with marginal hardware cost.
Ya-Shuai Lü, Hui Guo 0004, Libo Huang 0002, Qi Yu 0003, Li Shen 0007, Nong Xiao 0001, Zhiying Wang 0003
ACM Trans. Archit. Code Optim.2
2020 Coordinated Page Prefetch and Eviction for Memory Oversubscription Management in GPUs
abstract
The adoption of unified memory and demand paging has simplified programming and eased memory management in discrete GPUs. However, long-latency page faults cause significant performance overhead. While several software-based mechanisms have been proposed to address this issue, they suffer from inefficiency when page prefetching and pre-eviction are combined. For example, a state-of-the-art page replacement policy, hierarchical page eviction (HPE), is inefficient when prefetching is enabled. Furthermore, the prefetcher semantics-aware pre-evicting policy, which pre-evicts continuous pages in bulk the way they were brought in by the prefetcher, may cause thrashing for some irregular applications.In this paper, coordinated page prefetch and eviction (CPPE) is proposed to manage memory oversubscription in GPUs with unified memory. CPPE incorporates a modified page eviction policy, MHPE, and an access pattern-aware prefetcher in a fine-grained manner: MHPE is aware of prefetch semantics and the prefetcher prefetches pages according to access patterns in eviction candidates selected by MHPE. Simulation results show that, when the GPU memory is 75% and 50% oversubscribed, CPPE achieves an average speedup of 1.56x and 1.64x (up to 10.97x) over the state-of-the-art baseline, which combines a sequential-local prefetcher and LRU pre-eviction policy. CPPE also outperforms other approaches, including Random/reserved LRU with the sequential-local prefetcher, and simply disabling prefetching under memory oversubscription.
Qi Yu 0003, Bruce R. Childers, Libo Huang 0002, Cheng Qian 0006, Hui Guo 0004, Zhiying Wang 0003
IPDPS5
2014 Improving Speculation Accuracy with Inter-thread Fetching Value Prediction
Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009
ICA3PP (2)4
2013 HEUSPEC: A Software Speculation Parallel Model
abstract
Conventional software speculative parallel models are facing challenge due to the increasing number of the processor core and the diversification of the application. The performance of the guest program under the software speculative parallel execution model is closely related to the speculation accuracy, the control overhead and the rollback overhead of the model. In order to improve the speculative accuracy and the load balance, as well as improve the overhead of the conventional model, in this paper, we proposed a novel speculative parallel model named HEUSPEC. The HEUSPEC includes 2 key techniques, the heuristic value prediction(HVP) and the dynamic task granularity resizing(DTGR). We have implemented the runtime system of the model in ANSI C language. The experiment results show that when the speedup of the HEUSPEC model can reach 4.51 on the average (12% higher than conventional model) when speculative depth equals to 7. Besides, it shows good scalability and lower memory cost.
Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009
ICPP4