VLDB 2026 Research / reviewers in the wild / expert
Ruobing Han
dblp:236/4877
· DBLP profile ↗
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-3090-3951ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU ClustersabstractThe growing demand for GPU resources has led to widespread shortages in data centers, prompting the exploration of CPUs as an alternative for executing GPU programs. While prior research supports executing GPU programs on single CPUs, these approaches struggle to achieve competitive performance due to the computational capacity gap between GPUs and CPUs. Ruobing Han, Hyesoon Kim |
PPoPP | 1 |
| 2025 | Multiway Merge Partitioning for Sparse-Sparse Matrix Multiplication on GPUsabstractSparse-sparse matrix multiplication (SpGEMM) is a well-studied problem on CPUs, GPUs, accelerators (e.g. FPGAs), and distributed systems. The main computational bottleneck in SpGEMM is the reduction process, which involves matching indices to accumulate partial products that map to the same output locations and requires irregular memory accesses. Efficient implementations must use the memory hierarchy effectively so that this reduction is done in fast local (cache) memory as much as possible. This is challenging, especially on GPUs, where the local memory is managed explicitly, as different rows of the result may have vastly different numbers of nonzero elements and vastly different numbers of partial products required to produce the output row which may not fit in the local memory. We demonstrate an SpGEMM implementation on GPUs which perfectly partitions the partial products into equal-size blocks such that each block maps to disjoint output locations. In this way, the block size can be chosen to maximize local memory, all reductions happen in local memory, and each block can be processed independently without communication or data from other blocks. We show how this partitioning can be achieved by solving many instances of a multiway merge partitioning problem. There are several algorithms in the literature for solving this problem. We present the mathematical formulation, missing from all papers providing algorithms for this problem, and show that this is a useful framework for parallelizing it efficiently for GPUs. To our knowledge, this partitioning scheme has never been applied to SpGEMM. We evaluate our SpGEMM implementation, MMSpGEMM, with respect to state-of-the-art implementations including cuSPARSE, TileSpGEMM, spECK, and AC-SpGEMM, achieving speedups of up to $5.3 x, 10.0 x, 1.3 x$, and $1.1 x$, respectively, on a select set of 20 SuiteSparse matrices, and much higher on synthetic matrix multiplications with different left and right matrices. Finally, we discuss improvements and limitations and suggest ways to incorporate the idea into more general SpGEMM routines. Eric Lorimer, Ruobing Han, Sung Ha Kang, Hyesoon Kim |
PACT | 2 |
| 2025 | SoftCUDA: Running CUDA on Softcore GPUabstractField-Programmable Gate Arrays (FPGAs) have been extensively employed to accelerate parallel applications by allowing designers to customize their hardware for maximum performance. However, most FPGA-based designs are constrained to specific kernels, limiting their suitability across diverse workloads. As GPU workloads grow in complexity and their requirements diverge, softcore GPU(SoftGPU) designs have emerged to exploit FPGA reconfigurability for accelerating a broader range of parallel applications. Despite their potential, these designs have seen limited adoption due to the lack of comprehensive software stack support. In a CUDA-dominated development landscape, translating CUDA source code to alternative programming models can be challenging and often lacks direct feature parity. This paper introduces SoftCUDA, a novel framework that delivers comprehensive, end-to-end CUDA support on our SoftGPU, Vortex. By fully leveraging the reconfigurable architecture of SoftGPU and maintaining a user-friendly CUDA interface, SoftCUDA enables seamless integration and execution of unmodified CUDA applications on FPGA-based platforms. Chihyo Ahn, Ruobing Han, Udit Subramanya, Jisheng Zhao, Blaise-Pascal Tine, Hyesoon Kim |
FCCM | 2 |
| 2024 | Exponentially Expanding the Phase-Ordering Search Space via Dormant InformationabstractApplying compilation transformations in optimal sequences can significantly improve program speed and reduce code size. However, finding these optimal sequences—a problem known as the phase-ordering problem—remains a long-standing challenge. Specifically, modern compilers offer hundreds of available transformations, making the search space too large to explore efficiently within a reasonable timeframe. Existing solutions address this problem by grouping transformations into short sequences based on prior knowledge from human experts, and then searching for optimal orders among these sequences. Such pruning methods are aggressive, potentially excluding optimal solutions from the search space. Additionally, they rely on prior knowledge and lack scalability when applied to new transformations. Ruobing Han, Hyesoon Kim |
CC | 1 |
| 2024 | Enabling Fine-Grained Incremental Builds by Making Compiler StatefulabstractIncremental builds are commonly employed in software development, involving minor changes to existing source code that is then frequently recompiled. Speeding up incremental builds not only enhances the software development workflow but also improves CI/CD systems by enabling faster verification steps. Current solutions for incremental builds primarily rely on build systems that analyze file dependencies to avoid unnecessary recompilation of unchanged files. However, for the files that do undergo changes, these build systems simply invoke compilers to recompile them from scratch. This approach reveals a fundamental asymmetry in the system: while build systems operate in a stateful manner, compilers are stateless. As a result, incremental builds are applied only at a coarse-grained level, focusing on entire source files, rather than at a more fine-grained level that considers individual code sections. In this paper, we propose an innovative approach for enabling the fine-grained incremental build by introducing statefulness into compilers. Under this paradigm, the compiler leverages its profiling history to expedite the compilation process of modified source files, thereby reducing overall build time. Specifically, the stateful compiler retains dormant information of compiler passes executed in previous builds and uses this data to bypass dormant passes during subsequent incremental compilations. We also outline the essential changes needed to transform conventional stateless compilers into stateful ones. For practical evaluation, we modify the Clang compiler to adopt a stateful architecture and evaluate its performance on real-world C++ projects. Our comparative study indicates that the stateful version outperforms the standard Clang compiler in incremental builds, accelerating the end-to-end build process by an average of 6.72%. Ruobing Han, Jisheng Zhao, Hyesoon Kim |
CGO | 1 |
| 2024 | Unleashing CPU Potential for Executing GPU Programs Through Compiler/Runtime Optimizations
Ruobing Han, Jisheng Zhao, Hyesoon Kim |
MICRO | 1 |
| 2024 | CuPBoP: Making CUDA a Portable LanguageabstractCUDA is designed specifically for NVIDIA GPUs and is not compatible with non-NVIDIA devices. Enabling CUDA execution on alternative backends could greatly benefit the hardware community by fostering a more diverse software ecosystem. To address the need for portability, our objective is to develop a framework that meets key requirements, such as extensive coverage, comprehensive end-to-end support, superior performance, and hardware scalability. Existing solutions that translate CUDA source code into other high-level languages, however, fall short of these goals. In contrast to these source-to-source approaches, we present a novel framework, CuPBoP , which treats CUDA as a portable language in its own right. Compared to two commercial source-to-source solutions, CuPBoP offers a broader coverage and superior performance for the CUDA-to-CPU migration. Additionally, we evaluate the performance of CuPBoP against manually optimized CPU programs, highlighting the differences between CPU programs derived from CUDA and those that are manually optimized. Furthermore, we demonstrate the hardware scalability of CuPBoP by showcasing its successful migration of CUDA to AMD GPUs. To promote further research in this field, we have released CuPBoP as an open-source resource. Ruobing Han, Jun Chen 0038, Bhanu Garg, Xule Zhou, John Lu, Jeffrey Young 0001, Jaewoong Sim, Hyesoon Kim |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2023 | CuPBoP: A Framework to Make CUDA PortableabstractCUDA, as one of the most popular choices for GPU programming, can be executed only on NVIDIA GPUs. To execute CUDA on non-NVIDIA devices, researchers have proposed to translate CUDA to other programming languages. However, this approach cannot achieve high coverage due to the challenges in source-to-source translation. Ruobing Han, Jun Chen 0038, Bhanu Garg, Jeffrey Young 0001, Jaewoong Sim, Hyesoon Kim |
PPoPP | 1 |
| 2022 | COX : Exposing CUDA Warp-level Functions to CPUsabstractAs CUDA becomes the de facto programming language among data parallel applications such as high-performance computing or machine learning applications, running CUDA on other platforms becomes a compelling option. Although several efforts have attempted to support CUDA on devices other than NVIDIA GPUs, due to extra steps in the translation, the support is always a few years behind CUDA’s latest features. In particular, the new CUDA programming model exposes the warp concept in the programming language, which greatly changes the way the CUDA code should be mapped to CPU programs. In this article, hierarchical collapsing that correctly supports CUDA warp-level functions on CPUs is proposed. To verify hierarchical collapsing , we build a framework, COX , that supports executing CUDA source code on the CPU backend. With hierarchical collapsing , 90% of kernels in CUDA SDK samples can be executed on CPUs, much higher than previous works (68%). We also evaluate the performance with benchmarks for real applications and show that hierarchical collapsing can generate CPU programs with comparable or even higher performance than previous projects in general. Ruobing Han, Jaewoong Sim, Hyesoon Kim |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | GradientFlow: Optimizing Network Performance for Large-Scale Distributed DNN TrainingabstractIt is important to scale out deep neural network (DNN) training for reducing model training time. The high communication overhead is one of the major performance bottlenecks for distributed DNN training across multiple GPUs. Our investigations have shown that popular open-source DNN systems could only achieve 2.5 speedup ratio on 64 GPUs connected by 56 Gbps network. To address this problem, we propose a communication backend named GradientFlow for distributed DNN training, and employ a set of network optimization techniques. First, we integrate ring-based allreduce, mixed-precision training, and computation/communication overlap into GradientFlow. Second, we propose lazy allreduce to improve network throughput by fusing multiple communication operations into a single one, and design coarse-grained sparse communication to reduce network traffic by only transmitting important gradient chunks. When training AlexNet and ResNet-50 on the ImageNet dataset using 512 GPUs, our approach could achieve 410.2 and 434.1 speedup ratio, respectively. Peng Sun 0006, Yonggang Wen 0001, Ruobing Han, Wansen Feng, Shengen Yan |
IEEE Trans. Big Data | 3 |
| 2021 | Dynamic scaling for low-precision learningabstractIn recent years, distributed deep learning is becoming popular in industry and academia. Although researchers want to use distributed systems for training, it has been reported that the communication cost for synchronizing gradients can be a bottleneck. Using low-precision gradients is a promising technique for reducing the bandwidth requirement. In this work, we propose Auto Precision Scaling (APS), an algorithm that can improve the accuracy when we communicate gradients by low-precision floating-point values. APS can improve the accuracy for all precisions with a trivial communication cost. Our experimental results show that for both image classification and segmentation, applying APS can train the state-of-the-art models by 8-bit floating-point gradients with no or only a tiny accuracy loss (<0.05%). Furthermore, we can avoid any accuracy loss by designing a hybrid-precision technique. Finally, we propose a performance model to evaluate the proposed method. Our experimental results show that APS can get a significant speedup over the state-of-the-art method. To make it available to researchers and developers, we design and implement a high-performance system for customized precision Deep Learning(CPD), which can simulate the training process using an arbitrary low-precision customized floating-point format. We integrate CPD into PyTorch and make it open-source to the public1. Ruobing Han, Min Si, James Demmel, Yang You 0001 |
PPoPP | 1 |