EDBT 2026 Demo / reviewers in the wild / expert
Guoping Long
dblp:11/2292
· DBLP profile ↗
32ranked-venue papers
5as first author
4since 2021 · last 2026
0009-0006-3176-7572ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 8Databases, data management, data science and information retrieval · 6Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsabstractMeeting growing demands for low latency and cost efficiency in production-grade large language model (LLM) serving systems requires integrating advanced optimization techniques. However, dynamic and unpredictable input-output lengths of LLM, compounded by these optimizations, exacerbate the issues of workload variability, making it difficult to maintain high efficiency on AI accelerators, especially DSAs with tile-based programming models. To address this challenge, we introduce XY-Serve, a versatile, Ascend NPU native, end-to-end production LLM-serving system. The core idea is an abstraction mechanism that smooths out the workload variability by decomposing computations into unified, hardware-friendly, fine-grained meta primitives. Then, kernels can efficiently execute without concerning the irregularity of workload. After this abstraction mechanism, for Attention, we propose a meta-kernel that computes the basic pattern of GEMM-Softmax-GEMM with architectural-aware tile sizes. For Linear, we introduce a virtual padding scheme that adapts to dynamic shape changes while using highly efficient GEMM primitives with assorted fixed tile sizes. XY-Serve sits harmoniously with vLLM. Experimental results show up to 95% end-to-end throughput improvement compared with current publicly available baselines on Ascend NPUs. We also set a new performance record for Linear (average 14.6% faster) and Attention (average 21.5% faster) kernels relative to existing libraries. Lastly, we demonstrate the generality of our technologies on GPU platform. Mingcong Song, Xinru Tang, Fengfan Hou, Yipeng Ma, Runqiu Xiao, Hongjie Si, Dingcheng Jiang, Shouyi Yin, Yang Hu 0001, Guoping Long |
ASPLOS (1) | 12 |
| 2022 | AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesabstractThis work reveals that memory-intensive computation is a rising performance-critical factor in recent machine learning models. Due to a unique set of new challenges, existing ML optimizing compilers cannot perform efficient fusion under complex two-level dependencies combined with just-in-time demand. They face the dilemma of either performing costly fusion due to heavy redundant computation, or skipping fusion which results in massive number of kernels. Furthermore, they often suffer from low parallelism due to the lack of support for real-world production workloads with irregular tensor shapes. To address these rising challenges, we propose AStitch, a machine learning optimizing compiler that opens a new multi-dimensional optimization space for memory-intensive ML computations. It systematically abstracts four operator-stitching schemes while considering multi-dimensional optimization objectives, tackles complex computation graph dependencies with novel hierarchical data reuse, and efficiently processes various tensor shapes via adaptive thread mapping. Finally, AStitch provides just-in-time support incorporating our proposed optimizations for both ML training and inference. Although AStitch serves as a stand-alone compiler engine that is portable to any version of TensorFlow, its basic ideas can be generally applied to other ML frameworks and optimization compilers. Experimental results show that AStitch can achieve an average of 1.84x speedup (up to 2.73x) over the state-of-the-art Google's XLA solution across five production workloads. We also deploy AStitch onto a production cluster for ML workloads with thousands of GPUs. The system has been in operation for more than 10 months and saves about 20,000 GPU hours for 70,000 tasks per week. Zhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long, Kai Zhu 0004, Feiwen Zhu, Wenyi Zhao, Jun Yang 0052, Jidong Zhai, Shuaiwen Song, Wei Lin 0016 |
ASPLOS | 4 |
| 2022 | Efficient Pipeline Planning for Expedited Distributed DNN TrainingabstractTo train modern large DNN models, pipeline parallelism has recently emerged, which distributes the model across GPUs and enables different devices to process different microbatches in pipeline. Earlier pipeline designs allow multiple versions of model parameters to co-exist (similar to asynchronous training), and cannot ensure the same model convergence and accuracy performance as without pipelining. Synchronous pipelining has recently been proposed which ensures model performance by enforcing a synchronization barrier between training iterations. Nonetheless, the synchronization barrier requires waiting for gradient aggregation from all microbatches and thus delays the training progress. Optimized pipeline planning is needed to minimize such wait and hence the training time, which has not been well studied in the literature. This paper designs efficient, near-optimal algorithms for expediting synchronous pipeline-parallel training of modern large DNNs over arbitrary inter-GPU connectivity. Our algorithm framework comprises two components: a pipeline partition and device mapping algorithm, and a pipeline scheduler that decides processing order of microbatches over the partitions, which together minimize the per-iteration training time. We conduct thorough theoretical analysis, extensive testbed experiments and trace-driven simulation, and demonstrate our scheme can accelerate training up to 157% compared with state-of-the-art designs. Ziyue Luo, Xiaodong Yi 0001, Guoping Long, Shiqing Fan, Chuan Wu 0001, Jun Yang 0052, Wei Lin 0016 |
INFOCOM | 3 |
| 2021 | DAPPLE: a pipelined data parallel approach for training large modelsabstractIt is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an effective approach for improving device utilization. However, there are still several tricky issues to address: improving computing efficiency while ensuring convergence, and reducing memory usage without incurring additional computing costs. We propose DAPPLE, a synchronous training framework which combines data parallelism and pipeline parallelism for large DNN models. It features a novel parallelization strategy planner to solve the partition and placement problems, and explores the optimal hybrid strategies of data and pipeline parallelism. We also propose a new runtime scheduling algorithm to reduce device memory usage, which is orthogonal to re-computation approach and does not come at the expense of training throughput. Experiments show that DAPPLE planner consistently outperforms strategies generated by PipeDream's planner by up to 3.23× speedup under synchronous training scenarios, and DAPPLE runtime outperforms GPipe by 1.6× speedup of training throughput and saves 12% of memory consumption at the same time. Shiqing Fan, Zongyan Cao, Siyu Wang 0006, Zhen Zheng, Chuan Wu 0001, Guoping Long, Jun Yang 0052, Lixue Xia, Lansong Diao, Wei Lin 0016 |
PPoPP | 8 |
| 2020 | Optimizing distributed training deployment in heterogeneous GPU clustersabstractThis paper proposes HeteroG, an automatic module to accelerate deep neural network training in heterogeneous GPU clusters. To train a deep learning model with large amounts of data, distributed training using data or model parallelism has been widely adopted, mostly over homogeneous devices (GPUs, network bandwidth). Heterogeneous training environments may often exist in shared clusters with GPUs of different models purchased in different batches and network connections of different bandwidth availability (e.g., due to contention). Classic data parallelism does not work well in a heterogeneous cluster, while model-parallel training is hard to plan. HeteroG enables highly-efficient distributed training over heterogeneous devices, by automatically converting a single-GPU training model to a distributed one according to the deep learning graph and available resources. HeteroG embraces operation-level hybrid parallelism, communication architecture selection and execution scheduling, based on a carefully designed strategy framework exploiting both GNN-based learning and combinatorial optimization. We compare HeteroG with existing parallelism schemes and show that it achieves up-to 222% training speed-up. HeteroG also enables efficient training of large models over a set of heterogeneous devices where simple parallelism is infeasible. Xiaodong Yi 0001, Shiwei Zhang 0002, Ziyue Luo, Guoping Long, Lansong Diao, Chuan Wu 0001, Zhen Zheng, Jun Yang 0052, Wei Lin 0016 |
CoNEXT | 4 |
| 2020 | Fast Training of Deep Learning Models over Multiple GPUsabstractThis paper proposes FastT, a transparent module to work with the TensorFlow framework for automatically identifying a satisfying deployment and execution order of operations in DNN models over multiple GPUs, for expedited model training. We propose white-box algorithms to compute the strategies with small computing resource consumption in a short time. Recently, similar studies have been done to optimize device placement using reinforcement learning. Compared to those works which learn to optimize device placement of operations in several hours using large amounts of computing resources, our approach can find excellent device placement and execution order within minutes using the same computing node as for training. We design a list of scheduling algorithms to compute the device placement and execution order for each operation and also design an algorithm to split operations in the critical path to support fine-grained (mixed) data and model parallelism to further improve the training speed in each iteration. We compare FastT with representative strategies and obtain insights on the best strategies for training different types of DNN models based on extensive testbed experiments. Xiaodong Yi 0001, Ziyue Luo, Mengdi Wang 0001, Guoping Long, Chuan Wu 0001, Jun Yang 0052, Wei Lin 0016 |
Middleware | 5 |
| 2020 | Online Bayesian max-margin subspace learning for multi-view classification and regression
Jia He 0001, Changying Du, Fuzhen Zhuang, Qing He 0003, Guoping Long |
Mach. Learn. | 6 |
| 2017 | Nonlinear Maximum Margin Multi-View Learning with Adaptive KernelabstractExisting multi-view learning methods based on kernel function either require the user to select and tune a single predefined kernel or have to compute and store many Gram matrices to perform multiple kernel learning. Apart from the huge consumption of manpower, computation and memory resources, most of these models seek point estimation of their parameters, and are prone to overfitting to small training data. This paper presents an adaptive kernel nonlinear max-margin multi-view learning model under the Bayesian framework. Specifically, we regularize the posterior of an efficient multi-view latent variable model by explicitly mapping the latent representations extracted from multiple data views to a random Fourier feature space where max-margin classification constraints are imposed. Assuming these random features are drawn from Dirichlet process Gaussian mixtures, we can adaptively learn shift-invariant kernels from data according to Bochners theorem. For inference, we employ the data augmentation idea for hinge loss, and design an efficient gradient-based MCMC sampler in the augmented space. Having no need to compute the Gram matrix, our algorithm scales linearly with the size of training set. Extensive experiments on real-world datasets demonstrate that our method has superior performance. Jia He 0001, Changying Du, Changde Du, Fuzhen Zhuang, Qing He 0003, Guoping Long |
IJCAI | 6 |
| 2016 | Online Bayesian Max-Margin Subspace Multi-View Learning
Jia He 0001, Changying Du, Fuzhen Zhuang, Qing He 0003, Guoping Long |
IJCAI | 6 |
| 2016 | Online variational Bayesian Support Vector RegressionabstractTraditional Support Vector Regression (SVR) solvers require user pre-specified penalty (regularization) parameter as input and typically model the training data with maximum a posterior (MAP) principle. The resultant point estimates can be affected seriously by inappropriate regularization, outliers and noise, especially when training online. In this paper, we address the aforementioned problems by developing a Bayesian SVR model with the pseudo-likelihood and data augmentation idea. Then we perform variational posterior inference in an augmented variable space and the approximate posterior of model weights, rather than point estimates as in traditional SVR, are used to make robust predictions. Besides, once the approximate posterior is obtained from a given set of data, we can regard it as model prior when dealing with new arrival data, which leads to a natural way to extend our batch model to the online scenario. Experiments on several benchmark regression problems as well as a real vehicle accident rate prediction task show that our models have superior performance while inferring penalty parameter automatically. Siqi Deng, Kan Gao, Changying Du, Wenjing Ma, Guoping Long, Yucheng Li 0002 |
IJCNN | 5 |
| 2016 | GPU-FV: Realtime Fisher Vector and Its Applications in Video MonitoringabstractFisher vector has been widely used in many multimedia retrieval and visual recognition applications with good performance. However, the computation complexity prevents its usage in real-time video monitoring. In this work, we proposed and implemented GPU-FV, a fast Fisher vector extraction method with the help of modern GPUs. The challenge of implementing Fisher vector on GPUs lies in the data dependency in feature extraction and expensive memory access in Fisher vector computing. To handle these challenges, we carefully designed GPU-FV in a way that utilizes the computing power of GPU as much as possible, and applied optimizations such as loop tiling to boost the performance. GPU-FV is about 12 times faster than the CPU version, and 50\% faster than a non-optimized GPU implementation. For standard video input (320*240), GPU-FV can process each frame within 34ms on a model GPU. Our experiments show that GPU-FV obtains a similar recognition accuracy as traditional FV on VOC 2007 and Caltech 256 image sets. We also applied GPU-FV for realtime video monitoring tasks and found that GPU-FV outperforms a number of previous works. Especially, when the number of training examples are small, GPU-FV outperforms the recent popular deep CNN features borrowed from ImageNet. Wenjing Ma, Liangliang Cao, Lei Yu 0012, Guoping Long, Yucheng Li 0002 |
ICMR | 4 |
| 2016 | Bayesian Group Feature Selection for Support Vector Learning Machines
Changde Du, Changying Du, Shandian Zhe, A-Li Luo, Qing He 0003, Guoping Long |
PAKDD (1) | 6 |
| 2016 | Efficient Bayesian Maximum Margin Multiple Kernel Learning
Changying Du, Changde Du, Guoping Long, Xin Jin 0004, Yucheng Li 0002 |
ECML/PKDD (1) | 3 |
| 2016 | Learning Beyond Predefined Label Space via Bayesian Nonparametric Topic Modelling
Changying Du, Fuzhen Zhuang, Jia He 0001, Qing He 0003, Guoping Long |
ECML/PKDD (1) | 5 |
| 2016 | Online Bayesian Multiple Kernel Bipartite Ranking
Changying Du, Changde Du, Guoping Long, Qing He 0003, Yucheng Li 0002 |
UAI | 3 |
| 2016 | Bridging Semantic Gap Between App Names: Collective Matrix Factorization for Similar Mobile App Recommendation
Ning Bu, Shuzi Niu, Lei Yu 0012, Wenjing Ma, Guoping Long |
WISE (2) | 5 |
| 2016 | Highly Optimized Code Generation for Stencil Codes with Computation Reuse for GPUs
Wenjing Ma, Kan Gao, Guoping Long |
J. Comput. Sci. Technol. | 3 |
| 2015 | PE-TLD: Parallel Extended Tracking-Learning-Detection for Multi-target Tracking
Chenggang Zhou, Qiankun Dong, Wenjing Ma, Guoping Long, Tao Li 0022 |
ICA3PP (2) | 4 |
| 2015 | Listwise Approach for Rank Aggregation in CrowdsourcingabstractInferring a gold-standard ranking over a set of objects, such as documents or images, is a key task to build test collections for various applications like Web search and recommender systems. Crowdsourcing services provide an efficient and inexpensive way to collect judgments via labeling by sets of annotators. We thus study the problem of finding a consensus ranking from crowdsourced judgments. In contrast to conventional rank aggregation methods which minimize the distance between predicted ranking and input judgments from either pointwise or pairwise perspective, we argue that it is critical to consider the distance in a listwise way to emphasize the position importance in ranking. Therefore, we introduce a new listwise approach in this paper, where ranking measure based objective functions are utilized for optimization. In addition, we also incorporate the annotator quality into our model since the reliability of annotators can vary significantly in crowdsourcing. For optimization, we transform the optimization problem to the Linear Sum Assignment Problem, and then solve it by a very efficient algorithm named CrowdAgg guaranteeing the optimal solution. Experimental results on two benchmark data sets from different crowdsourcing tasks show that our algorithm is much more effective, efficient and robust than traditional methods. Shuzi Niu, Yanyan Lan, Jiafeng Guo, Xueqi Cheng 0001, Lei Yu 0012, Guoping Long |
WSDM | 6 |
| 2013 | StreamScan: fast scan algorithms for GPUs without global barrier synchronizationabstractScan (also known as prefix sum) is a very useful primitive for various important parallel algorithms, such as sort, BFS, SpMV, compaction and so on. Current state of the art of GPU based scan implementation consists of three consecutive Reduce-Scan-Scan phases. This approach requires at least two global barriers and 3N (N is the problem size) global memory accesses. In this paper we propose StreamScan, a novel approach to implement scan on GPUs with only one computation phase. The main idea is to restrict synchronization to only adjacent workgroups, and thereby eliminating global barrier synchronization completely. The new approach requires only 2N global memory accesses and just one kernel invocation. On top of this we propose two important op-timizations to further boost performance speedups, namely thread grouping to eliminate unnecessary local barriers, and register optimization to expand the on chip problem size. We designed an auto-tuning framework to search the parameter space automatically to generate highly optimized codes for both AMD and Nvidia GPUs. We implemented our technique with OpenCL. Compared with previous fast scan implementations, experimental results not only show promising performance speedups, but also reveal dramatic different optimization tradeoffs between Nvidia and AMD GPU platforms. Shengen Yan, Guoping Long, Yunquan Zhang |
PPoPP | 2 |
| 2013 | MPFFT: An Auto-Tuning FFT Library for OpenCL GPUs
Yan Li 0005, Yunquan Zhang, Yiqung Liu 0005, Guoping Long, Haipeng Jia |
J. Comput. Sci. Technol. | 4 |
| 2012 | GPURoofline: A Model for Guiding Performance Optimizations on GPUs
Haipeng Jia, Yunquan Zhang, Guoping Long, Jianliang Xu, Shengen Yan, Yan Li 0005 |
Euro-Par | 3 |
| 2012 | An Insightful Program Performance Tuning Chain for GPU Computing
Haipeng Jia, Yunquan Zhang, Guoping Long, Shengen Yan |
ICA3PP (1) | 3 |
| 2011 | CRSD: Application Specific Auto-tuning of SpMV for Diagonal Sparse Matrices
Xiangzheng Sun, Yunquan Zhang, Guoping Long, Xianyi Zhang, Yan Li 0005 |
Euro-Par (2) | 4 |
| 2011 | Automatic FFT Performance Tuning on OpenCL GPUsabstractMany fields of science and engineering, such as astronomy, medical imaging, seismology and spectroscopy, have been revolutionized by Fourier methods. The fast Fourier transform (FFT) is an efficient algorithm to compute the discrete Fourier transform (DFT) and its inverse. The emerging class of high performance computing architectures, such as GPU, seeks to achieve much higher performance and efficiency by exposing a hierarchy of distinct memories to programmers. However, the complexity of GPU programming poses a significant challenge for programmers. In this paper, based on the Kronecker product form multi-dimensional FFTs, we propose an automatic performance tuning framework for various OpenCL GPUs. Several key techniques of GPU programming on AMD and NVIDIA GPUs are also identified. Our OpenCL FFT library achieves up to 1.5 to 4 times, 1.5 to 40 times and 1.4 times the performance of clAmdFft 1.0 for 1D, 2D and 3D FFT respectively on an AMD GPU, and the overall performance is within 90% of CUFFT 4.0 on two NVIDIA GPUs. Yan Li 0005, Yunquan Zhang, Haipeng Jia, Guoping Long |
ICPADS | 4 |
| 2010 | Minimal Multi-threading: Finding and Removing Redundant Instructions in Multi-threaded ProcessorsabstractParallelism is the key to continued performance scaling in modern microprocessors. Yet we observe that this parallelism can often contain a surprising amount of instruction redundancy. We propose to exploit this redundancy to improve performance and decrease energy consumption. We propose a multi-threading micro-architecture, Minimal Multi-Threading (MMT), that leverages register renaming and the instruction window to combine the fetch and execution of identical instructions between threads in SPMD applications. While many techniques exploit intra-thread similarities by detecting when a later instruction may use an earlier result, MMT exploits inter-thread similarities by, whenever possible, fetching instructions from different threads together and only splitting them if the computation is unique. With two threads, our design achieves a speedup of 1.15 (geometric mean) over a two-thread traditional SMT with a trace cache. With four threads, our design achieves a speedup of 1.25 (geometric mean) over a traditional SMT processor with four-threads and a trace cache. These correspond to speedups of 1.5 and 1.84 over a traditional out-of-order processor. Moreover, our performance increases in most applications with no power increase because the increase in overhead is countered with a decrease in cache accesses, leading to a decrease in energy consumption for all applications. Guoping Long, Diana Franklin, Susmit Biswas, Pablo J. Ortiz, Jason Oberg, Dongrui Fan, Fred Chong |
MICRO | 1 |
| 2009 | Characterizing and Understanding the Bandwidth Behavior of Workloads on Multi-core Processors
Guoping Long, Dongrui Fan, Junchao Zhang 0004 |
Euro-Par | 1 |
| 2009 | Architectural support for cilk computations on many-core architecturesabstractNo abstract available. Guoping Long, Dongrui Fan, Junchao Zhang 0004 |
PPoPP | 1 |
| 2009 | Godson-T: An Efficient Many-Core Architecture for Parallel Program Executions
Dongrui Fan, Nan Yuan, Junchao Zhang 0004, Yongbin Zhou, Wei Lin 0004, Fenglong Song, Xiaochun Ye, Lei Yu 0012, Guoping Long, Hao Zhang 0009 |
J. Comput. Sci. Technol. | 10 |
| 2008 | A Performance Model of Dense Matrix Operations on Many-Core Architectures
Guoping Long, Dongrui Fan, Junchao Zhang 0004, Fenglong Song, Nan Yuan, Wei Lin 0004 |
Euro-Par | 1 |
| 2008 | Location Consistency Model Revisited: Problem, Solution and ProspectsabstractLocation consistency (LC) is a weak memory consistency model which is defined entirely on partial order execution semantics of parallel programs. Compared with sequential consistency (SC), LC is scalable and provides ample theoretical parallelism. This makes LC an interesting memory model in the upcoming many-core parallel processing era. Previous work has pointed out that LC does not guarantee SC execution behavior for all data race free programs. In this paper, we compare the semantics of LC with PRAM consistency and memory coherence, and prove that LC is strictly weaker than PRAM consistency. For data race free programs, we prove that the semantics of LC is equivalent to memory coherence. In addition, by introducing memory ordering semantics into LC judiciously, we prove that the enhanced model is equivalent to SC for data race free programs. Finally, we discuss possible solutions for adding reasoning rules for LC-like weak memory models. Guoping Long, Nan Yuan, Dongrui Fan |
PDCAT | 1 |
| 2007 | Design and Implementation of Floating Point Stack on General RISC ArchitectureabstractThis paper presents a framework for implementing the X86 FP stack used in an x86-compliant processor based on a general RISC architecture. Architectural supports are added to a typical RISC architecture to maintain the FP stack status. Some speculative techniques are applied to the decode stage to enable pipelined and efficient FP operations. An optimized register renaming scheme is proposed to eliminate redundant micro-ops in FP programs, resulting in an increased performance while mitigating the burden on register rename table. The simulation results show that on average more than 10% fmov micro-ops are removed. Elimination of micro-ops significantly speeds up the execution of programs. The IPC increases are as high as 30% for some programs, and near 10% on average Xuehai Qian, Hao Zhang 0009, Guoping Long, Junchao Zhang 0004, Dongrui Fan |
PDP | 4 |