EDBT 2026 Demo / reviewers in the wild / expert
Xiaolong Xie
dblp:83/9047
· DBLP profile ↗
24ranked-venue papers
9as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 26% Graph learning · 23% Representation and self-supervised learning · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
GPUs and heterogeneous computing · 41% Memory systems · 22% Parallel and multicore computing · 13% | |
| Network and information security
1 paper |
Hardware security and side channels · 33% Systems and software security · 33% Authentication and access control · 33% | |
| Databases, data mining, and information retrieval
3 papers |
Database system architecture and tuning · 42% Query processing and optimization · 42% Data mining · 16% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
1.0 | 1 | 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning · AAAI 2026 |
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement |
1.0 | 1 | 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning · AAAI 2026 |
Computer vision › 3D vision › 3d object detection › label-efficient 3d object detection
unsupervised 3d object detection |
1.0 | 1 | 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning · AAAI 2026 |
Machine learning › Graph learning › graph pre-training
graph neural network pre-training |
0.9 | 1 | 2025 | Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Graph learning
graph representation learning |
0.9 | 1 | 2025 | Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Representation and self-supervised learning
information bottleneck |
0.9 | 1 | 2025 | Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks · IEEE Trans. Knowl. Data Eng. 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization Experiences · ISCA 2025 |
GPUs and heterogeneous computing
GPU computing |
0.8 | 3 | 2022 | CuLDA: Solving Large-scale LDA Problems on GPUs · HPDC 2019 CuMF_SGD: Parallelized Stochastic Gradient Descent for Matrix Factorization on GPUs · HPDC 2017 Accelerating multi-way joins on the GPU · VLDB J. 2022 |
Memory systems › cache management › cache insertion policy
cache bypassing |
0.8 | 3 | 2018 | Optimizing Cache Bypassing and Warp Scheduling for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 An Efficient Compiler Framework for Cache Bypassing on GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 Coordinated static and dynamic cache bypassing for GPUs · HPCA 2015 |
GPUs and heterogeneous computing
GPU architecture |
0.8 | 3 | 2018 | CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUs · IEEE Trans. Computers 2018 Enabling coordinated register allocation and thread-level parallelism optimization for GPUs · MICRO 2015 Coordinated static and dynamic cache bypassing for GPUs · HPCA 2015 |
Query processing and optimization › join processing
multi-way join |
0.6 | 1 | 2022 | Accelerating multi-way joins on the GPU · VLDB J. 2022 |
Systems and software security › database security
encrypted database |
0.6 | 1 | 2022 | Operon: An Encrypted Database for Ownership-Preserving Data Management · Proc. VLDB Endow. 2022 |
Hardware security and side channels
trusted execution environments |
0.6 | 1 | 2022 | Operon: An Encrypted Database for Ownership-Preserving Data Management · Proc. VLDB Endow. 2022 |
GPUs and heterogeneous computing › GPU memory management
GPU cache management |
0.5 | 2 | 2018 | Optimizing Cache Bypassing and Warp Scheduling for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 An Efficient Compiler Framework for Cache Bypassing on GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
0.5 | 2 | 2018 | CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUs · IEEE Trans. Computers 2018 Enabling coordinated register allocation and thread-level parallelism optimization for GPUs · MICRO 2015 |
Memory systems
cache management |
0.4 | 3 | 2018 | Coordinated static and dynamic cache bypassing for GPUs · HPCA 2015 Optimizing Cache Bypassing and Warp Scheduling for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 An Efficient Compiler Framework for Cache Bypassing on GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Machine learning › Efficient and distributed learning › resource allocation
workload partitioning |
0.4 | 1 | 2019 | CuLDA: Solving Large-scale LDA Problems on GPUs · HPDC 2019 |
GPUs and heterogeneous computing
GPU-accelerated machine learning |
0.4 | 1 | 2019 | CuLDA_CGS: solving large-scale LDA problems on GPUs · PPoPP 2019 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.3 | 1 | 2018 | Optimizing Cache Bypassing and Warp Scheduling for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Parallel and multicore computing › parallel programming runtimes › thread management
thread throttling |
0.3 | 2 | 2018 | Enabling coordinated register allocation and thread-level parallelism optimization for GPUs · MICRO 2015 CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUs · IEEE Trans. Computers 2018 |
Robotics › Autonomous driving
perception |
0.3 | 1 | 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning · AAAI 2026 |
Machine learning › Representation and self-supervised learning
matrix factorization |
0.3 | 1 | 2017 | CuMF_SGD: Parallelized Stochastic Gradient Descent for Matrix Factorization on GPUs · HPDC 2017 |
Machine learning › Optimization for machine learning › distributed optimization
parallel stochastic gradient descent |
0.3 | 1 | 2017 | CuMF_SGD: Parallelized Stochastic Gradient Descent for Matrix Factorization on GPUs · HPDC 2017 |
Parallel and multicore computing › parallelization strategies
model and data parallelism |
0.3 | 1 | 2017 | CuMF_SGD: Parallelized Stochastic Gradient Descent for Matrix Factorization on GPUs · HPDC 2017 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.3 | 1 | 2025 | Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks · IEEE Trans. Knowl. Data Eng. 2025 |
Compilers and program optimization
compiler analysis |
0.2 | 1 | 2015 | Coordinated static and dynamic cache bypassing for GPUs · HPCA 2015 |
Compilers and program optimization › accelerator compilation
GPU compiler optimization |
0.2 | 1 | 2015 | An Efficient Compiler Framework for Cache Bypassing on GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Memory systems › cache management
cache contention mitigation |
0.2 | 1 | 2015 | Enabling coordinated register allocation and thread-level parallelism optimization for GPUs · MICRO 2015 |
GPUs and heterogeneous computing
GPU query processing |
0.2 | 1 | 2022 | Accelerating multi-way joins on the GPU · VLDB J. 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
case-based reasoning |
0.2 | 1 | 2013 | Handling missing values and unmatched features in a CBR system for hydro-generator design · Comput. Aided Des. 2013 |
Methods — techniques the papers use, named apart from their topics
data compression · 1.5Intel SGX · 1.1GPU acceleration · 1.1FPGA-based TEE · 1.1synchronization · 1.1sampling optimization · 1.1self-training · 1.0occupancy guided warm-up · 1.0large model priors reasoning · 1.0self-supervised pretraining · 0.9model-chip co-design · 0.9information bottleneck · 0.9workload partitioning · 0.4collapsed gibbs sampling · 0.4dynamic scheduling · 0.3dynamic register allocation · 0.3compile-time register allocation · 0.3model parallelism · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors ReasoningabstractUnsupervised 3D object detection leverages heuristic algorithms to discover potential objects, offering a promising route to reduce annotation costs in autonomous driving. Existing approaches mainly generate pseudo labels and refine them through self-training iterations. However, these pseudo-labels are often incorrect at the beginning of training, resulting in misleading the optimization process. Moreover, effectively filtering and refining them remains a critical challenge. In this paper, we propose $\textbf{OWL}$ for unsupervised 3D object detection by occupancy guided warm-up and large-model priors reasoning. OWL first employs an Occupancy Guided Warm-up (OGW) strategy to initialize the backbone weight with spatial perception capabilities, mitigating the interference of incorrect pseudo-labels on network convergence. Furthermore, OWL introduces an Instance-Cued Reasoning (ICR) module that leverages the prior knowledge of large models to assess pseudo-label quality, enabling precise filtering and refinement. Finally, we design a WAS (Weight-adapted Self-training) strategy to dynamically re-weight pseudo-labels, improving the performance through self-training. Extensive experiments on Waymo Open Dataset (WOD) and KITTI demonstrate that OWL outperforms state-of-the-art unsupervised methods by over 15.0\% mAP, revealing the effectiveness of our method. Xusheng Guo, Wanfa Zhang, Shijia Zhao, Qiming Xia, Xiaolong Xie, Chenglu Wen |
AAAI | 5 |
| 2025 | Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization ExperiencesabstractThe rapid growth of AI workloads at Meta has motivated our inhouse development of AI chips, aiming to significantly reduce the total cost of ownership and mitigate risks posed by unpredictable GPU supplies.At ISCA'23, we presented Meta's first-generation AI chip, MTIA 1.This paper describes its successor, MTIA 2i, now deployed at scale and serving billions of users.MTIA 2i significantly improves upon MTIA 1, reducing total cost of ownership by 44% compared to GPUs while delivering competitive performance per watt.A key differentiator is its memory hierarchy: instead of costly HBM, it uses large SRAM alongside LPDDR.Although there has been a proliferation of publications on AI chips, they often focus on architectural design and overlook three critical aspects:(1) co-designing and optimizing ML models to work effectively with the AI chip; (2) demonstrating sufficient flexibility to support a wide range of models; and (3) during the productionization process, addressing challenges unanticipated or decisions deferred at design time, such as dealing with memory errors, safe overclocking, reducing provisioned power, and implementing real-time firmware updates to mitigate silicon design defects.A key contribution of this paper is sharing our experience with these aspects, based on our journey of productionizing MTIA 2i at scale. Joel Coburn, Chunqiang Tang, Sameer Abu Asal, Neeraj Agrawal, Raviteja Chinta, Harish Dattatraya Dixit, Brian Dodds, Saritha Dwarakapuram, Amin Firoozshahian, Cao Gao, Kaustubh Gondkar, Tyler Graf, Junhan Hu, Sterling Hughes, Adam Hutchin, Bhasker Jakka, Guoqiang Jerry Chen, Indu Kalyanaraman, Ashwin Kamath, Pankaj Kansal, Erum Kazi, Roman Levenstein, Mahesh Maddury, Alex Mastro, Siji Medaiyese, Pritesh Modi, Jack Montgomery, Nadathur Satish, Amit Nagpal, Ashwin Narasimha, Maxim Naumov, Eleanor Ozer, Jongsoo Park, Poorvaja Ramani, Harikrishna Reddy, David Reiss, Deboleena Roy, Sathish Sekar, Pavan Shetty, Aravind Sukumaran-Rajam, Eran Tal, Mike Tsai, Shreya Varshini, Richard Wareing, Olívia Wu, Xiaolong Xie, Hangchen Yu, Tanmay Zargar, Zitong Zeng, Feixiong Zhang, Ajit Mathews, Jiyuan Zhang 0008, Emmanuel Menage, Truls Edvard Stokke, Mohammed Sourouri |
ISCA | 48 |
| 2025 | Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural NetworksabstractPre-training GNNs to extract transferable knowledge and apply it to downstream tasks has become the de facto standard of graph representation learning. Recent works focused on designing self-supervised pre-training tasks to extract useful and universal transferable knowledge from large-scale unlabeled data. However, they have to face an inevitable question: traditional pre-training strategies that aim at extracting useful information about pre-training tasks, may not extract all useful information about the downstream task. In this paper, we reexamine the pre-training process within traditional pre-training and fine-tuning frameworks from the perspective of Information Bottleneck (IB) and confirm that the forgetting phenomenon in pre-training phase may cause detrimental effects on downstream tasks. Therefore, we propose a novelDelayedBottleneckingPre-training (DBP) framework which maintains as much as possible mutual information between latent representations and training data during pre-training phase by suppressing the compression operation and delays the compression operation to fine-tuning phase to make sure the compression can be guided with labeled fine-tuning data and downstream tasks. To achieve this, we design two information control objectives that can be directly optimized and further integrate them into the actual model design. Extensive experiments on both chemistry and biology domains demonstrate the effectiveness of DBP. Zhe Zhao 0008, Pengkun Wang 0001, Xu Wang 0029, Haibin Wen, Xiaolong Xie, Zhengyang Zhou, Qingfu Zhang 0001, Yang Wang 0015 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Operon: An Encrypted Database for Ownership-Preserving Data ManagementabstractThe past decade has witnessed the rapid development of cloud computing and data-centric applications. While these innovations offer numerous attractive features for data processing, they also bring in new issues about the loss of data ownership. Though some encrypted databases have emerged recently, they can not fully address these concerns for the data owner. In this paper, we propose an ownership-preserving database (OPDB), a new paradigm that characterizes different roles' responsibilities from nowadays applications and preserves data ownership throughout the entire application. We build Operon to follow the OPDB paradigm, which utilizes the trusted execution environment (TEE) and introduces a behavior control list (BCL). Different from access controls that merely handle accessibility permissions, BCL further makes data operation behaviors under control. Besides, we make Operon practical for real-world applications, by extending database capabilities towards flexibility, functionality and ease of use. Operon is the first database framework with which the data owner exclusively controls its data across different roles' subsystems. We have successfully integrated Operon with different TEEs, i.e. , Intel SGX and an FPGA-based implementation, and various database services on Alibaba Cloud, i.e. , PolarDB and RDS PostgreSQL. The evaluation shows that Operon achieves 71% - 97% of the performance of plaintext databases under the TPC-C benchmark while preserving the data ownership. Sheng Wang 0011, Huorong Li, Feifei Li 0001, Chengjin Tian, Le Su, Yanshan Zhang, Yubing Ma, Lie Yan, Xuntao Cheng, Xiaolong Xie |
Proc. VLDB Endow. | 12 |
| 2022 | Accelerating multi-way joins on the GPU
Zhuohang Lai, Xibo Sun, Qiong Luo 0001, Xiaolong Xie |
VLDB J. | 4 |
| 2019 | CuLDA: Solving Large-scale LDA Problems on GPUsabstractLatent Dirichlet Allocation(LDA) is a popular topic model. Given the fact that the input corpus of LDA algorithms consists of millions to billions of tokens, the LDA training process is very time-consuming, which prevents the adoption of LDA in many scenarios, e.g., online service. GPUs have benefited modern machine learning algorithms and big data analysis as they can provide high memory bandwidth and tremendous computation power. Therefore, many frameworks, e.g. TensorFlow, Caffe, CNTK, support GPUs for accelerating various data-intensive machine learning algorithms. However, we observe that the performance of existing LDA solutions on GPUs is not satisfying. In this paper, we present CuLDA, a GPU-based efficient and scalable approach to accelerate large-scale LDA problems. CuLDA is designed to efficiently solve LDA problems at high throughput. To this end, we first delicately design workload partitioning and synchronization mechanism to exploit multiple GPUs. Then, we offload the LDA sampling process to each individual GPU by optimizing from the sampling algorithm, parallelization, and data compression perspectives. Experiment evaluations show that compared with the state-of-the-art LDA solutions, CuLDA outperforms them by a large margin (up to 7.3X) on a single GPU. CuLDA is able to achieve an extra 7.5X speedup on 8 GPUs for large data sets. Xiaolong Xie, Yun Liang 0001, Wei Tan 0001 |
HPDC | 1 |
| 2019 | Efficient Data-Parallel Primitives on Heterogeneous SystemsabstractData-parallel primitives, such as gather, scatter, scan, and split, are widely used in data-intensive applications. However, it is challenging to optimize them on a system consisting of heterogeneous processors. In this paper, we study and compare the existing implementations and optimization strategies for a set of data-parallel primitives on three processors: GPU, CPU and Xeon Phi co-processor. Our goal is to identify the key performance factors in the implementations of data-parallel primitive operations on different architectures and develop general strategies for implementing these primitives efficiently on various platforms. We introduce a portable and efficient sequential memory access pattern, which eliminates the cost of adjusting the memory access pattern for individual device. With proper tuning, our optimized primitive implementations can achieve comparable performance to the native versions. Moreover, our profiling results show that the CPU and the Phi co-processor share most optimization strategies whereas the GPU differs from them significantly, due to the hardware differences among these devices, such as efficiency of vectorization, data and TLB caching, and data prefetching. We summarize these factors and deliver common primitive optimization strategies for heterogeneous systems. Zhuohang Lai, Qiong Luo 0001, Xiaolong Xie |
ICPP | 3 |
| 2019 | CuLDA_CGS: solving large-scale LDA problems on GPUsabstractGPUs have benefited many ML algorithms. However, we observe that the performance of existing Latent Dirichlet Allocation(LDA) solutions on GPUs are not satisfying. We present CuLDA_CGS, an efficient approach to accelerate large-scale LDA problems. We delicately design workload partition and synchronization mechanism to exploit multiple GPUs. We also optimize the algorithm from the sampling algorithm, parallelization, and data compression perspectives. Experiment evaluations show that compared with the state-of-the-art LDA solutions, CuLDA_CGS outperforms them by a large margin (up to 7.3X) on a single GPU. Xiaolong Xie, Yun Liang 0001, Wei Tan 0001 |
PPoPP | 1 |
| 2018 | CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUsabstractThe key to the high performance on GPUs lies in the massive threading to enable thread switching and hide long latencies. GPUs are equipped with a large register file to enable fast context switch. However, thread throttling techniques that are designed to mitigate cache contention, lead to under-utilization of registers. Register allocation is a significant factor for performance as it not just determines the single-thread performance, but indirectly affects the TLP. In this paper, we propose Coordinated Register Allocation and Thread-level parallelism (CRAT) to explore the optimization space of register allocation and TLP management on GPUs. CRATemploys both compile-time(CRAT-static) and run-time techniques(CRAT-dyn) to exhaust the design space. CRAT-static works statically to explore TLP and register allocation trade-off and CRAT-dyn exploits dynamic register allocation for further improvement. Experiments indicate that CRAT-static achieves an average 1.25X speedup over existing TLP management technique. On four register-limited applications, CRAT-dyn further improves the performance speedup of CRAT-static from 1.51X to 1.70X. Xiaolong Xie, Yun Liang 0001, Yudong Wu, Guangyu Sun 0003, Tao Wang 0004, Dongrui Fan |
IEEE Trans. Computers | 1 |
| 2018 | Optimizing Cache Bypassing and Warp Scheduling for GPUsabstractThe massive parallel architecture enables graphics processing units (GPUs) to boost performance for a wide range of applications. Initially, GPUs only employ scratchpad memory as on-chip memory. Recently, to broaden the scope of applications that can be accelerated by GPUs, GPU vendors have used caches as on-chip memory in the new generations of GPUs. Unfortunately, GPU caches face many performance challenges that arise due to the excessive thread contention for cache resource. Cache bypassing, where the memory requests can selectively bypass the cache, is one of the solutions that can help to mitigate the cache resource contention problem. In this paper, we propose coordinated static and dynamic cache bypassing to improve the GPU application performance. At compile-time, we identify the global loads that indicate strong preferences for caching or bypassing and encode the classification into the application binary. For the rest global loads, our dynamic cache bypassing has the flexibility to cache only a fraction of threads. In addition to coordinated bypassing, we also develop a bypass-aware warp scheduler to adaptively adjust the scheduling policy based on the cache performance. Evaluations show that our coordinated static and dynamic cache bypassing technique achieves up to$2.28\boldsymbol \times $(average$1.32\boldsymbol \times $) performance speedup for a variety of GPU applications. When we combine the coordinated cache bypassing with the bypass-aware scheduler, the average speedup is further improved to$1.38\boldsymbol \times $. Yun Liang 0001, Xiaolong Xie, Yu Wang 0002, Guangyu Sun 0003, Tao Wang 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | CuMF_SGD: Parallelized Stochastic Gradient Descent for Matrix Factorization on GPUsabstractStochastic gradient descent (SGD) is widely used by many machine learning algorithms. It is efficient for big data ap- plications due to its low algorithmic complexity. SGD is inherently serial and its parallelization is not trivial. How to parallelize SGD on many-core architectures (e.g. GPUs) for high efficiency is a big challenge. In this paper, we present cuMF_SGD, a parallelized SGD solution for matrix factorization on GPUs. We first design high-performance GPU computation kernels that accelerate individual SGD updates by exploiting model parallelism. We then design efficient schemes that parallelize SGD updates by exploiting data parallelism. Finally, we scale cuMF SGD to large data sets that cannot fit into one GPU's memory. Evaluations on three public data sets show that cuMF_SGD outperforms existing solutions, including a 64- node CPU system, by a large margin using only one GPU card. Xiaolong Xie, Wei Tan 0001, Liana L. Fong, Yun Liang 0001 |
HPDC | 1 |
| 2017 | Exploring cache bypassing and partitioning for multi-tasking on GPUsabstractGraphics Processing Units (GPUs) computing has become ubiquitous for embedded system, evidenced by its wide adoption for various general purpose applications. As more and more applications are accelerated by GPUs, multi-tasking scenario starts to emerge. Multi-tasking allows multiple applications to simultaneously execute on the same GPU and share the resource. This brings new challenges due to the contention among the different applications for the shared resources such as caches. However, the caches on GPUs are difficult to use. If used inappropriately, it may hurt the performance instead of improving it. In this paper, we propose to use cache partitioning together with cache bypassing as the shared cache management mechanism for multi-tasking on GPUs. The combined approach aims to reduce the interference among the tasks and preserve the locality for each task. However, the interplay among the cache partitioning and bypassing brings greater challenges. On one hand, the partitioned cache space to each task affects its cache bypassing decision. On the other hand, cache bypassing affects the cache capacity required for each task. To address this, we propose a two-step approach. First, we use cache partitioning to assign dedicated cache space to each task to reduce the interference among the tasks. During this process, we compare cache partitioning with coarse-grained cache bypassing. Then, we use fine-grained cache bypassing to selectively bypass certain data requests and threads for each task. We explore different cache partitioning and bypassing designs and demonstrate the potential benefits of this approach. Experiments using a wide range of applications demonstrate that our technique improves the overall system throughput by 52% on average compared to the default multi-tasking solution on GPUs. Yun Liang 0001, Xiaolong Xie |
ICCAD | 3 |
| 2017 | Random forests-based extreme learning machine ensemble for multi-regime time series prediction
Lin Lin 0014, Xiaolong Xie |
Expert Syst. Appl. | 3 |
| 2017 | Genetic algorithm optimized double-reservoir echo state network for multi-regime time series prediction
Xiaolong Xie, Lin Lin 0014 |
Neurocomputing | 2 |
| 2016 | Performance-centric register file design for GPUs using racetrack memoryabstractThe key to high performance for GPU architecture lies in massive threading to drive the large number of cores and enable overlapping of threading execution. However, in reality, the number of threads that can simultaneously execute is often limited by the size of the register file on GPUs. The traditional SRAM-based register file costs so large amount of chip area that it cannot scale to meet the increasing demand of massive threading for GPU applications. Racetrack memory is a promising technology for designing large capacity register file on GPUs due to its high data storage density. However, without careful deployment of registers, the lengthy shift operation of racetrack memory may hurt the performance. In this paper, we explore racetrack memory for designing high performance register file for GPU architecture. High storage density racetrack memory helps to improve the thread level parallelism, i.e., the number of threads that simultaneously execute. However, if the bits of the registers are not aligned to the ports, shift operations are required to move the bits to the ports. To mitigate the shift operation overhead problem, we develop a register file preshifting strategy and a compile-time managed register mapping algorithm. Experimental results demonstrate that our technique achieves up to 24% (19% on average) improvement in performance for a variety of GPU applications. Shuo Wang 0009, Yun Liang 0001, Chao Zhang 0007, Xiaolong Xie, Guangyu Sun 0003, Yongpan Liu, Yu Wang 0002 |
ASP-DAC | 4 |
| 2015 | Coordinated static and dynamic cache bypassing for GPUsabstractThe massive parallel architecture enables graphics processing units (GPUs) to boost performance for a wide range of applications. Initially, GPUs only employ scratchpad memory as on-chip memory. Recently, to broaden the scope of applications that can be accelerated by GPUs, GPU vendors have used caches in conjunction with scratchpad memory as on-chip memory in the new generations of GPUs. Unfortunately, GPU caches face many performance challenges that arise due to excessive thread contention for cache resource. Cache bypassing, where memory requests can selectively bypass the cache, is one solution that can help to mitigate the cache resource contention problem. In this paper, we propose coordinated static and dynamic cache bypassing to improve application performance. At compile-time, we identify the global loads that indicate strong preferences for caching or bypassing through profiling. For the rest global loads, our dynamic cache bypassing has the flexibility to cache only a fraction of threads. In CUDA programming model, the threads are divided into work units called thread blocks. Our dynamic bypassing technique modulates the ratio of thread blocks that cache or bypass at run-time. We choose to modulate at thread block level in order to avoid the memory divergence problems. Our approach combines compile-time analysis that determines the cache or bypass preferences for global loads with run-time management that adjusts the ratio of thread blocks that cache or bypass. Our coordinated static and dynamic cache bypassing technique achieves up to 2.28X (average I.32X) performance speedup for a variety of GPU applications. Xiaolong Xie, Yun Liang 0001, Yu Wang 0002, Guangyu Sun 0003, Tao Wang 0004 |
HPCA | 1 |
| 2015 | Enabling coordinated register allocation and thread-level parallelism optimization for GPUsabstractThe key to high performance on GPUs lies in the massive threading to enable thread switching and hide the latency of function unit and memory access. However, running with the maximum thread-level parallelism (TLP) does not necessarily lead to the optimal performance due to the excessive thread contention for cache resource. As a result, thread throttling techniques are employed to limit the number of threads that concurrently execute to preserve the data locality. On the other hand, GPUs are equipped with a large register file to enable fast context switch between threads. However, thread throttling techniques that are designed to mitigate cache contention, lead to under utilization of registers. Register allocation is a significant factor for performance as it not just determines the single-thread performance, but indirectly affects the TLP. Xiaolong Xie, Yun Liang 0001, Yudong Wu, Guangyu Sun 0003, Tao Wang 0004, Dongrui Fan |
MICRO | 1 |
| 2015 | Novel informative feature samples extraction model using cell nuclear pore optimization
Lin Lin 0014, Xiaolong Xie |
Eng. Appl. Artif. Intell. | 3 |
| 2015 | Two-layer random forests model for case reuse in case-based reasoning
Xiaolong Xie, Lin Lin 0014 |
Expert Syst. Appl. | 2 |
| 2015 | Novel adaptive hybrid rule network based on TS fuzzy rules using an improved quantum-behaved particle swarm optimization
Lin Lin 0014, Xiaolong Xie |
Neurocomputing | 3 |
| 2015 | An Efficient Compiler Framework for Cache Bypassing on GPUsabstractGraphics processing units (GPUs) have become ubiquitous for general purpose applications due to their tremendous computing power. Initially, GPUs only employ scratchpad memory as on-chip memory. Though scratchpad memory benefits many applications, it is not ideal for those general purpose applications with irregular memory accesses. Hence, GPU vendors have introduced caches in conjunction with scratchpad memory in the recent generations of GPUs. The caches on GPUs are highly configurable. The programmer or compiler can explicitly control cache access or bypass for global load instructions. This highly configurable feature of GPU caches opens up the opportunities for optimizing the cache performance. In this paper, we propose an efficient compiler framework for cache bypassing on GPUs. Our objective is to efficiently utilize the configurable cache and improve the overall performance for general purpose GPU applications. In order to achieve this goal, we first characterize GPU cache utilization and develop performance metrics to estimate the cache reuses and memory traffic. Next, we present efficient algorithms that judiciously select global load instructions for cache access or bypass. Finally, we present techniques to explore the unified cache and shared memory design space. We integrate our techniques into an automatic compiler framework that leverages parallel thread execution instruction set architecture to enable cache bypassing for GPUs. Experiments evaluation on NVIDIA GTX680 using a variety of applications demonstrates that compared to cache-all and bypass-all solutions, our techniques improve the performance from 4.6% to 13.1% for 16 KB L1 cache. Yun Liang 0001, Xiaolong Xie, Guangyu Sun 0003, Deming Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Process Takagi-Sugeno model: A novel approach for handling continuous input and output functions and its application to time series prediction
Xiaolong Xie, Lin Lin 0014 |
Knowl. Based Syst. | 1 |
| 2013 | An efficient compiler framework for cache bypassing on GPUsabstractGraphics Processing Units (GPUs) have become ubiquitous for general purpose applications due to their tremendous computing power. Initially, GPUs only employ scratchpad memory as on-chip memory. Though scratchpad memory benefits many applications, it is not ideal for those general purpose applications with irregular memory accesses. Hence, GPU vendors have introduced caches in conjunction with scratchpad memory in the recent generations of GPUs. The caches on GPUs are highly-configurable. The programmer or the compiler can explicitly control cache access or bypass for global load instructions. This highly-configurable feature of GPU caches opens up the opportunities for optimizing the cache performance. In this paper, we propose an efficient compiler framework for cache bypassing on GPUs. Our objective is to efficiently utilize the configurable cache and improve the overall performance for general purpose GPU applications. In order to achieve this goal, we first characterize GPU cache utilization and develop performance metrics to estimate the cache reuses and memory traffic. Next, we present efficient algorithms that judiciously select global load instructions for cache access or bypass. Finally, we integrate our techniques into an automatic compiler framework that leverages PTX instruction set architecture. Experiments evaluation demonstrates that compared to cache-all and bypass-all solutions, our techniques can achieve considerable performance improvement. Xiaolong Xie, Yun Liang 0001, Guangyu Sun 0003, Deming Chen |
ICCAD | 1 |
| 2013 | Handling missing values and unmatched features in a CBR system for hydro-generator design
Xiaolong Xie, Lin Lin 0014 |
Comput. Aided Des. | 1 |