EDBT 2026 Demo / reviewers in the wild / expert
Xiaoxin Tang
dblp:65/10312
· DBLP profile ↗
7ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0002-7404-2073ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 34% Cloud and datacenter computing · 27% GPUs and heterogeneous computing · 21% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › inference serving
DNN serving |
0.5 | 1 | 2021 | E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services · IEEE Trans. Parallel Distributed Syst. 2021 |
GPUs and heterogeneous computing
GPU resource management |
0.5 | 1 | 2021 | E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services · IEEE Trans. Parallel Distributed Syst. 2021 |
Parallel and multicore computing
KNN search |
0.4 | 2 | 2015 | Scalable Multicore k-NN Search via Subspace Clustering for Filtering · IEEE Trans. Parallel Distributed Syst. 2015 Data filtering for scalable high-dimensional k-NN search on multicore systems · HPDC 2014 |
Parallel and multicore computing › parallel algorithms › shared-memory parallel algorithms
multicore algorithms |
0.4 | 2 | 2015 | Scalable Multicore k-NN Search via Subspace Clustering for Filtering · IEEE Trans. Parallel Distributed Syst. 2015 Data filtering for scalable high-dimensional k-NN search on multicore systems · HPDC 2014 |
Memory systems
memory wall |
0.2 | 1 | 2015 | Scalable Multicore k-NN Search via Subspace Clustering for Filtering · IEEE Trans. Parallel Distributed Syst. 2015 |
Memory systems › memory management
memory footprint reduction |
0.2 | 1 | 2014 | Data filtering for scalable high-dimensional k-NN search on multicore systems · HPDC 2014 |
Cloud and datacenter computing
quality of service |
0.1 | 1 | 2021 | E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services · IEEE Trans. Parallel Distributed Syst. 2021 |
Algorithms and data structures › similarity search
nearest neighbor search |
0.1 | 1 | 2015 | Scalable Multicore k-NN Search via Subspace Clustering for Filtering · IEEE Trans. Parallel Distributed Syst. 2015 |
Data mining
high-dimensional data |
0.1 | 1 | 2014 | Data filtering for scalable high-dimensional k-NN search on multicore systems · HPDC 2014 |
Methods — techniques the papers use, named apart from their topics
subspace clustering · 0.8elastic batch scheduling · 0.5data filtering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SHA: QoS-Aware Software and Hardware Auto-Tuning for Database Systems
Quan Chen 0002, Xiaoxin Tang, Minyi Guo |
J. Comput. Sci. Technol. | 3 |
| 2021 | E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning ServicesabstractWe aim to tackle existing problems about deep learning serving on GPUs in the view of the system. GPUs have been widely adopted to serve online deep learning-based services that have stringent QoS(Quality-of-Service) requirements. However, emerging deep learning serving systems often result in poor responsiveness and low throughput of the inferences that damage user experience and increase the number of GPUs required to host an online service. Our investigation shows that the poor batching operation and the lack of data transfer-computation overlap are the root causes of the poor responsiveness and low throughput. To this end, we propose E2bird, a deep learning serving system that is comprised of a GPU-resident memory pool, a multi-granularity inference engine, and an elastic batch scheduler. The memory pool eliminates the unnecessary waiting of the batching operation and enables data transfer-computation overlap. The inference engine enables concurrent execution of different batches, improving the GPU resource utilization. The batch scheduler organizes inferences elasticallyto guarantee the QoS. Our experimental results on an Nvidia Titan RTXGPU show that E2bird reduces the response latency of inferences by up to 82.4 percent and improves the throughput by up to 62.8 percent while guaranteeing the QoS target compared with TensorFlow Serving. Weihao Cui, Quan Chen 0002, Han Zhao 0005, Mengze Wei, Xiaoxin Tang, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Ebird: Elastic Batch for Improving Responsiveness and Throughput of Deep Learning ServicesabstractGPUs have been widely adopted to serve online deep learning-based services that have stringent QoS requirements. However, emerging deep learning serving systems often result in long latency and low throughput of the inference requests that damage user experience and increase the number of GPUs required to host an online service. Our investigation shows that the poor batching operation and the lacking of data transfer-computation overlap are the root causes of the long latency and low throughput. To this end, we propose Ebird, a deep learning serving system that is comprised of a GPU-resident memory pool, a multi-granularity inference engine, and an elastic batch scheduler. The memory pool eliminates the unnecessary waiting of the batching operation and enables data transfer-computation overlap. The inference engine enables concurrent execution of different batches, improving the GPUs resource utilization. The batch scheduler organizes inference requests elastically. Our experimental results on an Nvidia Titan RTX GPU show that Ebird reduces the response latency of inferences by up to 70.9% and improves the throughput by up to 49.3% while guaranteeing the QoS target compared with TensorFlow Serving. Weihao Cui, Mengze Wei, Quan Chen 0002, Xiaoxin Tang, Jingwen Leng, Li Li 0012, Mingyi Guo |
ICCD | 4 |
| 2015 | Efficient Selection Algorithm for Fast k-NN Search on GPUsabstractk Nearest Neighbours (k-NN) search is a fundamental problem in many computer vision and machine learning tasks. These tasks frequently involve a large number of high-dimensional vectors, which require intensive computations. Recent research work has shown that the Graphics Processing Unit (GPU) is a promising platform for solving k-NN search. However, these search algorithms often meet a serious bottleneck on GPUs due to a selection procedure, called k-selection, which is the final stage of k-NN and significantly affects the overall performance. In this paper, we propose new data structures and optimization techniques to accelerate k-selection on GPUs. Three key techniques are proposed: Merge Queue, Buffered Search and Hierarchical Partition. Compared with previous works, the proposed techniques can significantly improve the computing efficiency of k-selection on GPUs. Experimental results show that our techniques can achieve an up to 4:2× performance improvement over the state-of-the-art methods. Xiaoxin Tang, Zhiyi Huang 0001, David M. Eyers, Steven Mills, Minyi Guo |
IPDPS | 1 |
| 2015 | Scalable Multicore k-NN Search via Subspace Clustering for Filteringabstractk Nearest Neighbors (k-NN) search is a widely used category of algorithms with applications in domains such as computer vision and machine learning. Despite the desire to process increasing amounts of high-dimensional data within these domains, k-NN algorithms scale poorly on multicore systems because they hit a memory wall. In this paper, we propose a novel data filtering strategy for k-NN search algorithms on multicore platforms. By excluding unlikely features during the k-NN search process, this strategy can reduce the amount of computation required as well as the memory footprint. It is complementary to the data selection strategies used in other state-of-the-art k-NN algorithms. A Subspace Clustering for Filtering (SCF) method is proposed to implement the data filtering strategy. Experimental results on four k-NN algorithms show that SCF can significantly improve their performance on three modern multicore platforms with only a small loss of search precision. Xiaoxin Tang, Zhiyi Huang 0001, David M. Eyers, Steven Mills, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Data filtering for scalable high-dimensional k-NN search on multicore systemsabstractK Nearest Neighbors (k-NN) search is a widely used category of algorithms with applications in domains such as computer vision and machine learning. With the rapidly increasing amount of data available, and their high dimensionality, k-NN algorithms scale poorly on multicore systems because they hit a memory wall. In this paper, we propose a novel data filtering strategy, named Subspace Clustering for Filtering (SCF), for k-NN search algorithms on multicore platforms. By excluding unlikely features in k-NN search, this strategy can reduce memory footprint as well as computation. Experimental results on four k-NN algorithms show that SCF can improve their performance on two modern multicore platforms with insignificant loss of search precision. Xiaoxin Tang, Steven Mills, David M. Eyers, Kai-Cheung Leung, Zhiyi Huang 0001, Minyi Guo |
HPDC | 1 |
| 2013 | Performance Tuning on Multicore Systems for Feature Matching within Image CollectionsabstractParallel programming is the mainstream for today's HPC applications. Programmers need to parallelize their programs to achieve better performance on multicore systems. However, due to a lack of good understanding of parallelism in algorithms, scheduling policy in runtime systems, and multicore architectures, programmers usually find it very hard to write high-performance, scalable programs on these parallel platforms. Although using a parallelized library written by experts can reduce the amount of work for coding, it does not automatically guarantee good performance according to our study. A better understanding of parallelism in algorithms, the OS/runtime systems, and hardware architectures is necessary if programmers wish to further improve performance. In this paper, we use SIFT-based feature matching within large-scale image collections to show the importance of three factors-the level of parallelism, scheduling policy, and memory architecture-that affect the performance of large-scale feature matching on multicore systems. We demonstrate experimental results using programs based on OpenCV and OpenMP, which are executed on both 16-core and 64-core machines. From our experimental results, we find that images with a large number of features achieve poor scalability on the 64-core machine due to a poor cache utilization. To address this issue of cache performance, we propose a Divide-and-Merge algorithm that divides the feature space into several small sub-spaces so that they fit within the cache. Our experiments show that the performance tuning addressing all of the three factors improves the speedup of feature matching from 10.6× to 21.5× on the 64-core machine. While the speedup is improved by 103%, the scalability of the feature matching algorithm is improved by up to 6.45 times on the 64-core machine with our performance tuning. Our study indicates that performance tuning on multicore systems is very challenging even for a simple image processing algorithm. Xiaoxin Tang, Steven Mills, David M. Eyers, Zhiyi Huang 0001, Kai-Cheung Leung, Minyi Guo |
ICPP | 1 |