EDBT 2026 Demo / reviewers in the wild / expert
Guanxian Jiang
dblp:199/6310
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-9837-0904ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hitcher: Efficient GPU-based Vector Search via Cluster-Centric Kernel and Hitch-Ride OrderingabstractSimilarity-based vector search, which retrieves the most similar vectors to a given query vector from a large vector dataset, underlies many applications such as search, recommendation, and Large Language Models (LLMs). Some systems run vector search on GPUs to enjoy GPU's high parallelism, but we observe that they are limited in query throughput and latency. In particular, their query-centric GPU kernel conducts computation independently for each query, failing to reuse data loaded to the GPU shared memory across queries and leading to a low GPU compute utilization. While their batch-based task reordering rearranges computation for queries in a batch to reduce CPU-GPU data transfer, but latency is prolonged since each query needs to wait for its slowest task. To tackle these problems, we propose Hitcher. Specifically, to reuse data across queries and improve GPU utilization, Hitcher implements a cluster-centric GPU kernel to batch computation on the same data for multiple queries. To reduce query latency, Hitcher adopts the hitch-ride ordering, which preserves the arrival order for query processing while batching computation across queries to improve efficiency. Hitcher can also offload computation tasks to the CPU to reduce CPU-GPU data transfer and utilize multiple GPUs. Experimental results show that Hitcher achieves up to 22× lower P99 query latency and 9× higher query throughput when compared with the state-of-the-art GPU-based vector query processing systems. Qihui Zhou, Changji Li, Guanxian Jiang, Chenhao Ma 0001, Xiao Yan 0002, Yu Mao 0001, Ming-Chang Yang, James Cheng |
KDD (1) | 3 |
| 2024 | Wings: Efficient Online Multiple Graph Pattern MatchingabstractFinding query patterns in a graph is fundamental for graph data analytics. Existing works mostly focus on either finding a single query pattern or finding patterns in a static graph. However, many applications today need to match multiple query patterns against a dynamically changing graph, i.e., online multiple graph pattern matching (online multi-GPM). Online multi-GPM is challenging as it requires quick responses for timely business decision-making. This paper proposes Wings — a distributed system for online multi-GPM. The key to efficient multi-GPM is a query planner that optimizes query plans by maximizing computation sharing among multiple queries and minimizing intermediate matching results. In addition, we also design an efficient query executor for Wings with memory footprint control and runtime redundant processing elimination. Our experimental results verified that Wings' designs are efficient for online multi-GPM. Guanxian Jiang, Yunjian Zhao, Zhi Liu 0002, Tatiana Jin, Wanying Zheng, Boyang Li 0016, James Cheng |
ICDE | 1 |
| 2024 | GE2: A General and Efficient Knowledge Graph Embedding Learning SystemabstractGraph embedding learning computes an embedding vector for each node in a graph and finds many applications in areas such as social networks, e-commerce, and medicine. We observe that existing graph embedding systems (e.g., PBG, DGL-KE, and Marius) have long CPU time and high CPU-GPU communication overhead, especially when using multiple GPUs. Moreover, it is cumbersome to implement negative sampling algorithms on them, which have many variants and are crucial for model quality. We propose a new system called GE 2 , which achieves both generality and efficiency for graph embedding learning. In particular, we propose a general execution model that encompasses various negative sampling algorithms. Based on the execution model, we design a user-friendly API that allows users to easily express negative sampling algorithms. To support efficient training, we offload operations from CPU to GPU to enjoy high parallelism and reduce CPU time. We also design COVER, which, to our knowledge, is the first algorithm to manage data swap between CPU and multiple GPUs for small communication costs. Extensive experimental results show that, comparing with the state-of-the-art graph embedding systems, GE 2 trains consistently faster across different models and datasets, where the speedup is usually over 2x and can be up to 7.5x. Chenguang Zheng, Guanxian Jiang, Xiao Yan 0002, Peiqi Yin, Qihui Zhou, James Cheng |
Proc. ACM Manag. Data | 2 |
| 2024 | Atom: An Efficient Query Serving System for Embedding-based Knowledge Graph Reasoning with Operator-level BatchingabstractKnowledge graph reasoning (KGR) answers logical queries over a knowledge graph (KG), and embedding-based KGR (EKGR) becomes popular recently, which embeds both queries and KG entities such that the vector embeddings of a query and its answer entities are similar. Compared with traditional KGR methods based on subgraph matching, EKGR produces fewer intermediate results and is more robust to missing and noisy information in the KG. However, existing systems are inefficient for serving online EKGR queries because they can only batch queries of the same type for execution (i.e., query-level batching ) and hence have limited batching opportunities due to the heterogeneity of queries. To serve EKGR queries efficiently, we propose the Atom system with operator-level batching, which decomposes queries into operators and batches operators of the same type from different queries for execution. The insight is that the types of operators are far fewer than the types of queries, and thus different queries typically share common operators, yielding more batching opportunities. To schedule the operators, Atom adopts a hybrid policy, which improves system throughput and avoids starving rare operators. For efficiency, Atom incorporates system optimizations including two-level pipeline, opportunistic submission, pre-allocated memory buffer, and tailored GPU kernels. Experiment results show that compared with existing systems, Atom can improve query throughput by over 20x and reduce query latency by over 5x. Micro experiments suggest that the designs and optimizations are effective in improving system performance. Qihui Zhou, Peiqi Yin, Xiao Yan 0002, Changji Li, Guanxian Jiang, James Cheng |
Proc. ACM Manag. Data | 5 |
| 2022 | VSGM: View-Based GPU-Accelerated Subgraph Matching on Large GraphsabstractSubgraph matching is a fundamental building block in graph analytics. Due to its high time complexity, GPU-based solutions have been proposed for sub graph matching. Most existing GPU-based works can only cope with relatively small graphs that fit in GPU memory. To support efficient subgraph matching on large graphs, we propose a view-based method to hide communication overhead and improve GPU utilization. We develop VSGM, a sub graph matching framework that supports efficient pipelined execution and multi-GPU architecture. Ex-tensive experimental evaluation shows that VSGM significantly outperforms the state-of-the-art solutions. Guanxian Jiang, Qihui Zhou, Tatiana Jin, Boyang Li 0016, Yunjian Zhao, James Cheng |
SC | 1 |
| 2021 | Timestamped State Sharing for Stream AnalyticsabstractState access in existing distributed stream processing systems is restricted locally within each operator. However, in advanced stream analytics such as online learning and dynamic graph analytics, enabling state sharing across different operators makes application development easier and stream processing more efficient. In addition, when stream records are timestamped, proper time semantics should be defined for both state updates and fetches. We propose a new state abstraction to address the limitations of existing systems and develop a distributed stream processing system, Nova, with native support for timestamped state sharing. We validate the expressiveness and efficiency of Nova with extensive experiments. Yunjian Zhao, Zhi Liu 0002, Yidi Wu 0001, Guanxian Jiang, James Cheng, Kunlong Liu, Xiao Yan 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Improving resource utilization by timely fine-grained schedulingabstractMonotask is a unit of work that uses only a single type of resource (e.g., CPU, network, disk I/O). While monotask was primarily introduced as a means to reason about job performance, in this paper we show that this fine-grained, resource-oriented abstraction can be leveraged by job schedulers to maximize cluster resource utilization. Although recent cluster schedulers have significantly improved resource allocation, the utilization of the allocated resources is often not high due to inaccurate resource requests. In particular, we show that existing scheduling mechanisms are ineffective for handling jobs with dynamic resource usage, which exists in common workloads, and propose a resource negotiation mechanism between job schedulers and executors that makes use of monotasks. We design a new framework, called Ursa, which enables the scheduler to capture accurate resource demands dynamically from the execution runtime and to provide timely, fine-grained resource allocation based on monotasks. Ursa also enables high utilization of the allocated resources by the execution runtime. We show by experiments that Ursa is able to improve cluster resource utilization, which effectively translates to improved makespan and average JCT. Tatiana Jin, Zhenkun Cai, Boyang Li 0016, Chengguang Zheng, Guanxian Jiang, James Cheng |
EuroSys | 5 |
| 2019 | Tangram: Bridging Immutable and Mutable Abstractions for Distributed Data Analytics
Xiao Yan 0002, Guanxian Jiang, Tatiana Jin, James Cheng, An Xu, Zhanhao Liu, Shuo Tu |
USENIX ATC | 3 |
| 2017 | The Best of Both Worlds: Big Data Programming with Both Productivity and PerformanceabstractCoarse-grained operators such as map and reduce have been widely used for large-scale data processing. While they are easy to master, over-simplified APIs sometimes hinder programmers from fine-grained control on how computation is performed and hence designing more efficient algorithms. On the other hand, resorting to domain-specific languages (DSLs) is also not a practical solution, since programmers may need to learn how to use many systems that can be very different from each other, and the use of low-level tools may even result in bug-prone programming. Fan Yang 0091, Yunjian Zhao, Guanxian Jiang, James Cheng |
SIGMOD Conference | 5 |