EDBT 2026 Demo / reviewers in the wild / expert
Tatiana Jin
dblp:213/8884
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0001-7596-1146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Wings: Efficient Online Multiple Graph Pattern MatchingabstractFinding query patterns in a graph is fundamental for graph data analytics. Existing works mostly focus on either finding a single query pattern or finding patterns in a static graph. However, many applications today need to match multiple query patterns against a dynamically changing graph, i.e., online multiple graph pattern matching (online multi-GPM). Online multi-GPM is challenging as it requires quick responses for timely business decision-making. This paper proposes Wings — a distributed system for online multi-GPM. The key to efficient multi-GPM is a query planner that optimizes query plans by maximizing computation sharing among multiple queries and minimizing intermediate matching results. In addition, we also design an efficient query executor for Wings with memory footprint control and runtime redundant processing elimination. Our experimental results verified that Wings' designs are efficient for online multi-GPM. Guanxian Jiang, Yunjian Zhao, Zhi Liu 0002, Tatiana Jin, Wanying Zheng, Boyang Li 0016, James Cheng |
ICDE | 5 |
| 2023 | Circinus: Fast Redundancy-Reduced Subgraph MatchingabstractSubgraph matching is one of the most important problems in graph analytics. Many algorithms and systems have been proposed for subgraph matching. Most of these works follow Ullmann's backtracking approach as it is memory-efficient in handling an explosive number of intermediate matching results. However, they have largely overlooked an intrinsic problem of backtracking, namely repeated computation, which contributes to a large portion of the heavy computation in subgraph matching. This paper proposes a subgraph matching system, Circinus, which enables effective computation sharing by a new compression-based backtracking method. Our extensive experiments show that Circinus significantly reduces repeated computation, which transfers to up to several orders of magnitude performance improvement. Tatiana Jin, Boyang Li 0016, Qihui Zhou, Qianli Ma 0003, Yunjian Zhao, James Cheng |
Proc. ACM Manag. Data | 1 |
| 2022 | HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and OptimizationabstractGraph neural networks (GNNs) have shown to significantly improve graph analytics. Existing systems for GNN training are primarily designed for homogeneous graphs. In industry, however, most graphs are actually heterogeneous in nature (i.e., having multiple types of nodes and edges). Existing systems train a heterogeneous GNN (HetGNN) as a composition of homogeneous GNN (HomoGNN) and thus suffer from critical limitations such as lack of memory optimization and limited operator parallelism. To address these limitations, we propose HGL - a heterogeneity-aware system for GNN training. At the core of HGL is an intermediate representation, called HIR, which provides a holistic representation for GNNs and enables cross-relation optimization in HetGNN training. We devise tailored optimizations on HIR, including graph stitching, operator fusion and operator bundling. Compared with DGL and PyG, HGL achieves a speedup from 7 to 22 times for training HetGNNs. Yuntao Gui, Yidi Wu 0001, Han Yang 0002, Tatiana Jin, Boyang Li 0016, Qihui Zhou, James Cheng, Fan Yu 0004 |
SC | 4 |
| 2022 | VSGM: View-Based GPU-Accelerated Subgraph Matching on Large GraphsabstractSubgraph matching is a fundamental building block in graph analytics. Due to its high time complexity, GPU-based solutions have been proposed for sub graph matching. Most existing GPU-based works can only cope with relatively small graphs that fit in GPU memory. To support efficient subgraph matching on large graphs, we propose a view-based method to hide communication overhead and improve GPU utilization. We develop VSGM, a sub graph matching framework that supports efficient pipelined execution and multi-GPU architecture. Ex-tensive experimental evaluation shows that VSGM significantly outperforms the state-of-the-art solutions. Guanxian Jiang, Qihui Zhou, Tatiana Jin, Boyang Li 0016, Yunjian Zhao, James Cheng |
SC | 3 |
| 2021 | Seastar: vertex-centric programming for graph neural networksabstractGraph neural networks (GNNs) have achieved breakthrough performance in graph analytics such as node classification, link prediction and graph clustering. Many GNN training frameworks have been developed, but they are usually designed as a set of manually written, GNN-specific operators plugged into existing deep learning systems, which incurs high memory consumption, poor data locality, and large semantic gap between algorithm design and implementation. This paper proposes the Seastar system, which presents a vertex-centric programming model for GNN training on GPU and provides idiomatic python constructs to enable easy development of novel homogeneous and heterogeneous GNN models. We also propose novel optimizations to produce highly efficient fused GPU kernels for forward and backward passes in GNN training. Compared with the state-of-the art GNN systems, DGL and PyG, Seastar achieves better usability, up to 2 and 8 times less memory consumption, and 14 and 3 times faster execution, respectively. Yidi Wu 0001, Kaihao Ma, Zhenkun Cai, Tatiana Jin, Boyang Li 0016, Chengguang Zheng, James Cheng, Fan Yu 0004 |
EuroSys | 4 |
| 2021 | Vertex-Centric Visual Programming for Graph Neural NetworksabstractGraph neural networks (GNNs) have achieved remarkable performance in many graph analytics tasks such as node classification, link prediction and graph clustering. Existing GNN systems (e.g., PyG and DGL) adopt a tensor-centric programming model and train GNNs with manually written operators. Such design results in poor usability due to the large semantic gap between the API and the GNN models, and suffers from inferior efficiency because of high memory consumption and massive data movement. We demonstrateSeastar, a novel GNN training framework that adopts avertex-centric programming paradigm and supportsautomatic kernel generation, to simplify model development and improve training efficiency. We will (i) show how to express GNN models succinctly using a visual "drag-and-drop'' interface or Seastar's vertex-centric python API; (ii) demonstrate the performance advantage of Seastar over existing GNN systems in convergence speed, training throughput and memory consumption; and (iii) illustrate how Seastar's optimizations (e.g., operator fusion and constant folding) improve training efficiency by profiling the run-time performance. Yidi Wu 0001, Yuntao Gui, Tatiana Jin, James Cheng, Xiao Yan 0002, Peiqi Yin, Yufei Cai, Bo Tang 0016, Fan Yu 0004 |
SIGMOD Conference | 3 |
| 2021 | Scaling Large Production Clusters with Partitioned Synchronization
Yihui Feng, Zhi Liu 0002, Yunjian Zhao, Tatiana Jin, Yidi Wu 0001, James Cheng, Chao Li 0009 |
USENIX ATC | 4 |
| 2020 | Improving resource utilization by timely fine-grained schedulingabstractMonotask is a unit of work that uses only a single type of resource (e.g., CPU, network, disk I/O). While monotask was primarily introduced as a means to reason about job performance, in this paper we show that this fine-grained, resource-oriented abstraction can be leveraged by job schedulers to maximize cluster resource utilization. Although recent cluster schedulers have significantly improved resource allocation, the utilization of the allocated resources is often not high due to inaccurate resource requests. In particular, we show that existing scheduling mechanisms are ineffective for handling jobs with dynamic resource usage, which exists in common workloads, and propose a resource negotiation mechanism between job schedulers and executors that makes use of monotasks. We design a new framework, called Ursa, which enables the scheduler to capture accurate resource demands dynamically from the execution runtime and to provide timely, fine-grained resource allocation based on monotasks. Ursa also enables high utilization of the allocated resources by the execution runtime. We show by experiments that Ursa is able to improve cluster resource utilization, which effectively translates to improved makespan and average JCT. Tatiana Jin, Zhenkun Cai, Boyang Li 0016, Chengguang Zheng, Guanxian Jiang, James Cheng |
EuroSys | 1 |
| 2019 | Tangram: Bridging Immutable and Mutable Abstractions for Distributed Data Analytics
Xiao Yan 0002, Guanxian Jiang, Tatiana Jin, James Cheng, An Xu, Zhanhao Liu, Shuo Tu |
USENIX ATC | 4 |
| 2018 | FlexPS: Flexible Parallelism Control in Parameter Server ArchitectureabstractAs a general abstraction for coordinating the distributed storage and access of model parameters, the parameter server (PS) architecture enables distributed machine learning to handle large datasets and high dimensional models. Many systems, such as Parameter Server and Petuum, have been developed based on the PS architecture and widely used in practice. However, none of these systems supports changing parallelism during runtime, which is crucial for the efficient execution of machine learning tasks with dynamic workloads. We propose a new system, called FlexPS, which introduces a novel multi-stage abstraction to support flexible parallelism control. With the multi-stage abstraction, a machine learning task can be mapped to a series of stages and the parallelism for a stage can be set according to its workload. Optimizations such as stage scheduler, stage-aware consistency controller, and direct model transfer are proposed for the efficiency of multi-stage machine learning in FlexPS. As a general and complete PS systems, FlexPS also incorporates many optimizations that are not limited to multi-stage machine learning. We conduct extensive experiments using a variety of machine learning workloads, showing that FlexPS achieves significant speedups and resource saving compared with the state-of-the-art PS systems such as Petuum and Multiverso. Tatiana Jin, Yidi Wu 0001, Zhenkun Cai, Xiao Yan 0002, Fan Yang 0091, Yuying Guo, James Cheng |
Proc. VLDB Endow. | 2 |