EDBT 2026 Demo / reviewers in the wild / expert
Yunjian Zhao
dblp:194/7654
· DBLP profile ↗
11ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0001-5486-331XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Wings: Efficient Online Multiple Graph Pattern MatchingabstractFinding query patterns in a graph is fundamental for graph data analytics. Existing works mostly focus on either finding a single query pattern or finding patterns in a static graph. However, many applications today need to match multiple query patterns against a dynamically changing graph, i.e., online multiple graph pattern matching (online multi-GPM). Online multi-GPM is challenging as it requires quick responses for timely business decision-making. This paper proposes Wings — a distributed system for online multi-GPM. The key to efficient multi-GPM is a query planner that optimizes query plans by maximizing computation sharing among multiple queries and minimizing intermediate matching results. In addition, we also design an efficient query executor for Wings with memory footprint control and runtime redundant processing elimination. Our experimental results verified that Wings' designs are efficient for online multi-GPM. Guanxian Jiang, Yunjian Zhao, Zhi Liu 0002, Tatiana Jin, Wanying Zheng, Boyang Li 0016, James Cheng |
ICDE | 2 |
| 2023 | Circinus: Fast Redundancy-Reduced Subgraph MatchingabstractSubgraph matching is one of the most important problems in graph analytics. Many algorithms and systems have been proposed for subgraph matching. Most of these works follow Ullmann's backtracking approach as it is memory-efficient in handling an explosive number of intermediate matching results. However, they have largely overlooked an intrinsic problem of backtracking, namely repeated computation, which contributes to a large portion of the heavy computation in subgraph matching. This paper proposes a subgraph matching system, Circinus, which enables effective computation sharing by a new compression-based backtracking method. Our extensive experiments show that Circinus significantly reduces repeated computation, which transfers to up to several orders of magnitude performance improvement. Tatiana Jin, Boyang Li 0016, Qihui Zhou, Qianli Ma 0003, Yunjian Zhao, James Cheng |
Proc. ACM Manag. Data | 6 |
| 2022 | VSGM: View-Based GPU-Accelerated Subgraph Matching on Large GraphsabstractSubgraph matching is a fundamental building block in graph analytics. Due to its high time complexity, GPU-based solutions have been proposed for sub graph matching. Most existing GPU-based works can only cope with relatively small graphs that fit in GPU memory. To support efficient subgraph matching on large graphs, we propose a view-based method to hide communication overhead and improve GPU utilization. We develop VSGM, a sub graph matching framework that supports efficient pipelined execution and multi-GPU architecture. Ex-tensive experimental evaluation shows that VSGM significantly outperforms the state-of-the-art solutions. Guanxian Jiang, Qihui Zhou, Tatiana Jin, Boyang Li 0016, Yunjian Zhao, James Cheng |
SC | 5 |
| 2021 | Scaling Large Production Clusters with Partitioned Synchronization
Yihui Feng, Zhi Liu 0002, Yunjian Zhao, Tatiana Jin, Yidi Wu 0001, James Cheng, Chao Li 0009 |
USENIX ATC | 3 |
| 2021 | Timestamped State Sharing for Stream AnalyticsabstractState access in existing distributed stream processing systems is restricted locally within each operator. However, in advanced stream analytics such as online learning and dynamic graph analytics, enabling state sharing across different operators makes application development easier and stream processing more efficient. In addition, when stream records are timestamped, proper time semantics should be defined for both state updates and fetches. We propose a new state abstraction to address the limitations of existing systems and develop a distributed stream processing system, Nova, with native support for timestamped state sharing. We validate the expressiveness and efficiency of Nova with extensive experiments. Yunjian Zhao, Zhi Liu 0002, Yidi Wu 0001, Guanxian Jiang, James Cheng, Kunlong Liu, Xiao Yan 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | G-Miner: an efficient task-oriented graph mining systemabstractGraph mining is one of the most important areas in data mining. However, scalable solutions for graph mining are still lacking as existing studies focus on sequential algorithms. While many distributed graph processing systems have been proposed in recent years, most of them were designed to parallelize computations such as PageRank and Breadth-First Search that keep states on individual vertices and propagate updates along edges. Graph mining, on the other hand, may generate many subgraphs whose number can far exceed the number of vertices. This inevitably leads to much higher computational and space complexity rendering existing graph systems inefficient. We propose G-Miner, a distributed system with a new architecture designed for general graph mining. G-Miner adopts a unified programming framework for implementing a wide range of graph mining algorithms. We model subgraph processing as independent tasks, and design a novel task pipeline to streamline task processing for better CPU, network and I/O utilization. Our extensive experiments validate the efficiency of G-Miner for a range of graph mining tasks. Miao Liu 0006, Yunjian Zhao, Xiao Yan 0002, Da Yan 0001, James Cheng |
EuroSys | 3 |
| 2017 | Efficient Processing of Growing Temporal Graphs
Huanhuan Wu, Yunjian Zhao, James Cheng, Da Yan 0001 |
DASFAA (2) | 2 |
| 2017 | LoSHa: A General Framework for Scalable Locality Sensitive HashingabstractLocality Sensitive Hashing (LSH) algorithms are widely adopted to index similar items in high dimensional space for approximate nearest neighbor search. As the volume of real-world datasets keeps growing, it has become necessary to develop distributed LSH solutions. Implementing a distributed LSH algorithm from scratch requires high development costs, thus most existing solutions are developed on general-purpose platforms such as Hadoop and Spark. However, we argue that these platforms are both hard to use for programming LSH algorithms and inefficient for LSH computation. We propose LoSHa, a distributed computing framework that reduces the development cost by designing a tailor-made, general programming interface and achieves high efficiency by exploring LSH-specific system implementation and optimizations. We show that many LSH algorithms can be easily expressed in LoSHa's API. We evaluate LoSHa and also compare with general-purpose platforms on the same LSH algorithms. Our results show that LoSHa's performance can be an order of magnitude faster, while the implementations on LoSHa are even more intuitive and require few lines of code. James Cheng, Fan Yang 0091, Yunjian Zhao, Xiao Yan 0002, Ruihao Zhao |
SIGIR | 5 |
| 2017 | The Best of Both Worlds: Big Data Programming with Both Productivity and PerformanceabstractCoarse-grained operators such as map and reduce have been widely used for large-scale data processing. While they are easy to master, over-simplified APIs sometimes hinder programmers from fine-grained control on how computation is performed and hence designing more efficient algorithms. On the other hand, resorting to domain-specific languages (DSLs) is also not a practical solution, since programmers may need to learn how to use many systems that can be very different from each other, and the use of low-level tools may even result in bug-prone programming. Fan Yang 0091, Yunjian Zhao, Guanxian Jiang, James Cheng |
SIGMOD Conference | 3 |
| 2017 | LFTF: A Framework for Efficient Tensor Analytics at ScaleabstractTensors are higher order generalizations of matrices to model multi-aspect data, e.g., a set of purchase records with the schema (user_id, product_id, timestamp, feedback). Tensor factorization is a powerful technique for generating a model from a tensor, just like matrix factorization generates a model from a matrix, but with higher accuracy and richer information as more attributes are available in a higher- order tensor than a matrix. The data model obtained by tensor factorization can be used for classification, recommendation, anomaly detection, and so on. Though having a broad range of applications, tensor factorization has not been popularly applied compared with matrix factorization that has been widely used in recommender systems, mainly due to the high computational cost and poor scalability of existing tensor factorization methods. Efficient and scalable tensor factorization is particularly challenging because real world tensor data are mostly sparse and massive. In this paper, we propose a novel distributed algorithm, called Lock-Free Tensor Factorization (LFTF), which significantly improves the efficiency and scalability of distributed tensor factorization by exploiting asynchronous execution in a re-formulated problem. Our experiments show that LFTF achieves much higher CPU and network throughput than existing methods, converges at least 17 times faster and scales to much larger datasets. Fan Yang 0091, Fanhua Shang, James Cheng, Yunjian Zhao, Ruihao Zhao |
Proc. VLDB Endow. | 6 |
| 2016 | A comparison of general-purpose distributed systems for data processingabstractGeneral-purpose distributed systems for data processing become popular in recent years due to the high demand from industry for big data analytics. However, there is a lack of comprehensive comparison among these systems and detailed analysis on their performance. In this paper, we conduct an extensive performance study on four state-of-the-art general-purpose distributed computing systems. Our results reveal useful insights on the design and implementation, which help the improvement of existing systems and the development of better new systems. James Cheng, Yunjian Zhao, Fan Yang 0091, Haipeng Chen 0002, Ruihao Zhao |
IEEE BigData | 3 |