EDBT 2026 Demo / reviewers in the wild / expert
Yongli Cheng
dblp:174/2042
· DBLP profile ↗
21ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-5250-9437ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 5 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graphago: Accelerating SSD-based Graph Processing via Activity-Aware Graph PreprocessingabstractSSD-based graph processing systems have emerged as a cost-effective solution for handling the ever-growing, large-scale graphs that exceed the memory capacity of a single machine. However, the mismatch between the large SSD access granularity (e.g., 4KB) and the small size of the graph vertex data leads to significant read amplification and low I/O efficiency. Despite existing works proposing techniques like dynamic active data gathering or reordering-based graph preprocessing to tackle this challenge, they inevitably cause problems such as expensive on-line computation overheads, inefficient graph traversal, and I/O imbalance, thus degrading the performance of graph processing. Xianghao Xu, Gongxuan Zhang, Yongli Cheng, Fang Wang 0001 |
SC | 4 |
| 2024 | A disk I/O optimized system for concurrent graph processing jobs
Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Dan Feng 0001, Peng Fang 0002 |
Frontiers Comput. Sci. | 4 |
| 2024 | An efficient SSSP algorithm on time-evolving graphs with prediction of computation results
Yongli Cheng, Chuanjie Huang, Hong Jiang 0001, Xianghao Xu, Fang Wang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2024 | TgStore: An Efficient Storage System for Large Time-Evolving GraphsabstractExisting graph systems focus mainly on the execution efficiency of the graph analysis tasks, often ignoring the importance and efficiency of time-evolving graph storage. However, to effectively mine the potential application values, an efficient storage system is important for time-evolving graphs whose storage requirement scales with the increasing number of snapshots. Storage cost and snapshot access speed are the two most important performance indicators for a time-evolving graph storage system, which are challenging for designers of such systems because they are conflicting goals. In this article, we address these challenges by proposing an efficient storage scheme for the large time-evolving graphs. We first design aSnapshot-level Data Deduplication (SLDD)strategy to eliminate the large number of repeated vertices and edges among the snapshots, and then aStructure-Changing Graph Representation (SCGR)to significantly improve the snapshot access speed. We implement an efficient time-evolving graph storage system, TgStore, based on this scheme to effectively store large-scale time-evolving graphs, aiming to efficiently support the time-evolving graph analysis tasks. Experimental results show that TgStore can obtain a high compression ratio of 43.03:1 when storing 100 snapshots of Twitter, while with an average snapshot access speedup of 16×. Efficient storage scheme enables TgStore to efficiently support time-evolving graph algorithms. For example, when executing the Pagerank algorithm on the time-evolving graph of Twitter, TgStore outperforms Graphone, a state-of-the-art time-evolving graph storage system, by 15.9× in algorithm execution speed and 1.45× in memory usage. Yongli Cheng, Hong Jiang 0001, Lingfang Zeng, Fang Wang 0001, Xianghao Xu, Yuhang Wu 0005 |
IEEE Trans. Big Data | 1 |
| 2023 | Deep Metric Multi-View Hashing for Multimedia RetrievalabstractLearning the hash representation of multi-view heterogeneous data is an important task in multimedia retrieval. However, existing methods fail to effectively fuse the multi-view features and utilize the metric information provided by the dissimilar samples, leading to limited retrieval precision. Current methods utilize weighted sum or concatenation to fuse the multi-view features. We argue that these fusion methods cannot capture the interaction among different views. Furthermore, these methods ignored the information provided by the dissimilar samples. We propose a novel deep metric multi-view hashing (DMMVH) method to address the mentioned problems. Extensive empirical evidence is presented to show that gate-based fusion is better than typical methods. We introduce deep metric learning to the multi-view hashing problems, which can utilize metric information of dissimilar samples. On the MIR-Flickr25K, MS COCO, and NUS-WIDE, our method outperforms the current state-of-the-art methods by a large margin (up to 15.28 mean Average Precision (mAP) improvement). Xiaohu Ruan, Yongli Cheng, Zhangmin Huang, Lingfang Zeng |
ICME | 3 |
| 2023 | FPC-Net: Learning to detect face forgery by adaptive feature fusion of patch correlation with CG-LossabstractAbstract With the rapid development of manipulation technologies, the generation of Deep Fake videos is more accessible than ever. As a result, face forgery detection becomes a challenging task, attracting a significant amount of attention from researchers worldwide. However, most previous work, consisting of convolutional neural networks (CNN), is not sufficiently discriminative and cannot fully utilise subtle clues and similar textures during the process of facial forgery detection. Moreover, these methods cannot simultaneously consider accuracy and time efficiency. To address such problems, we propose a novel framework named FPC‐Net to extract some meaningful and unnatural expressions in local regions. This framework utilises CNN, long short‐term memory (LSTM), channel groups loss (CG‐Loss) and adaptive feature fusion to detect face forgery videos. First, the proposed method exploits spatial features by CNN, and a channel‐wise attention mechanism is employed to separate channels. Specifically, with the help of channel groups loss, the channels are divided into two groups, each representing a specific class. Second, LSTM is applied to learn the correlation of spatial features. Finally, the correlation of features is mapped into other latent spaces. Through a lot of experiments, the results are that the detection speed of the proposed method reaches 420 FPS and the auc scores achieve best performance of 99.7%, 99.9%, 94.7%, and 82.0% on Raw Celeb‐DF, Raw Face Forensics++, F2F and NT datasets respectively. The experimental results demonstrate that the proposed framework has great time efficiency performance while improving the detection performance compared with other frame‐level methods in most cases. Lichao Su, Yongli Cheng |
IET Comput. Vis. | 4 |
| 2023 | LOSC: A locality-optimized subgraph construction scheme for out-of-core graph processing
Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Yu Hua 0001, Dan Feng 0001, Yongxuan Zhang |
J. Parallel Distributed Comput. | 4 |
| 2022 | GraphSD: A State and Dependency aware Out-of-Core Graph Processing SystemabstractIn recent years, system researchers have proposed many out-of-core graph processing systems to efficiently handle graphs that exceed the memory capacity of a single machine. Through disk-friendly graph data organizations and well-designed execution engines, existing out-of-core graph processing systems can maintain sequential locality on disk access and greatly reduce disk I/Os during processing. However, they have not fully explored the characteristics of graph data and algorithm execution to further reduce disk I/Os, leaving significant room for performance improvement. In this paper, we present a novel out-of-core graph processing system called GraphSD, which optimizes the I/O traffic by simultaneously capturing the state and dependency of graph data during computation. At the heart of GraphSD is a state- and dependency-aware update strategy that includes two adaptive update models, selective cross-iteration update (SCIU) and full cross-iteration update (FCIU). These two update models are dynamically triggered at runtime to enable active-vertex aware processing and cross-iteration vertex value computation, which avoid loading inactive edges and reduce disk I/Os in the future iterations. Moreover, an efficient sub-block based buffering scheme is proposed to further minimize I/O overheads. Our evaluation results show that GraphSD outperforms two state-of-the-art out-of-core graph processing systems HUS-Graph and Lumos by up to 2.7 × and 3.9 × respectively. Xianghao Xu, Hong Jiang 0001, Fang Wang 0001, Yongli Cheng, Peng Fang 0002 |
ICPP | 4 |
| 2021 | MatchMaker: Aspect-Based Sentiment Classification via Mutual Information
Yongli Cheng, Fang Wang 0001, Xianghao Xu, Wenxiong Wu |
ICONIP (2) | 2 |
| 2021 | GraphCP: An I/O-Efficient Concurrent Graph Processing FrameworkabstractBig data applications increasingly rely on the analysis of large graphs. In order to analyze and process the large graphs with high cost efficiency, researchers have developed a number of out-of-core graph processing systems in recent years based on just one commodity computer. On the other hand, with the rapidly growing need of analyzing graphs in the real-world, graph processing systems have to efficiently handle massive concurrent graph processing (CGP) jobs. Unfortunately, due to the inherent design for single graph processing job, existing out-of-core graph processing systems usually incur redundant data accesses and storage and severe competition of I/O bandwidth when handling the CGP jobs, thus leading to very long waiting time experienced by users for the computing results. In this paper, we propose an I/O-efficient out-of-core graph processing system, GraphCP, to support the processing of CGP jobs. GraphCP proposes a benefit-aware sharing execution model that shares the I/O access and processing of graph data among the CGP jobs and adaptively schedules the loading of graph data, which efficiently overcomes above challenges faced by existing out-of-core graph processing systems. In addition, GraphCP organizes the graph data with a Source-Sorted Sub-Block graph representation for better processing capacity and I/O access locality. Extensive evaluation results show that GraphCP is 10.3x and 4.6x faster than two state-of-the-art out-of-core graph processing systems GridGraph and GraphZ respectively, and 2.1x faster than a CGP-oriented graph processing system Seraph. Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Dan Feng 0001, Yongxuan Zhang, Peng Fang 0002 |
IWQoS | 4 |
| 2021 | CIC-PIM: Trading spare computing power for memory space in graph processing
Yongxuan Zhang, Hong Jiang 0001, Fang Wang 0001, Yu Hua 0001, Dan Feng 0001, Yongli Cheng, Yuchong Hu, Renzhi Xiao |
J. Parallel Distributed Comput. | 6 |
| 2020 | A Hybrid Update Strategy for I/O-Efficient Out-of-Core Graph ProcessingabstractIn recent years, a number of out-of-core graph processing systems have been proposed to process graphs with billions of edges on just one commodity computer, due to their high cost efficiency. To obtain a better performance, these systems adopt a full I/O model that scans all edges during the computation to avoid the inefficiency of random I/Os. Although this model ensures good I/O access locality, it leads to a large number of useless edges to be loaded when running graph algorithms that only access a small portion of edges in each iteration. An intuitive method to solve this I/O inefficiency problem is the on-demand I/O model that only accesses the active edges. However, this method only works well for the graph algorithms with very few active edges, since the I/O cost will grow rapidly as the number of active edges increases due to the increasing amount of random I/Os. In this article, we present HUS-Graph, an efficient out-of-core graph processing system to address the above I/O issues and achieve a good balance between I/O traffic and I/O access locality. HUS-Graph adopts a hybrid update strategy including two update models, Row-oriented Push (ROP) and Column-oriented Pull (COP). It supports switching between ROP and COP adaptively, for the graph algorithms that have different computation and I/O features. For traversal-based algorithms, HUS-Graph also provides an immediate propagation-based vertex update scheme to accelerate the vertex state propagation and convergence speed. Furthermore, HUS-Graph adopts a locality-optimized dual-block representation to organize graph data and an I/O-based performance prediction method to enable the system to dynamically select the optimal update model between ROP and COP. To save the disk space and further reduce I/O traffic, HUS-Graph implements a space-efficient storage format by combining several graph compression methods. Extensive experimental results show that HUS-Graph outperforms two existing out-of-core systems GraphChi and GridGraph by 1.2x-52.8x. Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Dan Feng 0001, Yongxuan Zhang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | LOSC: efficient out-of-core graph processing with locality-optimized subgraph constructionabstractBig data applications increasingly rely on the analysis of large graphs. In recent years, a number of out-of-core graph processing systems have been proposed to process graphs with billions of edges on just one commodity computer, by efficiently using the secondary storage (e.g., hard disk, SSD). On the other hand, the vertex-centric computing model is extensively used in graph processing thanks to its good applicability and expressiveness. Unfortunately, when implementing vertex-centric model for out-of-core graph processing, the large number of random memory accesses required to construct subgraphs lead to a serious performance bottleneck that substantially weakens cache access locality and thus leads to very long waiting time experienced by users for the computing results. In this paper, we propose an efficient out-of-core graph processing system, LOSC, to substantially reduce the overhead of subgraph construction without sacrificing the underlying vertex-centric computing model. LOSC proposes a locality-optimized subgraph construction scheme that significantly improves the in-memory data access locality of the subgraph construction phase. Furthermore, LOSC adopts a compact edge storage format and a lightweight replication of vertices to reduce I/O traffic and improve computation efficiency. Extensive evaluation results show that LOSC is respectively 6.9x and 3.5x faster than GraphChi and GridGraph, two state-of-the-art out-of-core systems. Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Yu Hua 0001, Dan Feng 0001, Yongxuan Zhang |
IWQoS | 4 |
| 2019 | Using High-Bandwidth Networks Efficiently for Fast Graph ComputationabstractNowadays, high-bandwidth networks are more easily accessible than ever before. However, existing distributed graph-processing frameworks, such as GPS, fail to efficiently utilize the additional bandwidth capacity in these networks for higher performance, due to their inefficient computation and communication models, leading to very long waiting times experienced by users for the graph-computing results. The root cause lies in the fact that the computation and communication models of these frameworks generate, send and receive messages so slowly that only a small fraction of the available network bandwidth is utilized. In this paper, we propose a high-performance distributed graph-processing framework, called BlitzG, to address this problem. This framework fully exploits the available network bandwidth capacity for fast graph processing. Our approach aims at significant reduction in (i) the computation workload of each vertex for fast message generation by using a new slimmed-down vertex-centric computation model and (ii) the average message overhead for fast message delivery by designing a light-weight message-centric communication model. Evaluation on a 40Gbps Ethernet, driven by real-world graph datasets, shows that BlitzG outperforms GPS by up to 27x with an average of 20.7x. Yongli Cheng, Hong Jiang 0001, Fang Wang 0001, Yu Hua 0001, Dan Feng 0001, Wenzhong Guo, Yunxiang Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | HUS-Graph: I/O-Efficient Out-of-Core Graph Processing with Hybrid Update StrategyabstractIn recent years, a number of out-of-core graph processing systems have been proposed to process graphs with billions of edges on just one commodity computer, due to their high cost efficiency. To obtain the better performance, these systems adopt a full I/O model that accesses all edges during the computation to avoid the ineffectiveness of random I/Os. Although this model ensures good I/O access locality, it loads a large number of useless edges when running graph algorithms that only require a small portion of edges in each iteration. A natural method to solve this problem is the on-demand I/O model that only accesses the active edges. However, this method only works well for the graph algorithms with very few active edges, since the I/O cost will grow rapidly as the number of active edges increases due to larger amount of random I/Os. Xianghao Xu, Fang Wang 0001, Hong Jiang 0001, Yongli Cheng, Dan Feng 0001, Yongxuan Zhang |
ICPP | 4 |
| 2018 | A communication-reduced and computation-balanced framework for fast graph computation
Yongli Cheng, Fang Wang 0001, Hong Jiang 0001, Yu Hua 0001, Dan Feng 0001, Lingling Zhang 0006 |
Frontiers Comput. Sci. | 1 |
| 2018 | A highly cost-effective task scheduling strategy for very large graph computation
Yongli Cheng, Fang Wang 0001, Hong Jiang 0001, Yu Hua 0001, Dan Feng 0001, Yunxiang Wu, Tingwei Zhu, Wenzhong Guo |
Future Gener. Comput. Syst. | 1 |
| 2017 | BlitzG: Exploiting high-bandwidth networks for fast graph processingabstractNowadays, high-bandwidth networks are easily accessible in data centers. However, existing distributed graph-processing frameworks fail to efficiently utilize the additional bandwidth capacity in these networks for higher performance, due to their inefficient computation and communication models, leading to very long waiting times experienced by users for the graph-computing results. The root cause lies in the fact that the computation and communication models of these frameworks generate, send and receive messages so slowly that only a small fraction of the available network bandwidth is utilized. In this paper, we propose a high-performance distributed graph-processing framework, called BlitzG, to address this problem. This framework fully exploits the available network bandwidth capacity for fast graph processing. Our approach aims at significant reduction in (i) the computation workload of each vertex for fast message generation by using a new slimmed-down vertex-centric computation model and (ii) the average message overhead for fast message delivery by designing a lightweight message-centric communication model. Evaluation on a 40Gbps Ethernet, driven by real-world graph datasets, shows that BlitzG outperforms the state-of-the-art distributed graph-processing frameworks by up to 27x with an average of 20.7x. Yongli Cheng, Hong Jiang 0001, Fang Wang 0001, Yu Hua 0001, Dan Feng 0001 |
INFOCOM | 1 |
| 2017 | Efficient Anonymous Communication in SDN-Based Data Center NetworksabstractWith the rapid growth of application migration, the anonymity in data center networks becomes important in breaking attack chains and guaranteeing user privacy. However, existing anonymity systems are designed for the Internet environment, which suffer from high computational and network resource consumption and deliver low performance, thus failing to be directly deployed in data centers. In order to address this problem, this paper proposes an efficient and easily deployed anonymity scheme for software defined networking-based data centers, called mimic channel (MIC). The main idea behind MIC is to conceal the communication participants by modifying the source/destination addresses, such as media access control (MAC) and Internet protocol (IP) address at switch nodes, so as to achieve anonymity. Compared with the traditional overlay-based approaches, our in-network scheme has shorter transmission paths and less intermediate operations, thus achieving higher performance with less overhead. We also propose a collision avoidance mechanism to ensure the correctness of routing, and three mechanisms to enhance the traffic-analysis resistance. To enhance the practicality, we further propose solutions to enable MIC co-existing with some MIC-incompatible systems, such as packet analysis systems, intrusion detection systems, and firewall systems. Our security analysis demonstrates that MIC ensures unlinkability and improves traffic-analysis resistance. Our experiments show that MIC has extremely low overhead compared with the base-line transmission control protocol (TCP) (or secure sockets layer (SSL)), e.g., less than 1% overhead in terms of throughput. Experiments on MIC-based distributed file system show the applicability and efficiency of MIC. Tingwei Zhu, Dan Feng 0001, Fang Wang 0001, Yu Hua 0001, Qingyu Shi 0001, Yongli Cheng |
IEEE/ACM Trans. Netw. | 7 |
| 2016 | DD-Graph: A Highly Cost-Effective Distributed Disk-based Graph-Processing FrameworkabstractExisting distributed graph-processing frameworks, e.g.,GPS, Pregel and Giraph, handle large-scale graphs in the memory of clusters built of commodity compute nodes for better scalability and performance. While capable of scaling out according to the size of graphs up to thousands of compute nodes, for graphs beyond a certain size, these frameworks usually require the investments of machines that are either beyond the financial capability of or unprofitable for most small and medium-sized organizations. At the other end of the spectrum of graph-processing frameworks research, the single-node disk-based graph-processing frameworks, e.g., GraphChi, handle large-scale graphs on one commodity computer, leading to high efficiency in the use of hardware but at the cost of low user performance and limited scalability. Motivated by this dichotomy, in this paper we propose a distributed disk-based graph-processing framework, called DD-Graph, that can process super-large graphs on a small cluster while achieving the high performance of existing distributed in-memory graph-processing frameworks. Yongli Cheng, Fang Wang 0001, Hong Jiang 0001, Yu Hua 0001, Dan Feng 0001, XiuNeng Wang |
HPDC | 1 |
| 2016 | LCC-Graph: A high-performance graph-processing framework with low communication costsabstractWith the rapid growth of data, communication overhead has become an important concern in applications of data centers and cloud computing. However, existing distributed graph-processing frameworks routinely suffer from high communication costs, leading to very long waiting times experienced by users for the graph-computing results. In order to address this problem, we propose a new computation model with low communication costs, called LCC-BSP. We use this model to design and implement a high-performance distributed graph-processing framework called LCC-Graph. This framework eliminates the high communication costs in existing distributed graph-processing frameworks. Moreover, LCC-Graph also minimizes the computation workload of each vertex, significantly reducing the computation time for each superstep. Evaluation of LCC-Graph on a 32-node cluster, driven by real-world graph datasets, shows that it significantly outperforms existing distributed graph-processing frameworks in terms of runtime, particularly when the system is supported by a high-bandwidth network. For example, LCC-Graph achieves an order of magnitude performance improvement over GPS and GraphLab. Yongli Cheng, Fang Wang 0001, Hong Jiang 0001, Yu Hua 0001, Dan Feng 0001, XiuNeng Wang |
IWQoS | 1 |