Chuitian Rong

dblp:02/8098 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
5since 2021 · last 2026
0000-0003-2949-3892ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 5 first-author · 1 since 2021Systems, architecture and hardware · 6 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Metric-aware multi-objective deep reinforcement learning for database knob tuning
Chuitian Rong, Jitai Li, Chunbin Lin, Fang Du, Wei Lu 0015
Future Gener. Comput. Syst.1
2025 The Lightweight Optimization Techniques for Knobs Tuning of Database System Using Evolution Strategy
Jitai Li, Chuitian Rong, Yukun Cui, Tiegang Chen
IEEE Big Data2
2023 Distributed structural clustering on large graph
abstract
Summary Graph clustering is a primitive operation for graph data mining. It plays an important role to reveal community clusters, hubs, and outliers in complex networks. There are several graph clustering algorithms have been proposed based on the well‐studied SCAN algorithm in recent years. However, SCAN and its improved sequential variants are prohibitively slow due to their iterative computations. The parallel variants are focusing on improving the efficiency of graph clustering by utilizing multi‐cores computer architectures on single computing node with complex optimization techniques. Therefore, SCAN and its variants are not suitable for processing very large graphs due to the limitations of memory size and storage volume on a single node. In this article, we proposed a distributed parallel structural clustering algorithm using MapReduce. In order to improve the efficiency further, we proposed optimization techniques including partition based clustering and simplified combination with labels to accelerate the operations. We conducted extensive experiments on real world datasets. The experimental results showed our algorithm is high efficiency and scales well under different settings.
Chuitian Rong
Concurr. Comput. Pract. Exp.1
2023 Multivariate time series data imputation using attention-based mechanism
Jingqi Zhao, Chuitian Rong, Chunbin Lin, Xin Dang
Neurocomputing2
2022 Distributed Exact Structural Clustering on Large Graph
abstract
Graph clustering is an important technique to detect community clusters in complex networks. SCAN (Structural Clustering Algorithm for Networks) is a well-studied graph clustering algorithm that has been widely applied over the years. However, the processing time cost of sequential SCAN and its variants cannot be tolerable on large graphs. The existing parallel variants of SCAN are focusing on fully utilizing the computing capacity of multi-core computer architectures and inventing sophisticated optimization techniques on single computing node. As the objects and their relationships in cyberspace are varying over time, the scale of graph data is increasing with high rate. The graph clustering algorithms on single node are facing challenges from limited computing resources, such as computing performance, memory size and storage volume. The distributed processing algorithm is called for processing large graphs. This work presents a distributed structural graph clustering algorithm using Spark. Furthermore, the edge pruning technique and adaptive checking are optimized to improve clustering efficiency. And the label propagation clustering is simplified to reduce the communication cost in the distributed clustering iterations. It also conduct extensive experiments on real-world datasets to testify the efficiency and scalability of the distributed algorithm. Experimental results show that efficient clustering performance can be achieved and it scales well under different settings.
Chuitian Rong
ICPADS2
2020 Motif Discovery Using Similarity-Constraints Deep Neural Networks
Chuitian Rong, Ziliang Chen 0002, Chunbin Lin
DASFAA (1)1
2020 Boundary-connection deletion strategy based method for community detection in complex networks
Chuitian Rong, Qingshuang Yao
Appl. Intell.2
2020 Parallel time series join using spark
abstract
Summary A time series is a sequence of data points in successive temporal order. Time series data is produced in many applications scenarios, and the techniques for its analysis have generated substantial interest. Time series join is a primitive operation that retrieves all pairs of correlated subsequences from two given time series. As the Pearson correlation coefficient, a measure of the correlation between two variables, has multiple beneficial mathematical properties, for example, the fact that it is invariant with respect to scale and offset, it is used to measure the correlation between two time series. Considering the need to analyze big time series data, we focus on the study of scalable and distributed techniques to process massive data sets. Specifically, we propose a parallel approach to perform time series joins using Spark, a popular analytics engine for large‐scale data processing. Our solution builds on (1) a fast method to compute the fast Fourier transform on the times series to calculate the correlation between two time series, (2) a lossless partition method to divide the time series into multiple subsequences and enable a parallel and correct computation of the join result, and (3) optimization techniques to avoid redundant computations. We performed extensive tests and showed that the proposed approach is efficient and scalable across different data sets and test configurations.
Chuitian Rong, Yasin N. Silva
Concurr. Comput. Pract. Exp.1
2019 Similarity Grouping in Big Data Systems
Yasin N. Silva, Manuel Sandoval, Diana Prado, Xavier Wallace, Chuitian Rong
SISAP5
2019 Similarity joins for high-dimensional data using Spark
abstract
Summary Similarity join on high‐dimensional data is a primitive operation. It is used to find all data pairs that with distance no more than ϵ from the given data set according to a specific distance measure. As the data set scale and dimension increase, computation cost increases vastly. Hadoop and Spark have become the popular platforms for big‐data analysis. Because Spark has native advantages in iterative computations, we adopted it as our platform to perform similarity joins on high‐dimensional data sets. In order to resolve problems such as data imbalance, data duplication, and redundant computation of existing works, we have proposed a new algorithm based on Symbolic aggregation and vertical decomposition. We first conduct dimension‐reduction using symbolic aggregation method. Then, we applied vertical partition operation on processed data. The join operations are performed on each vertical partition in parallel manner and the proposed new filters are utilized to prune false positives in early stage. Finally, the partial results generated from each partition will be aggregated and verified to get final results. Our proposed algorithm can significantly improve the efficiency of similarity joins on high‐dimensional data. In order to verify the efficiency and scalability of our methods, we implemented it using MapReduce and Spark. We compared our methods with existing works on public data sets, and the experimental results showed that the new methods were more efficient and scalable under different running environments.
Chuitian Rong, Xiaohai Cheng, Ziliang Chen 0002, Na Huo
Concurr. Comput. Pract. Exp.1
2017 Fast and Scalable Distributed Set Similarity Joins for Big Data Analytics
abstract
Set similarity join is an essential operation in big data analytics, e.g., data integration and data cleaning, that finds similar pairs from two collections of sets. To cope with the increasing scale of the data, distributed algorithms are called for to support large-scale set similarity joins. Multiple techniques have been proposed to perform similarity joins using MapReduce in recent years. These techniques, however, usually produce huge amounts of duplicates in order to perform parallel processing successfully as MapReduce is a shared-nothing framework. The large number of duplicates incurs on both large shuffle cost and unnecessary computation cost, which significantly decrease the performance. Moreover, these approaches do not provide a load balancing guarantee, which results in a skewness problem and negatively affects the scalability properties of these techniques. To address these problems, in this paper, we propose a duplicatefree framework, called FS-Join, to perform set similarity joins efficiently by utilizing an innovative vertical partitioning technique. FS-Join employs three powerful filtering methods to prune dissimilar string pairs without computing their similarity scores. To further improve the performance and scalability, FS-Join integrates horizontal partitioning. Experimental results on three real datasets show that FS-Join outperforms the state-of-theart methods by one order of magnitude on average, which demonstrates the good scalability and performance qualities of the proposed technique.
Chuitian Rong, Chunbin Lin, Yasin N. Silva, Jianguo Wang 0001, Wei Lu 0015, Xiaoyong Du 0001
ICDE1
2017 String similarity join with different similarity thresholds based on novel indexing techniques
Chuitian Rong, Yasin N. Silva, Chunqing Li
Frontiers Comput. Sci.1
2016 An Experimental Survey of MapReduce-Based Similarity Joins
Yasin N. Silva, Jason M. Reed, Kyle Brown, Adelbert Wadsworth, Chuitian Rong
SISAP5
2015 String Similarity Join with Different Thresholds
abstract
String similarity join is an essential operation of many applications that need to find all similar string pairs from given two collections. The existing approaches are using the uniform and predefined similarity thresholds. While in real applications, regarding that the longer string pairs typically tolerate many more typos, it is necessary to apply variable thresholds to different strings instead of a constant one. Therefore, we proposed a solution for string similarity joins with different similarity thresholds in one procedure. In order to support different similarity thresholds, we devised the similarity aware index and index probing technique. To our best knowledge, it is the first work to address the problem. Experimental results on real-world datasets show that our solution can tackle with different similarity thresholds efficiently.
Chuitian Rong, Xiangling Zhang
KSEM1
2013 Efficient and exact duplicate detection on cloud
abstract
SUMMARY As the recent proliferation of social networks, mobile applications, and online services increased the rate of data gathering, to find near‐duplicate records efficiently has become a challenging issue. Related works on this problem mainly aim to propose efficient approaches on a single machine. However, when processing large‐scale dataset, the performance to identify duplicates is still far from satisfactory. In this paper, we try to handle the problem of duplicate detection applying MapReduce. We argue that the performance of utilizing MapReduce to detect duplicates mainly depends on the number of candidate record pairs and intermediate result size, which is related to the shuffle cost among different nodes in cluster. In this paper, we proposed a new signature scheme with new pruning strategies to minimize the number of candidate pairs and intermediate result size. The proposed solution is an exact one, which assures none duplicate record pair can be lost. The experimental results over both real and synthetic datasets demonstrate that our proposed signature‐based method is efficient and scalable. Copyright © 2012 John Wiley & Sons, Ltd.
Chuitian Rong, Wei Lu 0015, Xiaoyong Du 0001, Xiao Zhang 0001
Concurr. Comput. Pract. Exp.1
2013 Efficient and Scalable Processing of String Similarity Join
abstract
The string similarity join is a basic operation of many applications that need to find all string pairs from a collection given a similarity function and a user-specified threshold. Recently, there has been considerable interest in designing new algorithms with the assistant of an inverted index to support efficient string similarity joins. These algorithms typically adopt a two-step filter-and-refine approach in identifying similar string pairs: 1) generating candidate pairs by traversing the inverted index; and 2) verifying the candidate pairs by computing the similarity. However, these algorithms either suffer from poor filtering power (which results in high verification cost), or incur too much computational cost to guarantee the filtering power. In this paper, we propose a multiple prefix filtering method based on different global orderings such that the number of candidate pairs can be reduced significantly. We also propose a parallel extension of the algorithm that is efficient and scalable in a MapReduce framework. We conduct extensive experiments on both centralized and Hadoop systems using both real and synthetic data sets, and the results show that our proposed approach outperforms existing approaches in both efficiency and scalability.
Chuitian Rong, Wei Lu 0015, Xiaoli Wang 0002, Xiaoyong Du 0001, Yueguo Chen, Anthony K. H. Tung
IEEE Trans. Knowl. Data Eng.1
2011 Efficient Duplicate Detection on Cloud Using a New Signature Scheme
Chuitian Rong, Wei Lu 0015, Xiaoyong Du 0001, Xiao Zhang 0001
WAIM1
2010 Efficient Common Items Extraction from Multiple Sorted Lists
abstract
Given a set of lists, where items of each list are sorted by the ascending order of their values, the objective of this paper is to figure out the common items that appear in all of the lists efficiently. This problem is sometimes known as common items extraction from sorted lists. To solve this problem, one common approach is to scan all items of all lists sequentially in parallel until one of the lists is exhausted. However, we observe that if the overlap of items across all lists is not high, such sequential access approach can be significantly improved. In this paper, we propose two algorithms, MergeSkip and MergeESkip, to solve this problem by taking the idea of skipping as many items of lists as possible. As a result, a large number of comparisons among items can be saved, and hence the efficiency can be improved. We conduct extensive analysis of our proposed algorithms on one real dataset and two synthetic datasets with different data distributions. We report all our findings in this paper.
Wei Lu 0015, Chuitian Rong, Jinchuan Chen, Xiaoyong Du 0001, Gabriel Pui Cheong Fung, Xiaofang Zhou 0001
APWeb2