VLDB 2026 Research / reviewers in the wild / expert
Zujie Ren
dblp:83/7247
· DBLP profile ↗
21ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-9985-9805ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward real-world Table Agents: capabilities, workflows, and design principles for LLM-based table intelligence
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang 0001, Lingxin Wang, Lihua Yu, Zujie Ren, Gang Chen 0001, Junbo Zhao 0002 |
World Wide Web (WWW) | 7 |
| 2023 | Textile pattern recommendations with convolutional neural networks and autoencoderabstractAbstract Textile pattern design is a time‐consuming and tedious work. Mihui, our ongoing developing system, employs deep‐learning techniques to automatically generate huge volumes of patterns with the help of human guidance. However, trained as a black box, Mihui cannot provide customized service for each individual designer who shows unique aesthetics preferences. In this article, we introduce the recommendation module of Mihui. The module forwards all generated pattern images to a deep encoding network, where images are mapped into 128‐dimension vectors. For each user of Mihui, we create a profile by his/her purchased or downloaded history. A novel encoder network is proposed to learn a personal taste vector for each user, based on which, we recommend new patterns to him/her. Our records in Mihui show that the recommendation module effectively improve users' experience on Mihui. Kuang Mao, Sai Wu, Jiajia He, Haichao Huang, Yanlong Yin, Zujie Ren |
Concurr. Comput. Pract. Exp. | 6 |
| 2023 | A fast approximate method for k-edge connected component detection in graphs with high accuracy
Ting Yu 0004, Mengchi Liu, Zujie Ren, Ji Zhang 0001 |
Inf. Sci. | 3 |
| 2022 | A Parallel Framework for Streaming Graphs ComputingabstractStreaming computation for large graphs on parallel systems faces challenges in task decomposition, data skew, and resource scheduling. In this work, we propose a general parallel streaming framework for the node-centered graph algorithms to improve the computation efficiency. We construct the parallel procedure of the incremental maximal clique enumeration (IMCE) task and accelerate the incremental Candidate Map Constructor (CMC) algorithm through the framework for large-scale streaming graphs. Experimental results on three large real-world graphs show the framework’s positive effect on the algorithm’s execution time. Ting Jiang 0006, Ting Yu 0004, Zexian Hong, Zujie Ren, Ji Zhang 0001 |
IEEE Big Data | 4 |
| 2022 | IDGMS: a One-Stop Graph Mining System for Infectious DiseasesabstractData mining in infectious disease pandemic scenarios is a complex giant task involving data from various fields and requirements of real-time and dynamic. In this paper, we propose a graph mining system for the infectious disease pandemic, IDGMS, with one-stop, dynamic, and interactive characteristics. The system has been applied to solve problems from three view scales and performs well. The system is constructed as a loose coupling structure at the front and back ends and can be extended to more graph mining issues. To the best of our knowledge, we are the first graph system especially targeting data mining of infectious diseases. Zenghui Xu, Ting Yu 0004, Xingyun Hong, Mingzhang Li, Yang Zhang 0042, Zujie Ren, Ji Zhang 0001 |
IEEE Big Data | 6 |
| 2022 | Knowledge Tracing Based on Gated Heterogeneous Graph Convolutional NetworksabstractThe advancement of science and technology provides the possibility of personalized intelligent education. Representation learning of students’ behavior data is challenging because whether time sequences and interactive behaviors or the correlation between knowledge points and students carrying important information. Some researchers propose knowledge tracing to provide ideas for solving this dilemma. However, existing knowledge tracing methods are divided into machine learning and deep learning. Machine learning-based methods require manual feature extraction and a large amount of prior knowledge. Although deep learning-based methods can automatically extract features, most methods either only use the time series information of the data, or use the association between knowledge points. All the methods ignore the association between knowledge points and students. To fill this gap, we propose a Gated Heterogeneous Graph Convolutional Network (GHGCN) model. We utilize the encoder-decoder framework to predict student performance using the representations of nodes, which is learned from heterogeneous convolutional networks and gate recurrent unit. To validate the effectiveness of the proposed GHGCN model, we conduct the experiments on three public datasets: Simulated Data, Assistments 2009, and Assistments 2015. The results indicate that our method can achieve better performance compared with state-of-the-art algorithms. Yang Zhang 0042, Zhen Wang 0037, Ting Yu 0004, Mingming Lu, Zujie Ren, Ji Zhang 0001 |
IEEE Big Data | 5 |
| 2022 | Characterizing Co-Located Workloads in Alibaba Cloud DatacentersabstractWorkload characteristics are vital for both data center operation and job scheduling in co-located data centers, where online services and batch jobs are deployed on the same production cluster. In this article, a comprehensive analysis is conducted on Alibaba's cluster-trace-v2018 of a production cluster of 4034 machines. The findings and insights are the following: (1) The workload on the production cluster poses a daily cyclical fluctuation, in terms of CPU and disk I/O utilization, and the memory system has become the performance bottleneck of a co-located cluster. (2) Batch jobs including their tasks and derived instances can be approximated as Zipf distribution. However, for all batch jobs with directed acyclic graph dependency, they suffer from co-location with online services since the online services are highly prioritized. (3) The resource usages of containers have similar cyclical fluctuation consistent with the whole cluster, while their memory usages remain approximately constant. (4) The number of batch jobs co-located with online services is dependent on the mispredictions per kilo instructions of online services. In order to guarantee the QoS of online services, when the MPKI of online services rises, the number of batch jobs to be co-located on the same machine should decrease. Congfeng Jiang, Yitao Qiu, Weisong Shi, Zhefeng Ge, Shenglei Chen, Christophe Cérin, Zujie Ren, Guoyao Xu, Jiangbin Lin |
IEEE Trans. Cloud Comput. | 8 |
| 2021 | Improving Irregularly Sampled Time Series Learning with Time-Aware Dual-Attention Memory-Augmented NetworksabstractIrregularly, asynchronously and sparsely sampled multivariate time series (IASS-MTS) are characterized by sparse non-uniform time intervals between successive observations and different sampling rates amongst series. Those properties pose substantial challenges to mainstream machine learning models for learning complicated relations within and across IASS-MTS. This is because that most of the models assume that the time series in question are even, complete (fixed-dimensional features) and synchronous. To address these challenges, we present a novel time-aware Dual-Attention and Memory-Augmented Network (DAMA-Net). The proposed model can leverage both time irregularity, multi-sampling rates and global temporal patterns information inherent in IASS-MTS so as to learn more effective representations for improving prediction performance. Comprehensive experiments on real datasets show that the DAMA-Net outperforms the state-of-the-art methods in multivariate time series classification task. Zhen Wang 0037, Yang Zhang 0042, Ai Jiang, Ji Zhang 0001, Zhao Li 0007, Jun Gao 0003, Ke Li 0044, Chenhao Lu, Zujie Ren |
CIKM | 9 |
| 2020 | Recommendation on Heterogeneous Information Network with Type-Sensitive Sampling
Jinze Bai, Zhao Li 0007, Donghui Ding, Pengrui Hui, Jun Gao 0003, Ji Zhang 0001, Zujie Ren |
DASFAA (3) | 9 |
| 2019 | Priority-Based Optimization of I/O Isolation for Hybrid Deployed Services
Youhuizi Li, Li Zhou 0008, Zujie Ren, Jian Wan 0001 |
CollaborateCom | 4 |
| 2018 | Towards Building a Scalable Data Analytics System on Clouds: An Early Experience on AliCloudabstractWith the development of big data, big data processing systems, such as Hadoop and Spark, are widely used to handle large-scale data. To avoid the complexity and expensiveness of building a self-owned big data processing system, cloud providers tend to deploy big data processing tools as cloud services. Typical examples include Amazon EMR, Azure HDInsight and AliCloud E-MapReduce. However, how to build a cost-efficient system and scale the system is still challenging. In this paper, we have conducted a case study on AliCloud E-MapReduce, and analyzed the system performance upon local and remote file systems. We compared the scalability of Hadoop and Spark by using scaleout and scale-up strategies respectively. Based on the analysis results, we derive several observations and implications, which will contribute to guide the performance optimization. Congfeng Jiang, Zujie Ren, Youhuizi Li, Jian Wan 0001, Jiangbin Lin |
IEEE CLOUD | 3 |
| 2018 | How Good is Query Optimizer in Spark?
Zujie Ren, Na Yun, Youhuizi Li, Jian Wan 0001, Lihua Yu, Xinxin Fan |
CollaborateCom | 1 |
| 2018 | Characterizing the Effectiveness of Query Optimizer in SparkabstractIn the big data community, Spark has been widely used for processing interactive queries. Spark employs a query optimizer, called Catalyst, to provides a set of optimization rules and supports Cost-Based Optimization (CBO). In this paper, we investigated the effectiveness of the optimization rules and cost-based optimization in Catalyst. We conducted comprehensive validation experiments by varying the data volume and cluster scale, and found that the execution time of most TPC-H queries were reduced slightly even when query optimizations are applied. We derived some interesting observations on Catalyst, which can help the community better understand and improve the query optimizer of Spark in future. Zujie Ren, Na Yun, Weisong Shi, Youhuizi Li, Jian Wan 0001, Lihua Yu, Xinxin Fan |
SERVICES | 1 |
| 2017 | Quantifying the Isolation Characteristics in Container Environments
Yusen Wu 0001, Zujie Ren, Weisong Shi, Jian Wan 0001 |
NPC | 3 |
| 2017 | Realistic and Scalable Benchmarking Cloud File Systems: Practices and Lessons from AliCloudabstractThe past decade has witnessed the rapid boom of cloud computing. Many public cloud infrastructures have been implemented and serve millions of tenants. Cloud file systems, which take charge of petabyte-scale data storage, play a crucial role in the performance of cloud infrastructures. Typical cloud file systems, including GFS, HDFS and Ceph, have attracted notable research efforts for performance evaluation and optimization. However, due to the heterogeneity and complexity of I/O workload characteristics in cloud environments, it is still challenging to conduct an accurate and efficient performance evaluation. To address this problem, we collected a two-week I/O workload trace from a 2,500-node production cluster in AliCloud, which is one of the largest cloud providers in Asia. Using the AliCloud trace, we characterized the I/O workload and data distribution, and compared two cloud services in multiple perspectives, including the request arrival pattern, request size, data population and so on. A list of observations and implications were derived and applied to help design a cloud file system benchmarking suite, called Porcupine. Porcupine aims to deploy a scalable and efficient performance evaluation on cloud file systems using realistic I/O workloads. We conducted a group of validation experiments, which demonstrated that Porcupine can achieve high accuracy and scalability. This paper provides our experiences and lessons in generating I/O workloads and deploying performance tests on cloud file systems, which we believe will be insightful to the cloud computing community in general. Zujie Ren, Weisong Shi, Jian Wan 0001, Jiangbin Lin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | iGen: A Realistic Request Generator for Cloud File Systems BenchmarkingabstractBenchmarking is a traditional approach for system performance evaluation and optimization. Over the past decades, a variety of file systems, e.g., GFS, HDFS and Ceph, have been designed and implemented, serving as the key components in cloud infrastructures. With the mature of those cloud file systems, the demands for performance evaluation and comparison are also rising. However, due to the complexity and heterogeneity of I/O workloads in cloud infrastructures, it is still challenging to generate realistic I/O workloads. System developers often use traditional file system benchmarks and make inaccurate assumptions on workload generation, yielding to misleading results. To address this problem, we investigate the characteristics of I/O requests in a production cloud infrastructure at Alibaba Cloud Computing, which is one of the biggest cloud providers in Asia. We proposed a flexible framework iGen to mimic I/O request arrivals. One of the salient features of the iGen is that the request arrival process is modeled by three statistics properties, request arrival rate, inter-arrival time distribution, and request periodicity. According to these properties, the iGen can determine the sequence of requests and the inter-arrival time between two subsequent requests. We use the iGen to emulate a real workload that collected from Alibaba cloud platform. Experimental results show that high accuracy and flexibility of the iGen. Zujie Ren, Weisong Shi, Jiangbin Lin |
CLOUD | 1 |
| 2014 | Workload Analysis, Implications, and Optimization on a Production Hadoop Cluster: A Case Study on TaobaoabstractUnderstanding the characteristics of MapReduce workloads in a Hadoop cluster is the key to making optimal configuration decisions and improving the system efficiency and throughput. However, workload analysis on a Hadoop cluster, particularly in a large-scale e-commerce production environment, has not been well studied yet. In this paper, we performed a comprehensive workload analysis using the trace collected from a 2000-node Hadoop cluster at Taobao, which is the biggest online e-commerce enterprise in Asia, ranked 10th in the world as reported by Alexa. The results of the workload analysis are representative and generally consistent with the data warehouses for e-commerce web sites, which can help researchers and engineers understand the workload characteristics of Hadoop in their production environments. Based on the observations and implications derived from the trace, we designed a workload generator Ankus, to expedite the performance evaluation and debugging of new mechanisms. Ankus supports synthesizing an e-commerce style MapReduce workload at a low cost. Furthermore, we proposed and implemented a job scheduling algorithm, Fair4S , which is designed to be biased towards small jobs. Small jobs account for the majority of the workload, and most of them require instant and interactive responses, which is an important phenomenon at production Hadoop systems. The inefficiency of Hadoop fair scheduler for handling small jobs motivates us to design the Fair4S, which introduces pool weights and extends job priorities to guarantee the rapid responses for small jobs. Experimental evaluation verified that the Fair4S accelerates the average waiting times of small jobs by a factor of 7 compared with the fair scheduler. Zujie Ren, Jian Wan 0001, Weisong Shi, Xianghua Xu |
IEEE Trans. Serv. Comput. | 1 |
| 2012 | Dual-JT: Toward the high availability of JobTracker in HadoopabstractMapReduce is a state-of-the-art computation paradigm that is becoming widely used for processing large-scale datasets. Hadoop is an open-source implementation of MapReduce and follows a masterCslave architecture. This architecture makes Hadoop suffer from a single point of failure in the JobTracker. In this paper, we design a solution to resolve the single point of failure of the Job Tracker and then enhance its availability. In this solution, a standby Job Tracker is introduced to act as a hot backup node of the active Job Tracker. The standby Job Tracker synchronizes the job execution process with the active Job Tracker by collecting and parsing the job log. If the active Job Tracker fails, the standby Job Tracker can take over quickly. This solution is implemented in Hadoop 0.20.x. Extensive experiments illustrate that this solution effectively enhances the availability of Job Tracker. A big production cluster in a large e-Commerce company has adopted this solution, which avoids interrupting job submission and execution when the Job Tracker fails or restarts. Jian Wan 0001, Minggang Liu, Xixiang Hu, Zujie Ren, Weisong Shi |
CloudCom | 4 |
| 2011 | PISA: A framework for integrating uncooperative peers into P2P-based federated search
Gang Chen 0001, Zujie Ren, Lidan Shou, Ke Chen 0005, Yijun Bei |
Comput. Commun. | 2 |
| 2010 | HAPS: Supporting Effective and Efficient Full-Text P2P Search with Peer Dynamics
Zujie Ren, Ke Chen 0005, Lidan Shou, Gang Chen 0001, Yijun Bei |
J. Comput. Sci. Technol. | 1 |
| 2009 | PISA: Federated Search in P2P Networks with Uncooperative Peers
Zujie Ren, Lidan Shou, Gang Chen 0001, Chun Chen 0001, Yijun Bei |
DEXA | 1 |