VLDB 2026 Research / reviewers in the wild / expert
Kuan Zou
dblp:315/4341
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hesitation and Tolerance in Recommender SystemsabstractUsers’ interactions with recommender systems often involve more than simple acceptance or rejection. We highlight two overlooked states: hesitation, when people deliberate without certainty, and tolerance, when this hesitation escalates into unwanted engagement before ending in disinterest. Across two large-scale surveys (N = 6, 644 and N = 3, 864), hesitation was nearly universal, and tolerance emerged as a recurring source of wasted time, frustration, and diminished trust. Analyses of e-commerce and short-video platforms confirm that tolerance behaviors, such as clicking without purchase or shallow viewing, correlate with decreased activity. Finally, an online field study at scale shows that even lightweight strategies treating tolerance as distinct from interest can improve retention while reducing wasted effort. By surfacing hesitation and tolerance as consequential states, this work reframes how recommender systems should interpret feedback, moving beyond clicks and dwell time toward designs that respect user value, reduce hidden costs, and sustain engagement. Kuan Zou, Aixin Sun, Yitong Ji, Hao Zhang 0048, Jing Wang 0060, Zhuohao (Jerry) Zhang, Xuemeng Jiang |
CHI | 1 |
| 2026 | Cooperation dynamics on hypergraphs with punishment and Q-learning
Kuan Zou, Changwei Huang |
Expert Syst. Appl. | 1 |
| 2023 | Revisiting "code smell severity classification using machine learning techniques"abstractIn the context of limited maintenance resources, predicting the severity of code smells is more practically useful than simply detecting them. Fontana et al. first empirically investigated some classification algorithms and some regression algorithms, for severity prediction. Their results showed that random forest and decision tree performed well on Mean Absolute Error (MAE), Mean Squared Error (MSE), and Spearman and Kendall rank correlation coefficients. However, they did not consider the issue of imbalanced data distribution in the severity dataset, and used inappropriate performance evaluation metrics. Therefore, we revisit the effectiveness of 10 classification methods and 11 regression methods, for code severity prediction using Cumulative Lift Chart (CLC) and Severity@20% as the primary performance metrics and Accuracy as the secondary performance indicator. The results show that the Gradient Boosting Regression (GBR) method performs the best in terms of these metrics. Lei Liu 0062, Peixin Yang, Kuan Zou, Guancheng Lin, Jianwen Xiang |
COMPSAC | 4 |
| 2023 | On the relative value of imbalanced learning for code smell detectionabstractSummary Machine learning‐based code smell detection (CSD) has been demonstrated to be a valuable approach for improving software quality and enabling developers to identify problematic patterns in code. However, previous researches have shown that the code smell datasets commonly used to train these models are heavily imbalanced. While some recent studies have explored the use of imbalanced learning techniques for CSD, they have only evaluated a limited number of techniques and thus their conclusions about the most effective methods may be biased and inconclusive. To thoroughly evaluate the effect of imbalanced learning techniques for machine learning‐based CSD, we examine 31 imbalanced learning techniques with seven classifiers to build CSD models on four code smell data sets. We employ four evaluation metrics to assess the detection performance with the Wilcoxon signed‐rank test and Cliff's . The results show that (1) Not all imbalanced learning techniques significantly improve detection performance, but deep forest significantly outperforms the other techniques on all code smell data sets. (2) SMOTE (Synthetic Minority Over‐sampling TEchnique) is not the most effective technique for resampling code smell data sets. (3) The best‐performing imbalanced learning techniques and the top‐3 data resampling techniques have little time cost for code smell detection. Therefore, we provide some practical guidelines. First, researchers and practitioners should select the appropriate imbalanced learning techniques (e.g., deep forest) to ameliorate the class imbalance problem. In contrast, the blind application of imbalanced learning techniques could be harmful. Then, better data resampling techniques than SMOTE should be selected to preprocess the code smell data sets. Kuan Zou, Jacky W. Keung, Xiao Yu 0008, Shuo Feng 0003, Yan Xiao 0002 |
Softw. Pract. Exp. | 2 |
| 2022 | Multi-CPR: A Multi Domain Chinese Dataset for Passage RetrievalabstractPassage retrieval is a fundamental task in information retrieval (IR) research, which has drawn much attention recently. In the English field, the availability of large-scale annotated dataset (e.g, MS MARCO) and the emergence of deep pre-trained language models (e.g, BERT) has resulted in a substantial improvement of existing passage retrieval systems. However, in the Chinese field, especially for specific domains, passage retrieval systems are still immature due to quality-annotated dataset being limited by scale. Therefore, in this paper, we present a novel multi-domain Chinese dataset for passage retrieval (Multi-CPR). The dataset is collected from three different domains, including E-commerce, Entertainment video and Medical. Each dataset contains millions of passages and a certain amount of human annotated query-passage related pairs. We implement various representative passage retrieval methods as baselines. We find that the performance of retrieval models trained on dataset from general domain will inevitably decrease on specific domain. Nevertheless, a passage retrieval system built on in-domain annotated dataset can achieve significant improvement, which indeed demonstrates the necessity of domain labeled data for further optimization. We hope the release of the Multi-CPR dataset could benchmark Chinese passage retrieval task in specific domain and also make advances for future studies. Dingkun Long, Qiong Gao, Kuan Zou, Pengjun Xie, Ruijie Guo, Guanjun Jiang, Luxi Xing |
SIGIR | 3 |