VLDB 2026 Research / reviewers in the wild / expert
Jingyi Cui
dblp:216/3282
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Clinically Consistent Ordinal Learning for Early Pneumoconiosis Diagnosis
Binglu Wang, Jingyi Cui |
KSEM (2) | 2 |
| 2026 | HMMShard: An HMM-based adaptive sharding framework for IIoT blockchain
Guangquan Xu, Cong Wang 0004, Jingyi Cui, Zenghao Yang |
J. Netw. Comput. Appl. | 4 |
| 2026 | VL-HTR: Learning Human-Target Representation From Vision-Language ModelabstractHuman-gaze-target prediction aims to predict the target point or object that humans are looking at in images. However, existing methods predominantly rely on vision-only features, which often struggle to capture the semantic context of small or occluded objects and lack explicit priors for precise head direction regression, leading to slow convergence and suboptimal performance. Therefore, we introduce VL-HTR, a novel vision-language learning method for human-target representation, which integrates multimodal knowledge from vision-language models (VLMs) to construct robust human-target relationships. Unlike traditional approaches, extracting multimodal features via pretrained VLMs enhances the model's grasp of human-target knowledge through the learnable target class and direction context. Then, a language-guided query alignment (LQA) module is introduced to improve the semantic-aware object representation capability through vision-language query alignment. Finally, to accelerate the gaze point regression learning process, we design a language-guided direction prediction (LDP) module to introduce multimodal human gaze direction priors, thereby facilitating the human-target relationship construction. Extensive validations across two distinct tasks, i.e., gaze object prediction (GOP) and gaze target estimation, involving five challenging benchmarks, demonstrating that VL-HTR achieves superior performance and much faster training convergence. Binglu Wang, Jingyi Cui, Haisheng Xia, Guangyu Guo 0001, Zhijun Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2026 | Rubato: Efficient Post-Quantum Asynchronous Distributed Randomness Beacon With Integrated ConsensusabstractDistributed randomness beacons are essential for distributed systems (e.g., blockchain and MPC), providing un biased and unpredictable shared randomness. However, implementations in asynchronous networks often suffer from poor scal ability, low throughput, and high resource consumption when deployed as independent protocols. The state-of-the-art HashRand (CCS'24) achieves high throughput and low computational in tensity using only lightweight post-quantum cryptographic primitives for an independent asynchronous beacon. Building further on this, we propose Rubato, a low-overhead, high-throughput beacon protocol that leverages lightweight batched Asynchronous Complete Secret Sharing with Byzantine Atomic Broadcast-based state machine replication (via our tailored RubatoSMR). Rubato reduces communication complexity by an O(clogn) factor compared to HashRand and resolves the circular dependency between beacon and BAB-SMR in asynchronous settings. Experiments on AWSdemonstrate that Rubato achieves ideal overall performance in scalability, resource usage, and throughput; for instance, at n = 121nodes, it produces an average of 174 beacons per minute, with RubatoSMR further optimizing memory and bandwidth consumption. Linghe Yang, Tonghong Chong, Jian Liu 0004, Jingyi Cui, Guangquan Xu, Yude Bai, Lei Zhang 0024, Tao Luo 0010 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Beyond Interpretability: The Gains of Feature Monosemanticity on Model RobustnessabstractDeep learning models often suffer from a lack of interpretability due to \emph{polysemanticity}, where individual neurons are activated by multiple unrelated semantics, resulting in unclear attributions of model behavior. Recent advances in \emph{monosemanticity}, where neurons correspond to consistent and distinct semantics, have significantly improved interpretability but are commonly believed to compromise accuracy. In this work, we challenge the prevailing belief of the accuracy-interpretability tradeoff, showing that monosemantic features not only enhance interpretability but also bring concrete gains in model performance of {\color{black} robustness-related tasks}. Across multiple robust learning scenarios—including input and label noise, few-shot learning, and out-of-domain generalization—our results show that models leveraging monosemantic features significantly outperform those relying on polysemantic features. Furthermore, we provide empirical and theoretical understandings on the robustness gains of feature monosemanticity. Our preliminary analysis suggests that monosemanticity, by promoting better separation of feature representations, leads to more robust decision boundaries {\color{black} under noise}. This diverse evidence highlights the \textbf{generality} of monosemanticity in improving model robustness. As a first step in this new direction, we embark on exploring the learning benefits of monosemanticity beyond interpretability, supporting the long-standing hypothesis of linking interpretability and robustness. Code is available at \url{https://github.com/PKU-ML/Monosemanticity-Robustness}. Qi Zhang 0067, Yifei Wang 0001, Jingyi Cui, Xiang Pan 0001, Stefanie Jegelka, Yisen Wang 0001 |
ICLR | 3 |
| 2025 | An Augmentation-Aware Theory for Self-Supervised Contrastive LearningabstractSelf-supervised contrastive learning has emerged as a powerful tool in machine learning and computer vision to learn meaningful representations from unlabeled data. Meanwhile, its empirical success has encouraged many theoretical studies to reveal the learning mechanisms. However, in the existing theoretical research, the role of data augmentation is still under-exploited, especially the effects of specific augmentation types. To fill in the blank, we for the first time propose an augmentation-aware error bound for self-supervised contrastive learning, showing that the supervised risk is bounded not only by the unsupervised risk, but also explicitly by a trade-off induced by data augmentation. Then, under a novel semantic label assumption, we discuss how certain augmentation methods affect the error bound. Lastly, we conduct both pixel- and representation-level experiments to verify our proposed theoretical results. Jingyi Cui, Hongwei Wen, Yisen Wang 0001 |
ICML | 1 |
| 2024 | ID-SR: Privacy-Preserving Social Recommendation Based on Infinite Divisibility for Trustworthy AIabstractRecommendation systems powered by artificial intelligence (AI) are widely used to improve user experience. However, AI inevitably raises privacy leakage and other security issues due to the utilization of extensive user data. Addressing these challenges can protect users’ personal information, benefit service providers, and foster service ecosystems. Presently, numerous techniques based on differential privacy have been proposed to solve this problem. However, existing solutions encounter issues such as inadequate data utilization and a tenuous trade-off between privacy protection and recommendation effectiveness. To enhance recommendation accuracy and protect users’ private data, we propose ID-SR, a novel privacy-preserving social recommendation scheme for trustworthy AI based on the infinite divisibility of Laplace distribution. We first introduce a novel recommendation method adopted in ID-SR, which is established based on matrix factorization with a newly designed social regularization term for improving recommendation effectiveness. We then propose a differential privacy-preserving scheme tailored to the above method that leverages the Laplace distribution’s characteristics to safeguard user data. Theoretical analysis and experimentation evaluation on two publicly available datasets demonstrate that our scheme achieves a superior balance between privacy protection and recommendation effectiveness, ultimately delivering an enhanced user experience. Jingyi Cui, Guangquan Xu, Jian Liu 0004, Shicheng Feng, Jianli Wang, Hao Peng 0002, Shihui Fu, Zhaohua Zheng, James Xi Zheng, Shaoying Liu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | IGA : An Improved Genetic Algorithm to Construct Weightwise (Almost) Perfectly Balanced Boolean Functions with High Weightwise NonlinearityabstractThe Boolean functions satisfying secure properties on the restricted sets of inputs are studied recently due to their importance in the framework of the FLIP stream cipher. However, finding Boolean functions with optimal cryptographic properties is an open research problem in the cryptographic community. This paper presents an Improved Genetic Algorithm (IGA) with the directed changes that keep the weightwise balancedness of Boolean functions. A cross-protection strategy is proposed to ensure that the offspring has the same weightwise balancedness characteristics of the parents while implementing crossover. Then, a large number of weightwise (almost) perfectly balanced (W(A)PB) functions with a good nonlinearity profile are obtained based on IGA. Finally, we make comparisons between our constructions and relevant works. The comparisons show that IGA has a significant advantage for reaching the W(A)PB functions with high weightwise nonlinearity. Moreover, it is the first time to obtain the 8-variable WPB functions with the weightwise nonlinearity of 28 in the restricted sets of inputs with Hamming weight of 4, and list the statistical indicators of the weightwise nonlinearity for W(A)PB functions for input size n = 9, 10. Jingyi Cui, Jian Liu 0004, Guangquan Xu, Lidong Han, Alireza Jolfaei, James Xi Zheng |
AsiaCCS | 2 |
| 2023 | GDTM: Gaussian Differential Trust Mechanism for Optimal Recommender System
Lixiao Gong, Guangquan Xu, Jingyi Cui, Shihui Fu, James Xi Zheng, Shaoying Liu |
ICA3PP (6) | 3 |
| 2023 | Rethinking Weak Supervision in Helping Contrastive LearningabstractContrastive learning has shown outstanding performances in both supervised and unsupervised learning, and has recently been introduced to solve weakly supervised learning problems such as semi-supervised learning and noisy label learning. Despite the empirical evidence showing that semi-supervised labels improve the representations of contrastive learning, it remains unknown if noisy supervised information can be directly used in training instead of after manual denoising. Therefore, to explore the mechanical differences between semi-supervised and noisy-labeled information in helping contrastive learning, we establish a unified theoretical framework of contrastive learning under weak supervision. Specifically, we investigate the most intuitive paradigm of jointly training supervised and unsupervised contrastive losses. By translating the weakly supervised information into a similarity graph under the framework of spectral clustering based on the posterior probability of weak labels, we establish the downstream classification error bound. We prove that semi-supervised labels improve the downstream error bound whereas noisy labels have limited effects under such a paradigm. Our theoretical findings here provide new insights for the community to rethink the role of weak supervision in helping contrastive learning. Jingyi Cui, Weiran Huang 0001, Yifei Wang 0001, Yisen Wang 0001 |
ICML | 1 |
| 2023 | GenDroid: A query-efficient black-box android adversarial attack framework
Guangquan Xu, Hongfei Shao, Jingyi Cui, Hongpeng Bai, Guangdong Bai, Shaoying Liu, Weizhi Meng 0001, James Xi Zheng |
Comput. Secur. | 3 |
| 2021 | GBHT: Gradient Boosting Histogram Transform for Density EstimationabstractIn this paper, we propose a density estimation algorithm called \textit{Gradient Boosting Histogram Transform} (GBHT), where we adopt the \textit{Negative Log Likelihood} as the loss function to make the boosting procedure available for the unsupervised tasks. From a learning theory viewpoint, we first prove fast convergence rates for GBHT with the smoothness assumption that the underlying density function lies in the space $C^{0,\alpha}$. Then when the target density function lies in spaces $C^{1,\alpha}$, we present an upper bound for GBHT which is smaller than the lower bound of its corresponding base learner, in the sense of convergence rates. To the best of our knowledge, we make the first attempt to theoretically explain why boosting can enhance the performance of its base learners for density estimation problems. In experiments, we not only conduct performance comparisons with the widely used KDE, but also apply GBHT to anomaly detection to showcase a further application of GBHT. Jingyi Cui, Hanyuan Hang, Yisen Wang 0001, Zhouchen Lin |
ICML | 1 |
| 2021 | Leveraged Weighted Loss for Partial Label LearningabstractAs an important branch of weakly supervised learning, partial label learning deals with data where each instance is assigned with a set of candidate labels, whereas only one of them is true. Despite many methodology studies on learning from partial labels, there still lacks theoretical understandings of their risk consistent properties under relatively weak assumptions, especially on the link between theoretical results and the empirical choice of parameters. In this paper, we propose a family of loss functions named \textit{Leveraged Weighted} (LW) loss, which for the first time introduces the leverage parameter $\beta$ to consider the trade-off between losses on partial labels and non-partial ones. From the theoretical side, we derive a generalized result of risk consistency for the LW loss in learning from partial labels, based on which we provide guidance to the choice of the leverage parameter $\beta$. In experiments, we verify the theoretical guidance, and show the high effectiveness of our proposed LW loss on both benchmark and real datasets compared with other state-of-the-art partial label learning algorithms. Hongwei Wen, Jingyi Cui, Hanyuan Hang, Yisen Wang 0001, Zhouchen Lin |
ICML | 2 |