VLDB 2026 Research / reviewers in the wild / expert
Jin Cheng 0008
dblp:230/9055-8
· DBLP profile ↗
9ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-3141-3460ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trading Vector Data in Vector DatabasesabstractVector data trading is essential for cross-domain learning with vector databases, yet it remains largely unexplored. We study this problem under online learning, where sellers face uncertain retrieval costs and buyers provide stochastic feedback to posted prices. Three main challenges arise: (1) heterogeneous and partial feedback in configuration learning, (2) variable and complex feedback in pricing learning, and (3) inherent coupling between configuration and pricing decisions. We propose a hierarchical bandit framework that jointly optimizes retrieval configurations and pricing. Stage I employs contextual clustering with confidence-based exploration to learn effective configurations with logarithmic regret. Stage II adopts interval-based price selection with local Taylor approximation to estimate buyer responses and achieve sublinear regret. We establish theoretical guarantees with polynomial time complexity and validate the framework on four real-world datasets, demonstrating consistent improvements in cumulative reward and regret reduction compared with existing methods. Jin Cheng 0008, Xiangxiang Dai, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
ICDE | 1 |
| 2026 | BANCO: Drift-Aware Batched Bandits for Adaptive Proximity Graph PruningabstractProximity graphs are the state-of-the-art solution for approximate nearest neighbor (ANN) search, supporting applications such as Web search and retrieval-augmented generation (RAG). Sustaining long-term performance requires adaptive pruning as data and query workloads evolve. However, existing approaches are largely static and uniform. Adaptive pruning faces three key challenges: temporal drift in data and query distributions, spatial heterogeneity across graph regions, and costly feedback due to graph-level evaluations. We present BANCO, a bandit-based framework for adaptive proximity graph pruning. BANCO unifies diverse pruning strategies within a common decision space and optimizes them via a drift-aware batched bandit algorithm. It addresses temporal drift through drift-aware updates, captures spatial heterogeneity using contextual features for region-specific pruning, and reduces evaluation costs through batched feedback aggregation. We establish a dynamic regret bound with sublinear loss and polynomial computational complexity. Extensive experiments on four real-world datasets demonstrate that BANCO helps maintain long-term ANN search efficiency and accuracy under evolving data and workloads. Jin Cheng 0008, Xiangxiang Dai, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
WWW | 1 |
| 2026 | COTRA: A Data Trading Framework for Multi-Source Data Cooperation
Jin Cheng 0008, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2026 | Trading Continuous QueriesabstractIn the bigdata era, data trading significantly enhances data-driven decision-making by facilitating data sharing. Streaming data from sources such as mobile devices and social media platforms creates new opportunities and challenges for data trading. Traditional data trading methods, designed for one time queries over static data base snapshots, neglect the growing need for trading continuous queries over streaming data. If applied directly to continuous queries, existing methods often result in repeated and imprecise charges that reduce the seller's profit, as they do not consider computation sharing during continuous query execution. To address these challenges, we propose CQ Trade, the first mechanism for continuous query based data trading, which incorporates computation sharing in query execution and integrates seamlessly with existing trading mechanisms. Our contributions are threefold: (1) we provide a theoretical analysis of prevalent computation-sharing techniques, including costmodeling and closed-form computation-sharing strategy derivation; (2) we formulate a general optimization problem to maximize the seller's profit, adaptable to various computation-sharing techniques; (3) we identify that our op timization problem merges vector bin packing and multidimensional knapsack challenges, and we tackle this complexity with a tailored branch-and-price algorithm that decomposes the problem in to a master problem and multiple sub-problems, achieving a globally optimal solution. Evaluation shows CQ Trade improve strading success rate by 12.8% and increases seller profit by 28.7% compared to traditional methods. Jin Cheng 0008, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | OSTOR: Online Scheduling Framework for Trading Continuous QueriesabstractData trading significantly enhances data utility by enabling data sharing across diverse applications. Despite being crucial for real-time analytics and online machine learning, trading continuous queries with streaming data output remains largely unexplored. The inherent characteristics of trading continuous queries pose distinctive technical challenges in scheduling query execution. First, the streaming nature demands online scheduling under information uncertainty, where data utilities and execution costs vary unpredictably during query execution. Second, the intrinsic NP-hardness of the optimization problem, coupled with repeated invocation requirements, necessitates efficient algorithmic solutions to address computational complexity. We present OSTOR, the first online scheduling framework for trading continuous queries. OSTOR aims to maximize social welfare, defined as the difference between buyers' obtained utilities and sellers' execution costs, while achieving both theoretical guarantees and practical efficiency. To handle the information uncertainty, we present a primary-dual decomposition method that transforms the online scheduling problem into multiple one-round integer programming problems, enabling adaptive decision-making that only needs current system information. To address the computational complexity, we design an adaptive dual descent (ADD) algorithm that iteratively optimizes dual variables, achieving a bounded constant approximation ratio in polynomial time. We further enhance OSTOR through structureaware greedy optimization strategies with provable performance guarantees. Extensive experiments demonstrate that OSTOR substantially improves social welfare and reduces query execution costs on both real-world and synthetic datasets, compared to existing data trading methods. Jin Cheng 0008, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
ICDE | 1 |
| 2024 | Cooperative Multi-source Data TradingabstractIn the era of big data, data trading significantly enhances data-driven technologies by facilitating data sharing. Despite the clear advantages often experienced by data users when incorporating multiple sources, the topic of multi-source data trading remains largely unexplored. This paper designs a novel data trading framework, which enables multi-source data trading through multi-source cooperation. The proposed framework aims to improve data usage efficiency and increase seller revenue. In particular, we model data sellers’ cooperative decisions through the Nash bargaining framework and systematically outline the interactions between sellers and buyers as a two-stage Stackelberg game. A key contribution of this work is the consideration of coupling among diverse data products, which is essential but often overlooked in prior studies. We properly classify data’s utility into endogenous and relational categories to disentangle the coupling. Despite the inherent non-convex nature of the optimization problem, we methodically derive the closed-form optimal solutions by decomposing the problem into several subproblems. Interestingly, we reveal that, under our proposed framework, sellers’ revenue initially remains steady with the increase of product coupling level, but begins to rise once the level exceeds a certain threshold due to the substitute effect. Finally, experimental results show that our proposed framework can improve the seller’s profit by up to 46.32% compared to traditional data trading methods in the current data market. Jin Cheng 0008, Ningning Ding, John C. S. Lui, Jianwei Huang 0001 |
GLOBECOM | 1 |
| 2021 | MATEC: A lightweight neural network for online encrypted traffic classification
Jin Cheng 0008, Yulei Wu, Yuepeng E, Junling You, Tong Li 0012, Hui Li 0098, Jingguo Ge |
Comput. Networks | 1 |
| 2021 | VNE-HRL: A Proactive Virtual Network Embedding Algorithm Based on Hierarchical Reinforcement LearningabstractVirtual network embedding (VNE) that instantiates virtualized networks on a substrate infrastructure, is one of the key research problems for network virtualization. Most existing VNE approaches, however, focus on the current virtual network request (VNR) and treat all VNRs equally, which disregard the long-term impact and waste many resources on the process of embedding infeasible VNRs (i.e., VNRs that cannot be embedded completely). To address these problems, a proactive virtual network embedding algorithm based on hierarchical reinforcement learning, VNE-HRL, is proposed in this paper. Within our framework, the VNE task is performed by a two-level agent that considers both the long-term impact of a VNR and the short-term effect of an embedding action. For each processing, a high-level agent aims to select a currently feasible VNR with the maximum long-term reward from a window-based batch, and a low-level agent is assigned to embed the selected VNR on a substrate infrastructure by performing a series of embedding actions. Extensive simulation results indicate that our algorithm best performance on most metrics compared with existing state-of-the-art solutions, with up to 9.92% and 33.03% improvement on acceptance ratio and average revenue. Jin Cheng 0008, Yulei Wu, Yeming Lin, Yuepeng E, Fan Tang, Jingguo Ge |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | Real-Time Encrypted Traffic Classification via Lightweight Neural NetworksabstractThe fast growth of encrypted traffic puts forward burning requirements on the efficiency of traffic classification. Although deep learning models perform well in the classification, they sacrifice the efficiency to obtain high-precision results. To reduce the resource and time consumption, a novel and lightweight model is proposed in this paper. Our design principle is to “maximize the reuse of thin modules A thin module adopts the multi-head attention and the 1D convolutional network. Attributed to the one-step interaction of all packets and the parallelized computation of the multi-head attention mechanism, a key advantage of our model is that the number of parameters and running time are significantly reduced. In addition, the effectiveness and efficiency of 1D convolutional networks are proved in traffic classification. Besides, the proposed model can work well in a real time manner, since only three consecutive packets of a flow are needed. To improve the stability of the model, the designed network is trained with the aid of ResNet, layer normalization and learning rate warm up. The proposed model outperforms the state-of-the-art works based on deep learning on two public datasets. The results show that our model has higher accuracy and running efficiency, while the number of parameters used is 1.8% of the 1D convolutional network and the training time halves. Jin Cheng 0008, Runkang He, Yuepeng E, Yulei Wu, Junling You, Tong Li 0012 |
GLOBECOM | 1 |