EDBT 2026 Demo / reviewers in the wild / expert
Yuqiu Qian
dblp:147/1552
· DBLP profile ↗
20ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0007-8894-3271ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SC-DAG: Semantic-Constrained Diffusion Attacks for Stealthy Exposure Manipulation in Visually-Aware Recommender SystemsabstractVisually-aware recommender system (VARS) has become increasingly prevalent in various online services by integrating visual features of items to enhance recommendation quality. However, VARS introduces new security vulnerabilities and malicious attackers can perform visual shilling attacks to manipulate recommendation lists via uploading generated images with visually imperceptible perturbations. While prior research has explored such threats to help service providers enhance their systems, existing visual shilling attack methods still suffer from uncontrolled pixel-space perturbation, energy dispersion dilemma and semantic misalignment in reference selection. In this work, we present Semantic-Constrained Diffusion Adversarial Generation (SC-DAG) for visual shilling attacks. SC-DAG overcomes key limitations of previous methods by focusing perturbations on semantically meaningful image regions through contour-aware segmentation, guiding adversarial generation in latent space using a conditional diffusion process, and performing a hybrid reference image selection strategy that balances popularity and semantic similarity. Extensive experiments on performing visual shilling attacks against multiple VARS models show that SC-DAG achieves state-of-the-art attack performance in elevating target items' ranking, while maintaining strong perceptual indistinguishability and minimal impact on overall recommendation performance of the system. Our work offers insights into leveraging structured semantic priors for more sophisticated adversarial manipulations against VARS and also highlights the necessity for developing more robust VARS models resilient to visual shilling attacks. We provide our implementation at https://github.com/KDEGroup/SC-DAG. Yuqiu Qian, Xiaodong Li 0009, Ziyu Lyu, Hui Li 0057 |
CIKM | 2 |
| 2024 | Locally Private Set-Valued Data Analyses: Distribution and Heavy Hitters EstimationabstractIn many mobile applications, user-generated data are presented as set-valued data. To tackle potential privacy threats in analyzing these valuable data, local differential privacy has been attracting substantial attention. However, existing approaches only provide sub-optimal utility and are expensive in computation and communication for set-valued data distribution estimation and heavy-hitter identification. In this paper, we propose a utility-optimal and efficient set-valued data publication method (i.e.,Wheel mechanism). On the user side, the computational complexity is only$O(\min \lbrace m\log m, m e^\epsilon \rbrace )$and communication costs are$O(\epsilon +\log m)$bits, where$m$is the number of items,$d$is the domain size and$\epsilon$is the privacy budget, while existing approaches usually depend on$O(d)$or$O(\log d)$($d \gg m$). Our theoretical analyses reveal the estimation errors have been reduced from the previously known$O(\frac{m^{2} d}{n\epsilon ^{2}})$to the optimal rate$O(\frac{m d}{n\epsilon ^{2}})$. Additionally, for heavy-hitter identification, we present a variant of the Wheel mechanism as an efficient frequency oracle, entailing only$O(\sqrt{n})$computational complexity. This heavy-hitter protocol achieves an identification bar of$\tilde{O}(\frac{1}{\epsilon }\sqrt{\frac{m}{n} \log d})$, reducing by a factor of$\sqrt{m}$relative to existing protocols. Extensive experiments demonstrate our methods are 3-100x faster than existing approaches and have optimized statistical efficiency. Shaowei Wang 0003, Yuntong Li, Yusen Zhong, Kongyang Chen, Xianmin Wang, Zhili Zhou 0001, Fei Peng 0001, Yuqiu Qian, Jiachun Du, Wei Yang 0011 |
IEEE Trans. Mob. Comput. | 8 |
| 2023 | Fine-Grained Private Knowledge DistillationabstractKnowledge distillation has emerged as a scalable and effective way for privacy-preserving machine learning. One remaining drawback is that it consumes privacy in a client-level manner. In order to attain fine-grained privacy accountant and improve utility, this work proposes a model-free reverse k-NN labeling method towards record-level private knowledge distillation, where each private record is employed for labeling at most k queries. Theoretically, we provide bounds of labeling error rate under the centralized/local model of differential privacy. Experimentally, we demonstrate that it achieves new state-of-the-art accuracy in MNIST/SVHN/CIFAR-10 dataset with one order of magnitude lower of privacy loss. Yuntong Li, Shaowei Wang 0003, Jin Li 0002, Yuqiu Qian, Bangzhou Xin, Wei Yang 0011 |
ICASSP | 5 |
| 2023 | Analyzing Preference Data With Local Privacy: Optimal Utility and Enhanced RobustnessabstractOnline service providers benefit from collecting and analyzing preference data from users, including both implicit preference data (e.g., watched videos of a user) and explicit preference data (e.g., ranking data over candidates). However, it brings ethical and legal issues of data privacy at the same time. In this paper, we study the problem of aggregating individual's preference data in the local differential privacy (LDP) setting. One naive approach is to add Laplace random noises, which however suffers from low statistical utility and is fragile to LDP-specific poisoning attacks. Therefore, we propose a novel mechanism to improve the utility and the robustness simultaneously: theadditive mechanism. The additive mechanism randomly outputs a subset of candidates with a probability proportional to their total scores. For preference data with Borda rule over$d$items, its mean squared error bound is optimized from$O(\frac{d^{5}}{n\epsilon ^{2}})$to$O(\frac{d^{4}}{n\epsilon ^{2}})$, and its maximum poisoning risk bound is reduced from$+\infty$to$O(\frac{d^{2}}{n\epsilon })$. We also theoretically investigate minimax lower bounds of$\epsilon$-LDP preference data aggregation, and prove the error rate of$O(\frac{d^{4}}{n\epsilon ^{2}})$is optimal for the Borda rule. Experimental results validate that our proposed approaches averagely reduce estimation error by 50% and are more robust to adversarial poisoning attacks. Shaowei Wang 0003, Xuandi Luo, Yuqiu Qian, Jiachun Du, Wenqing Lin, Wei Yang 0011 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Shuffle Differential Private Data Aggregation for Random PopulationabstractBridging the advantages of differential privacy in both centralized model (i.e., high accuracy) and local model (i.e., minimum trust), the shuffle privacy model has potential applications in many privacy-sensitive scenarios, such as mobile user data aggregation and federated learning. Since messages from users are anonymized by semi-trusted shufflers (e.g., anonymous channels, edge servers), every user could hide message among other users’ messages and inject only part of noises (a.k.a. privacy amplification). However, existing works assume that the participating user population is known in advance, which is unrealistic for dynamic environments (e.g., mobile computing, vehicular networks). In this work, we study the shuffle privacy model with a random participating population, and give privacy amplification bounds for population size with commonly encountered binomial, Poisson, sub-Gaussian distribution and etc. For further improving accuracy, we formulate and derive optimal dummy sizes for both non-adaptive and adaptive dummies. Finally, to break the error barrier due to the constraint of sending one single message per user, we design a multi-message shuffle private protocol supporting random population. Experiment results show that our approaches reduce more than 60% error when compared to the local model and naive approaches. We hope this work provides tailored solutions of shuffle privacy for dynamic mobile/distributed computing. Shaowei Wang 0003, Xuandi Luo, Yuqiu Qian, Youwen Zhu, Kongyang Chen, Qi Chen 0024, Bangzhou Xin, Wei Yang 0011 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Fast and Secure Distributed Nonnegative Matrix FactorizationabstractNonnegative matrix factorization (NMF) has been successfully applied in several data mining tasks. Recently, there is an increasing interest in the acceleration of NMF, due to its high cost on large matrices. On the other hand, the privacy issue of NMF over federated data is worthy of attention, since NMF is prevalently applied in image and text analysis which may involve leveraging privacy data (e.g, medical image and record) across several parties (e.g., hospitals). In this paper, we study theaccelerationandsecurityproblems of distributed NMF. First, we propose adistributed sketched alternating nonnegative least squares(DSANLS) framework for NMF, which utilizes a matrix sketching technique to reduce the size of nonnegative least squares subproblems with a convergence guarantee. For the second problem, we show that DSANLS with modification can be adapted to the security setting, but only forone or limited iterations. Consequently, we propose four efficient distributed NMF methods in both synchronous and asynchronous settings with a security guarantee. We conduct extensive experiments on several real datasets to show the superiority of our proposed methods. The implementation of our methods is available athttps://github.com/qianyuqiu79/DSANLS. Yuqiu Qian, Conghui Tan, Danhao Ding, Hui Li 0057, Nikos Mamoulis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Hiding Numerical Vectors in Local Private and Shuffled MessagesabstractNumerical vector aggregation has numerous applications in privacy-sensitive scenarios, such as distributed gradient estimation in federated learning, and statistical analysis on key-value data. Within the framework of local differential privacy, this work gives tight minimax error bounds of O(d s/(n epsilon^2)), where d is the dimension of the numerical vector and s is the number of non-zero entries. An attainable mechanism is then designed to improve from existing approaches suffering error rate of O(d^2/(n epsilon^2)) or O(d s^2/(n epsilon^2)). To break the error barrier in the local privacy, this work further consider privacy amplification in the shuffle model with anonymous channels, and shows the mechanism satisfies centralized (14 ln(2/delta) (s e^epsilon+2s-1)/(n-1))^0.5, delta)-differential privacy, which is domain independent and thus scales to federated learning of large models. We experimentally validate and compare it with existing approaches, and demonstrate its significant error reduction. Shaowei Wang 0003, Jin Li 0002, Yuqiu Qian, Jiachun Du, Wenqing Lin, Wei Yang 0011 |
IJCAI | 3 |
| 2021 | Sequential Recommendation in Online Games with Multiple Sequences, Tasks and User LevelsabstractOnline gaming is growing faster than ever before, with increasing challenges of providing better user experience. Recommender systems (RS) for online games face unique challenges since they must fulfill players’ distinct desires, at different user levels, based on their action sequences of various action types. Although many sequential RS already exist, they are mainly single-sequence, single-task, and single-user-level. In this paper, we introduce a new sequential recommendation model for multiple sequences, multiple tasks, and multiple user levels (abbreviated as M3Rec) in Tencent Games platform, which can fully utilize complex data in online games. We leverage Graph Neural Network and multi-task learning to design M3Rec in order to model the complex information in the heterogeneous sequential recommendation scenario of Tencent Games. We verify the effectiveness of M3Rec on three online games of Tencent Games platform, in both offline and online evaluations. The results show that M3Rec successfully addresses the challenges of recommendation in online games, and it generates superior recommendations compared with state-of-the-art sequential recommendation approaches. Si Chen 0011, Yuqiu Qian, Hui Li 0057, Chen Lin 0001 |
SSTD | 2 |
| 2020 | An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing PlatformsabstractIn this paper, we propose a Deep Reinforcement Learning (RL) framework for task arrangement, which is a critical problem for the success of crowdsourcing platforms. Previous works conduct the personalized recommendation of tasks to workers via supervised learning methods. However, the majority of them only consider the benefit of either workers or requesters independently. In addition, they do not consider the real dynamic environments (e.g., dynamic tasks, dynamic workers), so they may produce sub-optimal results. To address these issues, we utilize Deep Q-Network (DQN), an RL-based method combined with a neural network to estimate the expected long-term return of recommending a task. DQN inherently considers the immediate and the future rewards and can be updated quickly to deal with evolving data and dynamic changes. Furthermore, we design two DQNs that capture the benefit of both workers and requesters and maximize the profit of the platform. To learn value functions in DQN effectively, we also propose novel state representations, carefully design the computation of Q values, and predict transition probabilities and future states. Experiments on synthetic and real datasets demonstrate the superior performance of our framework. Nikos Mamoulis, Reynold Cheng, Guoliang Li 0001, Xiang Li 0067, Yuqiu Qian |
ICDE | 6 |
| 2020 | Set-valued Data Publication with Local Privacy: Tight Error Bounds and Efficient MechanismsabstractMost user-generated data in online services are presented as set-valued data, e.g., visited website URLs, recently used Apps by a person, and etc. These data are of great value to service providers, but also bring privacy concerns if collected and analyzed directly. To tackle potential privacy threatens, local differential privacy (LDP) attracts increasing attention nowadays. However, existing approaches only provide sub-optimal error bound for set-valued data distribution estimation with LDP. Besides, it is computational expensive and communication expensive to use for high dimensional set-valued data, considering large domains in real scenarios. Thus, existing approaches are unpractical to use on resource-constrained user-side devices (e.g., smartphones and wearable devices). In this paper, we propose a utility-optimal and efficient set-valued data publication method (i.e., wheel mechanism ). On the user side, each user contributes only one numerical value to represent their privatized data. The computational complexity is O (min{ m log m , me ɛ }) and communication cost is O (log( me ɛ )) bits, while existing approaches usually depend on O ( d ) or O (log d ), where m is the number of items in the set-valued data ( m ≡ 1 for categorical data), d is the domain size (usually d ≫ m ) and ɛ is the privacy budget. On the server side, the estimator takes numerical values from users as input and derives an unbiased distribution estimation. Theoretical results show that estimation error bounds are improved from previously known [EQUATION] to the optimal rate [EQUATION]. Results on extensive experiments demonstrate that our proposed wheel mechanism is 3-100× faster than existing approaches, meanwhile has optimal statistical efficiency. Shaowei Wang 0003, Yuqiu Qian, Jiachun Du, Wei Yang 0011, Liusheng Huang, Hongli Xu 0001 |
Proc. VLDB Endow. | 2 |
| 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, Danilo Jimenez Rezende |
ICLR (Poster) | 3 |
| 2019 | HHMF: hidden hierarchical matrix factorization for recommender systems
Hui Li 0057, Yu Liu 0066, Yuqiu Qian, Nikos Mamoulis, Wenting Tu, David Wai-Lok Cheung |
Data Min. Knowl. Discov. | 3 |
| 2018 | Maximizing Social Influence for the Awareness Threshold Model
Haiqi Sun, Reynold Cheng, Xiaokui Xiao, Yudian Zheng, Yuqiu Qian |
DASFAA (1) | 6 |
| 2018 | DSANLS: Accelerating Distributed Nonnegative Matrix Factorization via SketchingabstractNonnegative matrix factorization (NMF) has been successfully applied in different fields, such as text mining, image processing, and video analysis. NMF is the problem of determining two nonnegative low rank matrices U and V, for a given input matrix M, such that m ≈ UV⊥. There is an increasing interest in parallel and distributed NMF algorithms, due to the high cost of centralized NMF on large matrices. In this paper, we propose a distributed sketched alternating nonnegative least squares(DSANLS) framework for NMF, which utilizes a matrix sketching technique to reduce the size of nonnegative least squares subproblems in each iteration for U and V. We design and analyze two different random matrix generation techniques and two subproblem solvers. Our theoretical analysis shows that DSANLS converges to the stationary point of the original NMF problem and it greatly reduces the computational cost in each subproblem as well as the communication cost within the cluster. DSANLS is implemented using MPI for communication, and tested on both dense and sparse real datasets. The results demonstrate the efficiency and scalability of our framework, compared to the state-of-art distributed NMF MPI implementation. Yuqiu Qian, Conghui Tan, Nikos Mamoulis, David Wai-Lok Cheung |
WSDM | 1 |
| 2018 | Location-aware query reformulation for search engines
Zhipeng Huang 0001, Yuqiu Qian, Nikos Mamoulis |
GeoInformatica | 2 |
| 2017 | Reverse k-Ranks Queries on Large Graphs
Yuqiu Qian, Hui Li 0057, Nikos Mamoulis, Yu Liu 0066, David Wai-Lok Cheung |
EDBT | 1 |
| 2017 | P-LAG: Location-Aware Group Recommendation for Passive Users
Yuqiu Qian, Nikos Mamoulis, David Wai-Lok Cheung |
SSTD | 1 |
| 2017 | Microscopic and Macroscopic Spatio-Temporal Topic Models for Check-in DataabstractTwitter, together with other online social networks, such as Facebook, and Gowalla have begun to collect hundreds of millions of check-ins. Check-in data captures the spatial and temporal information of user movements and interests. To model and analyze the spatio-temporal aspect of check-in data and discover temporal topics and regions, we first propose a spatio-temporal topic model, i.e., Upstream Spatio-Temporal Topic Model (USTTM). USTTM can discover temporal topics and regions, i.e., a user's choice of region and topic is affected by time in this model. We use continuous time to model check-in data, rather than discretized time, avoiding the loss of information through discretization. In addition, USTTM captures the property that user's interests and activity space will change overtime, and users have different region and topic distributions at different times in USTTM. However, both USTTM and other related models capture “microscopic patterns” within a single city, where users share POIs, and cannot discover “macroscopic” patterns in a global area, where users check-in to different POIs. Therefore, we also propose a macroscopic spatio-temporal topic model, MSTTM, employing words of tweets that are shared between cities to learn the topics of user interests. We perform an experimental evaluation on Twitter and Gowalla data sets from New York City and on a Twitter US data set. In our qualitative analysis, we perform experiments with USTTM to discover temporal topics, e.g., how topic “tourist destinations” changes over time, and to demonstrate that MSTTM indeed discovers macroscopic, generic topics. In our quantitative analysis, we evaluate the effectiveness of USTTM in terms of perplexity, accuracy of POI recommendation, and accuracy of user and time prediction. Our results show that the proposed USTTM achieves better performance than the state-of-the-art models, confirming that it is more natural to model time as an upstream variable affecting the other variables. Finally, the performance of the macroscopic model MSTTM is evaluated on a Twitter US dataset, demonstrating a substantial improvement of POI recommendation accuracy compared to the microscopic models. Yu Liu 0066, Martin Ester, Yuqiu Qian, Bo Hu 0012, David Wai-Lok Cheung |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Barzilai-Borwein Step Size for Stochastic Gradient DescentabstractOne of the major issues in stochastic gradient descent (SGD) methods is how to choose an appropriate step size while running the algorithm. Since the traditional line search technique does not apply for stochastic optimization methods, the common practice in SGD is either to use a diminishing step size, or to tune a step size by hand, which can be time consuming in practice. In this paper, we propose to use the Barzilai-Borwein (BB) method to automatically compute step sizes for SGD and its variant: stochastic variance reduced gradient (SVRG) method, which leads to two algorithms: SGD-BB and SVRG-BB. We prove that SVRG-BB converges linearly for strongly convex objective functions. As a by-product, we prove the linear convergence result of SVRG with Option I proposed in [10], whose convergence result has been missing in the literature. Numerical experiments on standard data sets show that the performance of SGD-BB and SVRG-BB is comparable to and sometimes even better than SGD and SVRG with best-tuned step sizes, and is superior to some advanced SGD variants. Conghui Tan, Shiqian Ma, Yu-Hong Dai, Yuqiu Qian |
NIPS | 4 |
| 2014 | Wear-Leveling Optimization of Android YAFFS2 File System for NAND Based Embedded Devices
Yuqiu Qian |
WASA | 1 |