Qiang Yan 0001

dblp:79/6531-1 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-7328-2278ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 13 · 4 first-authorDatabases, data management, data science and information retrieval · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Content- and Context-Aware Click Model Based on Dynamic Graph Neural Networks
abstract
Click modeling constitutes a pivotal area of study within information retrieval, as it provides insights into user search behavior and enables the extraction of valuable implicit relevance feedback from large-scale click logs. However, existing click models often rely on content-agnostic IDs to represent queries and documents. Additionally, they employ context-independent assumptions, such as the examination hypothesis, in modeling click probabilities. As a result, contemporary click models often fall short of capturing the influence of the diverse, multi-modal content found on modern Search Engine Result Pages (SERPs) and the intricate contextual interactions among heterogeneous search results. To address this issue, we propose a novel Dynamic Graph Neural Click Model (DGCM). The proposed model incorporates rich content and context information by jointly representing them as nodes in a dynamic graph and further leverages a dynamic graph attention network to predict users’ clicks at different timesteps. To demonstrate the effectiveness of the DGCM model, we conducted extensive experiments on two large-scale datasets with both content and click information: the public Sogou-SRR dataset and a proprietary dataset collected on the WeChat platform. The experimental results indicate that by capitalizing on the content and context information, DGCM outperforms existing click models in the click prediction and relevance estimation tasks.
Jiaxin Mao, Ziyuan Zhao, Qiang Yan 0001
ACM Trans. Inf. Syst.4
2026 Corrigendum: A Content- and Context-Aware Click Model Based on Dynamic Graph Neural Networks
abstract
This is a corrigendum for the article “A Content- and Context-Aware Click Model Based on Dynamic Graph Neural Networks” published in ACM Trans. Inf. Syst. 44, 3, Article 67 (March 2026), 27 pages.
Jiaxin Mao, Ziyuan Zhao, Qiang Yan 0001
ACM Trans. Inf. Syst.4
2025 Dynamic Service Knowledge Base Construction at WeChat
Haoyang Li 0002, Alexander Zhou 0001, Fengmei Jin, Qing Li 0001, Ziyuan Zhao, Hao Xin, Qiang Yan 0001, Tiezheng Mao, Xueling Lin, Zijian Li 0002, Lei Chen 0002
ADMA (4)7
2025 Addressing Personalized Bias for Unbiased Learning to Rank
abstract
Unbiased learning to rank (ULTR), which aims to learn unbiased ranking models from biased user behavior logs, plays an important role in Web search. Previous research on ULTR has studied a variety of biases in users' clicks, such as position bias, presentation bias, and outlier bias. However, existing work often assumes that the behavior logs are collected from an ''average'' user, neglecting the differences between different users in their search and browsing behaviors. In this paper, we introduce personalized factors into the ULTR framework, which we term the user-aware ULTR problem. Through a formal causal analysis of this problem, we demonstrate that existing user-oblivious methods are biased when different users have different preferences over queries and personalized propensities of examining documents. To address such a personalized bias, we propose a novel user-aware inverse-propensity-score estimator for learning-to-rank objectives. Specifically, our approach models the distribution of user browsing behaviors for each query and aggregates user-weighted examination probabilities to determine propensities. We theoretically prove that the user-aware estimator is unbiased under some mild assumptions and shows lower variance compared to the straightforward way of calculating a user-dependent propensity for each impression. Finally, we empirically verify the effectiveness of our user-aware estimator by conducting extensive experiments on two semi-synthetic datasets and a real-world dataset.
Zechun Niu, Lang Mei, Ziyuan Zhao, Qiang Yan 0001, Jiaxin Mao, Ji-Rong Wen
CIKM5
2023 K-DLM: A Domain-Adaptive Language Model Pre-Training Framework with Knowledge Graph
Jiaxin Zou, Zuotong Xie, Qiang Yan 0001, Hai-Tao Zheng 0002
ICANN (4)5
2023 Self-supervised Bidirectional Prompt Tuning for Entity-enhanced Pre-trained Language Model
abstract
With the promotion of the pre-training paradigm, researchers are increasingly focusing on injecting external knowledge, such as entities and triplets from knowledge graphs, into pre-trained language models (PTMs) to improve their understanding and logical reasoning abilities. This results in significant improvements in natural language understanding and generation tasks and some level of interpretability. In this paper, we propose a novel two-stage entity knowledge enhancement pipeline for Chinese pre-trained models based on “bidirectional” prompt tuning. The pipeline consists of a “forward” stage, in which we construct fine-grained entity type prompt templates to boost PTMs injected with entity knowledge, and a “backward” stage, where the trained templates are used to generate type-constrained context-dependent negative samples for contrastive learning. Experiments on six classification tasks in the Chinese Language Understanding Evaluation (CLUE) benchmark demonstrate that our approach significantly improves upon the baseline results in most datasets, particularly those that have a strong reliance on diverse and extensive knowledge.
Jiaxin Zou, Xianghong Xu 0001, Qiang Yan 0001, Hai-Tao Zheng 0002
IJCNN4
2023 A Self-Correcting Sequential Recommender
abstract
Sequential recommendations aim to capture users’ preferences from their historical interactions so as to predict the next item that they will interact with. Sequential recommendation methods usually assume that all items in a user’s historical interactions reflect her/his preferences and transition patterns between items. However, real-world interaction data is imperfect in that (i) users might erroneously click on items, i.e., so-called misclicks on irrelevant items, and (ii) users might miss items, i.e., unexposed relevant items due to inaccurate recommendations.
Yujie Lin 0001, Zhumin Chen, Zhaochun Ren, Xin Xin 0003, Qiang Yan 0001, Maarten de Rijke, Xiuzhen Cheng, Pengjie Ren
WWW6
2022 FL-Market: Trading Private Models in Federated Learning
abstract
Acquiring a sufficient amount of training data is a significant bottleneck for machine learning (ML) based data analytics. Recently, commoditizing ML models has been proposed as an economical and moderate solution to ML-oriented data acquisition. However, existing model marketplaces assume that the broker can access data owners’ private training data, which may not be realistic in practice. In this paper, to promote trustworthy data acquisition for ML tasks, we propose FL-Market, a locally private model marketplace that protects privacy against not only model buyers but also an untrusted broker. FL-Market decouples ML from the need to centrally gather training data on the broker’s side using federated learning, a privacy-preserving ML paradigm in which data owners collaboratively train an ML model by uploading local gradients (to be aggregated into a global gradient for model updating). Then, FL-Market enables data owners to locally perturb their gradients by local differential privacy and thus further prevents privacy risks. To drive FL-Market, we propose a deep learning-empowered auction mechanism for intelligently deciding the local gradients’ perturbation levels and an optimal aggregation mechanism for aggregating the perturbed gradients. Our auction and aggregation mechanisms can jointly maximize the global gradient’s accuracy, which optimizes model buyers’ utility. Our experiments verify the effectiveness of the proposed mechanisms.
Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa, Huizhong Li, Qiang Yan 0001
IEEE Big Data5
2022 MGAD: Learning Descriptional Representation Distilled from Distributional Semantics for Unseen Entities
abstract
Entity representation plays a central role in building effective entity retrieval models. Recent works propose to learn entity representations based on entity-centric contexts, which achieve SOTA performances on many tasks. However, these methods lead to poor representations for unseen entities since they rely on a multitude of occurrences for each entity to enable accurate representations. To address this issue, we propose to learn enhanced descriptional representations for unseen entities by distilling knowledge from distributional semantics into descriptional embeddings. Specifically, we infer enhanced embeddings for unseen entities based on descriptions by aligning the descriptional embedding space to the distributional embedding space with different granularities, i.e., element-level, batch-level and space-level alignment. Experimental results on four benchmark datasets show that our approach improves the performance over all baseline methods. In particular, our approach can achieve the effectiveness of the teacher model on almost all entities, and maintain such high performance on unseen entities.
Yuanzheng Wang, Xueqi Cheng 0001, Yixing Fan, Xiaofei Zhu, Huasheng Liang, Qiang Yan 0001, Jiafeng Guo
IJCAI6
2022 Accelerating Elliptic Curve Digital Signature Algorithms on GPUs
abstract
The Elliptic Curve Digital Signature Algorithm (ECDSA) is an essential building block of various cryptographic protocols. In particular, most blockchain systems adopt it to ensure transaction integrity. However, due to its high computational intensity, ECDSA is often the performance bottleneck in blockchain transaction processing. Recent work has accelerated ECDSA algorithms on the CPU; in contrast, success has been limited on the GPU, which has great potential for parallelization but is challenging for implementing elliptic curve functions. In this paper, we propose RapidEC, a GPU-based ECDSA implementation for SM2, a popular elliptic curve. Specifically, we design architecture-aware parallel primitives for elliptic curve point operations, and parallelize the processing of a single SM2 request as well as batches of requests. Consequently, our GPU-based RapidEC outperformed the state-of-the-art CPU-based algorithm by orders of magnitude. Additionally, our GPU-based modular arithmetic functions as well as point operation primitives can be applied to other computation tasks.
Zonghao Feng, Qipeng Xie, Qiong Luo 0001, Yujie Chen 0007, Huizhong Li, Qiang Yan 0001
SC7
2022 Efficient Two-stage Label Noise Reduction for Retrieval-based Tasks
abstract
The existence of noisy labels in datasets has always been an essential dilemma in deep learning studies. Previous works detected noisy labels by analyzing the predicted probability distribution generated by the model trained on the same data and calculating the probabilities of each label to be regarded as noise. However, the predicted probability distribution from the whole dataset may introduce overfitting, and the overfitting on noisy labels may induce the probability distribution of clean and noisy items to be not conditional independent, making identification more challenging. Additionally, label noise reduction on image datasets has received much attention, while label noise reduction on text datasets has not. This paper proposes a noisy label reduction method for text datasets, which could be applied at retrieval-based tasks by getting a conditional independent probability distribution to identify noisy labels accurately. The method first generates a candidate set containing noisy labels, predicts the category probabilities by the model trained on the rest cleaner data, and then identifies noisy items by analyzing a confidence matrix. Moreover, we introduce a warm-up module and a sharpened cross-entropy loss function for efficiently training in the first stage. Empirical results on different rates of uniform and random label noise in five text datasets demonstrate that our method can improve the label noise reduction accuracy and end-to-end classification accuracy. Further, we find that the iteration of the label noise reduction method is efficient to high-rate label noise datasets, and our method will not hurt clean datasets too much.
Mengmeng Kuang, Weiyan Wang, Lie Kang, Qiang Yan 0001
WSDM5
2022 DVREI: Dynamic Verifiable Retrieval Over Encrypted Images
abstract
The increasing awareness in privacy has partly contributed to the renewed interest in privacy-preserving encrypted image retrieval, and designing for outsourced images stored on cloud servers, etc. However, there are some limitations in these existing schemes such as low retrieval accuracy, low retrieval efficiency, and less efficient result verification in the dynamic setting. Therefore, in this paper we present a novel Dynamic Verifiable Retrieval over Encrypted Images (DVREI) scheme. First, a pre-trained Convolutional Neural Network (CNN) model is utilized to extract image features to improve retrieval accuracy. Then, an encrypted index based on the K-means clustering algorithm is designed to improve retrieval efficiency. Finally, a dynamic verification tree based on the chameleon hash is used to verify the correctness of the retrieval results and support dynamic updates. We theoretically and experimentally evaluate the security and performance of DVREI to demonstrate its practicability.
Yingying Li 0001, Jianfeng Ma 0001, Yinbin Miao, Huizhong Li, Qiang Yan 0001, Yue Wang 0063, Ximeng Liu, Kim-Kwang Raymond Choo
IEEE Trans. Computers5
2020 TransN: Heterogeneous Network Representation Learning by Translating Node Embeddings
abstract
Learning network embeddings has attracted growing attention in recent years. However, most of the existing methods focus on homogeneous networks, which cannot capture the important type information in heterogeneous networks. To address this problem, in this paper, we propose TransN, a novel multi-view network embedding framework for heterogeneous networks. Compared with the existing methods, TransN is an unsupervised framework which does not require node labels or user-specified meta-paths as inputs. In addition, TransN is capable of handling more general types of heterogeneous networks than the previous works. Specifically, in our framework TransN, we propose a novel algorithm to capture the proximity information inside each single view. Moreover, to transfer the learned information across views, we propose an algorithm to translate the node embeddings between different views based on the dual-learning mechanism, which can both capture the complex relations between node embeddings in different views, and preserve the proximity information inside each view during the translation. We conduct extensive experiments on real-world heterogeneous networks, whose results demonstrate that the node embeddings generated by TransN outperform those of competitors in various network mining tasks.
Zijian Li 0002, Wenhao Zheng 0001, Xueling Lin, Ziyuan Zhao, Zhe Wang 0019, Yue Wang 0012, Xun Jian 0001, Lei Chen 0002, Qiang Yan 0001, Tiezheng Mao
ICDE9
2020 Communication-Efficient Collaborative Learning of Geo-Distributed JointCloud from Heterogeneous Datasets
abstract
With the popularity of cloud computing, the use of services provided by the cloud is increasing. Due to the high communication cost, and data privacy issues, a new crosscloud collaborative computing model is demanded instead of a single giant cloud. Coping with federated learning and Joint-Cloud, we propose a federated learning-based collaborative learning framework, in which the distributed cloud entities are able to learn the same model collaboratively. As compared to traditional cloud-centric approaches, the framework for JointCloud can reduce the network bandwidth overhead and guarantee privacy. However, there are two crucial challenges in the federated manner: heterogeneity and high communication overhead. To address the heterogeneity, we propose a Teacher-Student mechanism, the key of which is a regularization term incorporated with the objective function so as to adjust the gradients from the clients among JointCloud with different data distribution. Then, based on the Teacher-Student mechanism, we further present a communication-efficient federated optimation approach via joint Identification-Verification to reduce the communication rounds. We conduct extensive experiments on Non-IID datasets. The experimental results demonstrate that the proposed framework can significantly reduce communication costs and improve performance.
Xiaoli Li 0016, Chuan Chen 0001, Zibin Zheng, Huizhong Li, Qiang Yan 0001
JCC6
2020 ModCon: a model-based testing platform for smart contracts
abstract
Unlike those on public permissionless blockchains, smart contracts on enterprise permissioned blockchains are not limited by resource constraints, and therefore often larger and more complex. Current testing and analysis tools lack support for such contracts, which demonstrate stateful behaviors and require special treatment in quality assurance. In this paper, we present a model-based testing platform, called ModCon, relying on user-specified models to define test oracles, guide test generation, and measure test adequacy. ModCon is Web-based and supports both permissionless and permissioned blockchain platforms. We demonstrate the usage and key features of ModCon on real enterprise smart contract applications.
Ye Liu 0012, Yi Li 0008, Shangwei Lin 0001, Qiang Yan 0001
ESEC/SIGSOFT FSE4
2019 Blockchain-Based Credible and Privacy-Preserving QoS-Aware Web Service Recommendation
Xiaoli Li 0016, Erxin Du, Chuan Chen 0001, Zibin Zheng, Ting Cai 0002, Qiang Yan 0001
BlockSys6
2018 Empirical Study of Face Authentication Systems Under OSNFD Attacks
abstract
Face authentication has been widely available on smartphones, tablets, and laptops. As numerous personal images are published in online social networks (OSNs), OSN-based facial disclosure (OSNFD) creates significant threat against face authentication. We make the first attempt to quantitatively measure OSNFD threat to real-world face authentication systems on smartphones, tablets, and laptops. Our results show that the percentage of vulnerable users that are subject to spoofing attacks is high, which is about 64 percent for laptop users, and 93 percent smartphone/tablet users. We investigate liveness detection methods in the real-world face authentication systems against OSNFD threat. We discover that under protection of liveness detection, the percentage of vulnerable images is 18.8 percent, but the percentage of vulnerable users is as high as 73.3 percent. This evidence suggests that the current face authentication systems are not strong enough under OSNFD attacks. Finally, we develop a risk estimation tool based on logistic regression, and analyze the impacts of key attributes of facial images on the OSNFD risk. Our statistical analysis reveals that the most influential attributes of facial images are image resolution, facial makeup, occluded eyes, and illumination. This tool can be used to evaluate OSNFD risk for OSN images to increase users' awareness of OSNFD.
Yan Li 0075, Yingjiu Li, Qiang Yan 0001, Robert H. Deng
IEEE Trans. Dependable Secur. Comput.4
2015 Seeing Your Face Is Not Enough: An Inertial Sensor-Based Liveness Detection for Face Authentication
abstract
Leveraging built-in cameras on smartphones and tablets, face authentication provides an attractive alternative of legacy passwords due to its memory-less authentication process. However, it has an intrinsic vulnerability against the media-based facial forgery (MFF) where adversaries use photos/videos containing victims' faces to circumvent face authentication systems. In this paper, we propose FaceLive, a practical and robust liveness detection mechanism to strengthen the face authentication on mobile devices in fighting the MFF-based attacks. FaceLive detects the MFF-based attacks by measuring the consistency between device movement data from the inertial sensors and the head pose changes from the facial video captured by built-in camera. FaceLive is practical in the sense that it does not require any additional hardware but a generic front-facing camera, an accelerometer, and a gyroscope, which are pervasively available on today's mobile devices. FaceLive is robust to complex lighting conditions, which may introduce illuminations and lead to low accuracy in detecting important facial landmarks; it is also robust to a range of cumulative errors in detecting head pose changes during face authentication.
Yan Li 0075, Yingjiu Li, Qiang Yan 0001, Hancong Kong, Robert H. Deng
CCS3
2015 Privacy leakage analysis in online social networks
Yan Li 0075, Yingjiu Li, Qiang Yan 0001, Robert H. Deng
Comput. Secur.3
2015 Leakage-resilient password entry: Challenges, design, and evaluation
Qiang Yan 0001, Jin Han 0002, Yingjiu Li, Jianying Zhou 0001, Robert H. Deng
Comput. Secur.1
2014 Understanding OSN-based facial disclosure against face authentication systems
abstract
Face authentication is one of promising biometrics-based user authentication mechanisms that have been widely available in this era of mobile computing. With built-in camera capability on smart phones, tablets, and laptops, face authentication provides an attractive alternative of legacy passwords for its memory-less authentication process. Although it has inherent vulnerability against spoofing attacks, it is generally considered sufficiently secure as an authentication factor for common access protection. However, this belief becomes questionable since image sharing has been popular in online social networks (OSNs). A huge number of personal images are shared every day and accessible to potential adversaries. This OSN-based facial disclosure (OSNFD) creates a significant threat against face authentication. In this paper, we make the first attempt to quantitatively measure the threat of OSNFD. We examine real-world face-authentication systems designed for both smartphones, tablets, and laptops. Interestingly, our results find that the percentage of vulnerable images that can used for spoofing attacks is moderate, but the percentage of vulnerable users that are subject to spoofing attacks is high. The difference between systems designed for smartphones/tablets and laptops is also significant. In our user study, the average percentage of vulnerable users is 64% for laptop-based systems, and 93% for smartphone/tablet-based systems. This evidence suggests that face authentication may not be suitable to use as an authentication factor, as its confidentiality has been significantly compromised due to OSNFD. In order to understand more detailed characteristics of OSNFD, we further develop a risk estimation tool based on logistic regression to extract key attributes affecting the success rate of spoofing attacks. The OSN users can use this tool to calculate risk scores for their shared images so as to increase their awareness of OSNFD.
Yan Li 0075, Qiang Yan 0001, Yingjiu Li, Robert H. Deng
AsiaCCS3
2014 Towards semantically secure outsourcing of association rule mining on categorical data
Junzuo Lai, Yingjiu Li, Robert H. Deng, Jian Weng 0001, Chaowen Guan, Qiang Yan 0001
Inf. Sci.6
2013 Launching Generic Attacks on iOS with Approved Third-Party Applications
Jin Han 0002, Su Mon Kywe, Qiang Yan 0001, Feng Bao 0001, Robert H. Deng, Debin Gao, Yingjiu Li, Jianying Zhou 0001
ACNS3
2013 Designing leakage-resilient password entry on touchscreen mobile devices
abstract
Touchscreen mobile devices are becoming commodities as the wide adoption of pervasive computing. These devices allow users to access various services at anytime and anywhere. In order to prevent unauthorized access to these services, passwords have been pervasively used in user authentication. However, password-based authentication has intrinsic weakness in password leakage. This threat could be more serious on mobile devices, as mobile devices are widely used in public places.
Qiang Yan 0001, Jin Han 0002, Yingjiu Li, Jianying Zhou 0001, Robert H. Deng
AsiaCCS1
2013 Comparing Mobile Privacy Protection through Cross-Platform Applications
Jin Han 0002, Qiang Yan 0001, Debin Gao, Jianying Zhou 0001, Robert H. Deng
NDSS2
2013 Think Twice before You Share: Analyzing Privacy Leakage under Privacy Control in Online Social Networks
Yan Li 0075, Yingjiu Li, Qiang Yan 0001, Robert H. Deng
NSS3
2012 On Limitations of Designing Leakage-Resilient Password Systems: Attacks, Principals and Usability
Qiang Yan 0001, Jin Han 0002, Yingjiu Li, Robert H. Deng
NDSS1
2011 A software-based root-of-trust primitive on multicore platforms
abstract
Software-based root-of-trust has been proposed to overcome the disadvantage of hardware-based root-of-trust, which is the high cost in deployment and upgrade (when vulnerabilities are discovered). However, prior research on software-based root-of-trust only focuses on uniprocessor platforms. The essential security properties of such software-based root-of-trust, as analyzed and demonstrated in our paper, can be violated on multicore platforms. Since multicore processors are becoming increasingly popular, it is imperative to explore the feasibility of software-based root-of-trust on them.
Qiang Yan 0001, Jin Han 0002, Yingjiu Li, Robert H. Deng, Tieyan Li
AsiaCCS1
2011 On Detection of Erratic Arguments
Jin Han 0002, Qiang Yan 0001, Robert H. Deng, Debin Gao
SecureComm2