VLDB 2026 Research / reviewers in the wild / expert
Shan Qu
dblp:218/2613
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power Echoes: Investigating Moderation Biases in Online Power-Asymmetric ConflictsabstractOnline power-asymmetric conflicts are prevalent, and most platforms rely on human moderators to conduct moderation currently. Previous studies have been continuously focusing on investigating human moderation biases in different scenarios, while moderation biases under power-asymmetric conflicts remain unexplored. Therefore, we aim to investigate the types of power-related biases human moderators exhibit in power-asymmetric conflict moderation (RQ1) and further explore the influence of AI’s suggestions on these biases (RQ2). For this goal, we conducted a mixed design experiment with 50 participants by leveraging the real conflicts between consumers and merchants as a scenario. Results suggest several biases towards supporting the powerful party within these two moderation modes. AI assistance alleviates most biases of human moderation, but also amplifies a few. Based on these results, we propose several insights into future research on human moderation and human-AI collaborative moderation systems for power-asymmetric conflicts. Yaqiong Li, Peng Zhang 0060, Peixu Hou, Kainan Tu, Guangping Zhang, Shan Qu, Wenshi Chen, Yan Chen 0033, Ning Gu 0001, Tun Lu |
CHI | 6 |
| 2025 | Conformal Depression PredictionabstractWhile existing depression prediction methods based on deep learning show promise, their practical application is hindered by the lack of trustworthiness, as these deep models are often deployed asblack boxmodels, leaving us uncertain on the confidence of their predictions. For high-risk clinical applications like depression prediction, uncertainty quantification is essential in decision-making. In this paper, we introduce conformal depression prediction (CDP), a depression prediction method with uncertainty quantification based on conformal prediction (CP), giving valid confidence intervals with theoretical coverage guarantees for the model predictions. CDP is a plug-and-play module that requires neither model retraining nor an assumption about the depression data distribution. As CDP provides only an average coverage guarantee across all inputs rather than per-input performance guarantee, we further propose CDP-ACC, an improved conformal prediction with approximate conditional coverage. CDP-ACC firstly estimates the prediction distribution through neighborhood relaxation, and then introduces a conformal score function by constructing nested sequences, so as to provide a tighter prediction interval adaptive to specific input. We empirically demonstrate the application of CDP in uncertainty-aware facial depression prediction, as well as the effectiveness and superiority of CDP-ACC on the AVEC 2013 and AVEC 2014 datasets. Shan Qu, Xiuzhuang Zhou |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | On Social Network De-Anonymization With Communities: A Maximum A Posteriori PerspectiveabstractA crucial privacy-driven issue nowadays is re-identifying anonymized social networks by mapping them to correlated cross-domain auxiliary networks. Prior works are typically based on modeling social networks as random graphs representing users and their relations, and subsequently quantify the quality of mappings through varied cost functions. However, many cost functions are empirically proposed without sufficient theoretical support. For some other works probing the theoretical bound, it remains unknown how to algorithmically meet the demand of such quantifications, i.e., to minimize the cost functions. Besides, only few prior works have discussed the de-anonymization of social networks with communities. We address those concerns in a social network modeling parameterized by community structures that can be leveraged as side information for de-anonymization. Based on the Maximum A Posteriori (MAP) estimation, our first contribution is a series of MAP-based cost functions, which, when minimized, enjoy superiority to previous ones in finding the correct mapping with the highest probability. The feasibility of the cost functions is then for the first time algorithmically characterized. We prove the general multiplicative inapproximability and thus propose two heuristics, which, respectively, enjoy an$\epsilon$-additive approximation and a conditional optimality in carrying out successful user re-identification. Our theoretical findings are also empirically validated under classical synthetic and real-wrold social networks. Both theoretical and empirical observations manifest the importance of community in enhancing privacy inferencing. Jiapeng Zhang 0001, Shan Qu, Huquan Kang, Luoyi Fu, Haisong Zhang, Xinbing Wang, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Tracing Truth and Rumor Diffusions Over Mobile Social Networks: Who are the Initiators?abstractWith the increasing popularity of mobile devices, each user is able to conveniently acquire messages from others, and share diverse forms of information, like texts, images, or videos through online mobile apps. The full freedom of speech makes a great amount of truth (i.e., true information) and rumor (i.e., false information) propagate rapidly in a hybrid way through mobile platforms. As a huge variety of information floods pouring over us each day, identifying the authenticity of massive events becomes a necessary task to maintain the stability of Mobile Social Networks (MSNs). An important way to realize it is to trace their diffusions and make judgements according to the reliability of sources. With this regard, this paper proposes a diffusion model that characterizes the simultaneous diffusion of both truth and rumor in realistic MSNs, and makes the first attempt to figure out their respective sources. The problem of interest can be stated as: Given an outcome of cascade of both truth and rumor in MSNs, i.e., a set of nodes that might be the ignorant, the spreader of truth or rumor, or simply the silent receiver, how can we infer both truth sources and rumor sources? Different from previous sources detection works considering single type of nodes, the interplay between truth diffusions and rumor diffusions makes the conventional methods not work. To answer this question, we aim to maximize thesimilarity index, i.e., the number of nodes possessing the same states between the resulting network triggered by our estimated sources with the proposed diffusion model and the given observation network. Compared with existing techniques to trace diffusions of truth or rumor, it is much harder to find two kinds of sets at the same time, including truth sources and rumor sources, due to two primary reasons: (i) our biset optimization makes the submodularity techniques fail; (ii) our objective function is proven to be non-bisubmodular. To overcome above limitations, we first convert the objectivesimilarity indexinto a bisubmodular function by virtue of set covering. Based on this, we propose an approximation algorithm called Truth and Rumor Sources Detection (TRSD) algorithm via multiple reverse samplings with a provable$\frac{1}{4(1+\epsilon)^2}$approximation ratio. Further, a novel “time reversal” sources optimization strategy is proposed to converge the number of output sources from TRSD to a steady state. The effectiveness of our models and algorithms are empirical validated in two various datasets, from which we observe an up to 15% ofsimilarity indexgain as well as a narrowed down gap 0.6% to the ground truth. Shan Qu, Hui Xu 0011, Luoyi Fu, Huan Long, Xinbing Wang, Guihai Chen, Chenghu Zhou |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Facial Depression Recognition by Deep Joint Label Distribution and Metric LearningabstractWhile existing prediction models built on popular deep architectures have shown promising results in facial depression recognition, they still lack sufficient discriminative power due to the issues of 1) limited amount of labeled depression data for deep representation learning and, 2) large variation in facial expression across different persons of the same depression score and the subtle difference in facial expression across different depression levels. In this article, we formulate the facial depression recognition as a label distribution learning (LDL) problem, and propose a deep joint label distribution and metric learning (DJ-LDML) method to address these issues. In DJ-LDML, LDL exploits label relevance inherent in depression data to implicitly increase the amount of training data associated with each depression level without actually enlarging the dataset, while deep metric learning (DML) aims at learning a deep ordinal embedding with a specifically designed label-aware histogram loss, allowing semantics similarity between video sequences (described by ordinal labels) to be preserved for discriminative feature learning. The two learning modules in our DJ-LDML work collaboratively to enhance the representation ability and discriminative power of the deeply learned spatiotemporal feature, leading to improved depression prediction. We empirically evaluate our method on two benchmark datasets and the results demonstrate the effectiveness of our formulation. Xiuzhuang Zhou, Zeqiang Wei, Min Xu 0003, Shan Qu, Guodong Guo |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Directional Total Variation Regularized High-Resolution Prestack AVA InversionabstractPrestack seismic inversion has emerged as a powerful technique for reconstructing parameters attribute to the subsurface properties and building the geophysical parameter models. However, the inversion algorithms always suffer from spatial blur and low resolution. Total variation (TV) regularization preserves the spatial variation boundary of data by highlighting the sparsity of the first-order difference, which is regarded as an important technical means for image restoration. However, when the data do not change along the spatial grid direction, TV regularization is prone to a staircase effect. In this article, a directional TV (DTV) method is proposed to conduct the prestack amplitude variation with offset/angle (AVO/AVA) inversion. The method consists of three essential steps: estimating the seismic slope attribute from the seismic data, introducing seismic slope attribute to the TV regularization to establish the objective function, and optimizing the objective function by the split-Bregman algorithm. Finally, the conventional and proposed methods are applied to the synthetic and the real seismic data. The comparison of different methods demonstrates that the proposed method is applicable to reveal the detailed subsurface models, alleviate the staircase effect or artifact substantially, and further upgrade the quality of prestack inversion results. Guangtan Huang, Xiaohong Chen 0003, Shan Qu, Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Measuring Social Network De-Anonymizability by Means of Morphism PropertyabstractAnonymization techniques have tranquilized current social network users in terms of privacy leakage, however, it does not radically prevent adversaries from de-anonymizing users, as they may map the users to an un-anonymized network. Till now, researchers share a common thread in such de-anonymization attack: unveiling conditions leading to successful de-anonymization under the chosen network model. However, it has not yet been well understand how the structural property in different network models intrinsically determines de-anonymizability. We address the above issue in this paper by making the two contributions: (i) We discover that the automorphic degree and homomorphic degree of social networks determine their de-anonymizability universally. The automorphic degree characterizes the distinguishability of the users in a network, while the homomorphic degree models the similarities of users between two networks. We conclude that a smaller automorphic degree and a larger homomorphic degree conduce to a higher de-anonymizability. Such model-independent phenomenon refreshes us with a latitudinal study as it generalizes the essential commonness of de-anonymization in different network models. (ii) We derive explicit parametric bounds of the de-anonymizability for three classic network models, showing that such bounds correspond well to our conclusion about morphism property. We then algorithmically and experimentally show that such theoretical results literally make sense to adversaries. Such longitudinal study, including welding theory, algorithm and validation, promises applicability of our results on morphism property in real cases. Luoyi Fu, Jiapeng Zhang 0001, Shan Qu, Huquan Kang, Xinbing Wang, Guihai Chen |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | MIERank: Co-ranking Individuals and Communities with Multiple Interactions in Evolving NetworksabstractRanking has significant applications in real life. It aims to evaluate the importance (or popularity) of two categories of objects, i.e., individuals and communities. Numerous efforts have been dedicated to these two types of rankings respectively. Instead, in this paper, we for the first time explore the co-ranking of both individuals and communities. Our insight lies in that co-ranking may enhance the mutual evaluation on both sides. To this end, we first establish an Evolving Coupled Graph that contains a series of smoothly weighted snapshots, each of which characterizes and couples the intricate interactions of both individuals and communities till a certain evolution time into a single graph. Then we propose an algorithm, called MIERank to implement the co-ranking of individuals and communities in the proposed evolving graph. The core idea of MIERank lies in a novel unbiased random walk, which, when sampling the interplay among nodes over different generation times, incorporates the preference knowledge of ranking by utilizing nodes' future actions. MIERank returns the co-ranking of both individuals and communities by iteratively alternating between their corresponding stationary probabilities of the unbiased random walk in a mutually-reinforcing manner. We prove the efficiency of MIERank in terms of its convergence, optimality and extensiblity. Our experiments on a big scholarly dataset of 606862 papers and 1215 fields further validate the superiority of MIERank with fast convergence and an up to 26% ranking accuracy gain compared with the separate counterparts. Shan Qu, Luoyi Fu, Xinbing Wang |
INFOCOM | 1 |
| 2021 | Seeking the Truth in a Decentralized MannerabstractIn networks where massive sources make observations of same entities, we intend to seek thetruth– the most trustworthy value of each entity from conflicting information claimed by multiple sources. Various methods are proposed for accurately inferring both source reliability and truths, yet relying heavily on centralized settings that incur tremendous overhead to source side. In this paper, we offer adecentralizeddesign of truth discovery task that can fit favorably to the environments with limited resources. Considering that sources forming the connected network and making individual observations, we undertake the joint maximum likelihood estimation (MLE) of truth and source reliability. Our decentralization framework simply allows each source to maintain local information exchange at a time, and computes very basic functions of data observations. To this end, we facilitate the decentralization by simplifying the MLE problem into optimizing an objective function. Upon the proof of NP-hardness, two proposed decentralized algorithms (exact and approximation) are decentralized and randomized via a combination of algorithms from their centralized counterparts that ensure performance guarantee. The derived time complexity features explicit data/network dependent terms, which leads to further acceleration in truth finding. Remarkably, in two well connected networks like random geometric and preferential attachment graphs, the accelerated approximation method enjoys logarithmic time complexity while preserving comparable accuracy to the centralized counterparts. The effectiveness of the proposed decentralizations are further empirically confirmed. Luoyi Fu, Jiasheng Xu, Shan Qu, Zhiying Xu, Xinbing Wang, Guihai Chen |
IEEE/ACM Trans. Netw. | 3 |
| 2020 | Joint Inference on Truth/Rumor and Their Sources in Social NetworksabstractIn the contemporary era of information explosion, we are often faced with the mixture of massive truth (true information) and rumor (false information) flooded over social networks. Under such circumstances, it is very essential to infer whether each claim (e.g., news, messages) is a truth or a rumor, and identify their sources, i.e., the users who initially spread those claims. While most prior arts have been dedicated to the two tasks respectively, this paper aims to offer the joint inference on truth/rumor and their sources. Our insight is that a joint inference can enhance the mutual performance on both sides.To this end, we propose a framework named SourceCR, which alternates between two modules, i.e., credibility-reliability training for truth/rumor inference and division-querying for source detection, in an iterative manner. To elaborate, the former module performs a simultaneous estimation of claim credibility and user reliability by virtue of an Expectation Maximization algorithm, which takes the source reliability outputted from the latter module as the initial input. Meanwhile, the latter module divides the network into two different subnetworks labeled via the claim credibility, and in each subnetwork launches source detection by applying querying of theoretical budget guarantee to the users selected via the estimated reliability from the former module. The proposed SourceCR is provably convergent, and algorithmic implementable with reasonable computational complexity. We empirically validate the effectiveness of the proposed framework in both synthetic and real datasets, where the joint inference leads to an up to 35% accuracy of credibility gain and 29% source detection rate gain compared with the separate counterparts. Shan Qu, Luoyi Fu, Xinbing Wang, Jun (Jim) Xu |
INFOCOM | 1 |
| 2018 | Multi-Rack Regenerating Codes for Hierarchical Distributed Storage SystemsabstractErasure codes provide higher reliability than replication for a same level of redundancy to store data in distributed storage systems, yet with more bandwidth overhead. Recently, regenerating codes are introduced, which significantly reduce the repair bandwidth by analyzing the fundamental tradeoff between storage capacity and repair bandwidth via the information flow graph. In reality, distributed storage systems with hierarchical structures are more common in data centers where data are organized in racks, and the cross-rack communication is more costly than the in-rack communication. Hence, in this paper, we introduce a class of codes to repair a failed node by downloading data from nodes in the same rack only, which are termed as multi-rack regenerating codes (MRC). Different with existing works, the cross-rack repair bandwidth under our codes can be reduced to zero. Meanwhile, we obtain the optimal tradeoff between storage and bandwidth of MRC, and present an explicit construction of MRC with the common product-matrix framework. Shan Qu, Jinbei Zhang, Haiwen Cao, Xinbing Wang |
ICC | 1 |
| 2018 | Asymmetric regenerating codes for heterogeneous distributed storage systemsabstractDistributed storage systems provide reliability by distributing data over multiple storage nodes. Once a node fails, a new node is introduced to the system to maintain the availability of the stored data. The new node downloads information from other surviving nodes called helper nodes to recover the lost data in the failed node. The number of helper nodes is called repair degree. Compared to traditional approaches, e.g., replication and erasure codes, the regenerating codes proposed recently can significantly reduce the repair bandwidth in homogeneous distributed storage systems. Most existing works focus on uniform settings (e.g., in terms of repair degree and repair bandwidth). However, due to network structures or connectivity limitations, for each failed node, the number of required helper nodes may be different for distinct failed nodes. Furthermore, considering the limits of network traffic of bandwidth, the amount of information allowed to be downloaded from each helper node could also vary. Thus we are motivated to investigate heterogeneous distributed storage systems where the repair degree and the amount of information downloaded from each helper node can be different. In order to obtain the minimal bandwidth to recover a failed node, we construct an information flow graph for such heterogeneous systems. By analyzing the cut-set bound of the information flow graph, the optimal tradeoff between storage capacity and repair bandwidth is derived. We then propose asymmetric regenerating codes that can achieve the curve of the optimal tradeoff. A linear construction of asymmetric regenerating codes is presented. Compared with previous regenerating codes, asymmetric regenerating codes are shown to have a lower repair bandwidth under a certain constraint condition, whose reduction can be up to 36.2%. Shan Qu, Jinbei Zhang, Xinbing Wang |
WiOpt | 1 |