EDBT 2026 Demo / reviewers in the wild / expert
Qingyuan Gong
dblp:146/0648
· DBLP profile ↗
20ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-7942-8752ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Computer networks · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Large-Scale Dataset of Interactions Between Weibo Users and Platform-Empowered LLM AgentabstractWe release a large-scale dataset that captures interactions between human users and CommentRobert, an LLM-based social media agent on Weibo. The dataset contains Weibo posts in which users actively mention the LLM agent account @CommentRobert, indicating that the users are interested in interacting with the platform-empowered LLM agent. The dataset contains 557,645 interactions from 304,400 unique users over 17 months. We detail our data collection methodology, user attributes, and content characteristics, underscoring the dataset's value in examining real-world human-LLM agent interactions. Our analysis offers insights into the demographic and behavioral traits of users interested in the selected LLM agent, interaction dynamics between humans and the agent, and linguistic patterns in comments. These interactions provide a unique lens through which to explore how humans perceive, trust, and communicate with LLMs. This dataset enables further research into modeling human intent understanding, improving LLM agent design, and studying the evolution of human-LLM agent relationships. Potential applications also include long-term user engagement prediction and AI-generated comment detection on social platforms. This constructed dataset is available at https://zenodo.org/records/16921462. Shaokui Gu, Qingyuan Gong, Fenghua Tong, Yipeng Zhou, Qiang Duan 0002, Yang Chen 0001 |
CIKM | 3 |
| 2025 | On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic WritingabstractThe rising popularity of large language models (LLMs) has raised concerns about potential abuse and harmful content. As a result, developing a highly generalizable and adaptable machine-generated text (MGT) detection system has become an urgent priority. Given that LLMs are most commonly misused in academic writing, this work investigates the generalization and adaptation capabilities of MGT detectors in three key aspects specific to academic writing: First, we construct MGT-Academic, a large-scale dataset comprising over 336M tokens and 749K samples. MGT-Academic focuses on academic writing, featuring human-written texts (HWTs) and MGTs across STEM, Humanities, and Social Sciences, paired with an extensible code framework for efficient benchmarking. Second, we benchmark the performance of various detectors for binary classification and text attribution tasks in both in-domain and cross-domain settings. This benchmark reveals the often-overlooked challenges of text attribution tasks. Third, we introduce a novel text attribution task in which models must adapt to new classes over time, with little or no access to prior training data, spanning both few-shot and many-shot scenarios. We implement a range of adaptation techniques to enhance performance across these settings. Our findings provide new insights into the generalization ability of MGT detectors and lay the foundation for building robust, adaptive detection systems. The code framework is available at https://github.com/Y-L-LIU/MGTBench-2.0. Yule Liu, Zhiyuan Zhong, Zhen Sun 0001, Jingyi Zheng, Jiaheng Wei, Qingyuan Gong, Fenghua Tong, Yang Chen 0001, Yang Zhang 0016, Xinlei He 0001 |
KDD (2) | 7 |
| 2025 | Prototypical clustered federated learning for heart rate predictionabstractPredicting future heart rate (HR) not only helps in detecting abnormal heart rhythms but also provides timely support for downstream health monitoring services. Existing methods for HR prediction encounter challenges, especially concerning privacy protection and data heterogeneity. To address these challenges, this paper proposes a novel HR prediction framework, PCFedH, which leverages personalized federated learning and prototypical contrastive learning to achieve stable clustering results and more accurate predictions. PCFedH contains two core modules: a prototypical contrastive learning-based federated clustering module, which characterizes data heterogeneity and enhances HR representation to facilitate more effective clustering, and a two-phase soft clustered federated learning module, which enables personalized performance improvements for each local model based on stable clustering results. Experimental results on two real-world datasets demonstrate the superiority of our approach over state-of-the-art methods, achieving an average reduction of 3.1% in the mean squared error across both datasets. Additionally, we conduct comprehensive experiments to empirically validate the effectiveness of the key components in the proposed method. Among these, the personalization component is identified as the most crucial aspect of our design, indicating its substantial impact on overall performance. Hui Ruan, Yang Chen 0001, Jiong Chen 0003, Ziyue Li 0002, Xiang Su 0001, Yipeng Zhou, Qingyuan Gong |
Frontiers Inf. Technol. Electron. Eng. | 8 |
| 2025 | DeepHole: Identifying Structural Hole Spanners in Online Social Networks Using Behavior EmbeddingabstractInfluential users in online social networks have been given a lot of attention, which could maximize information diffusion in a centralized dissemination process. However, there exist users that can significantly promote diversified information communication among users, which is a pluralistic interaction process. Structural hole (SH) spanners serve as a bridge between social groups, promoting interactions between users and allowing ordinary users to receive diverse information. Filling structural holes, SH spanners have more opportunities to obtain innovative content. Identifying SH spanners is challenging due to their inconspicuous characteristics and privacy policy restrictions. In this article, we are the first to study the SH spanner identification problem using the behavior data generated by users, releasing the dependency on the entire graph structure to find SH spanners. We propose DeepHole to design both the semantic and the sequential behavior analysis modules on users’ generated content, utilizing the TextCNN and Transformer models, respectively. The semantic and sequential embeddings are effective in distinguishing the SH spanners from ordinary users. We conduct comprehensive evaluations using the Yelp and Foursquare datasets simultaneously. Results show that DeepHole can achieve a satisfying performance in detecting SH spanners labeled by both constraint and effective size metrics, with AUC values of 0.923 and 0.933 for Yelp and 0.828 and 0.815 for Foursquare. Qingyuan Gong, Shaokui Gu, Xin Wang 0002, Pan Hui 0001, Jar-der Luo, Xiaoming Fu 0001, Yang Chen 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Understanding Work Rhythms in Software Development and Their Effects on Technical PerformanceabstractThe temporal patterns of code submissions, denoted as work rhythms, provide valuable insight into the work habits and productivity in software development. In this paper, we investigate the work rhythms in software development and their effects on technical performance by analyzing the profiles of developers and projects from 110 international organizations and their commit activities on GitHub. Using clustering, we identify four work rhythms among individual developers and three work rhythms among software projects. Strong correlations are found between work rhythms and work regions, seniority, and collaboration roles. We then define practical measures for technical performance and examine the effects of different work rhythms on them. Our findings suggest that moderate overtime is related to good technical performance, whereas fixed office hours are associated with receiving less attention. Furthermore, we survey 92 developers to understand their experience with working overtime and the reasons behind it. The survey reveals that developers often work longer than required. A positive attitude towards extended working hours is associated with situations that require addressing unexpected issues or when clear incentives are provided. In addition to the insights from our quantitative and qualitative studies, this work sheds light on tangible measures for both software companies and individual developers to improve the recruitment process, project planning, and productivity assessment. Jiayun Zhang, Qingyuan Gong, Yang Chen 0001, Yu Xiao 0001, Xin Wang 0002, Aaron Yi Ding |
IET Softw. | 2 |
| 2023 | Poster: A Privacy-preserving Heart Rate Prediction System for Drivers in Connected VehiclesabstractThe prediction of health metrics for drivers has become increasingly crucial due to the potential impact of drivers' health conditions on traffic accidents. Heart attack is one of the primary causes of health-related traffic tragedies. However, drivers' heart rate (HR) is considered highly-private data, which should not be collected by a centralized server for training the prediction model. To this end, we contribute FedHeart, a novel privacy-preserving federated learning (FL) system for HR prediction. We observe distinct HR changes when drivers are in steady-state and changing-state conditions, and thus we utilize FL to train two separate models for these states. To enhance the prediction accuracy, we incorporate contrastive learning to extract HR features. Through experiments on two real-world datasets, we validate the efficiency of the proposed system in accurately predicting HR during driving scenarios. Hui Ruan, Qingyuan Gong, Yang Chen 0001, Jiong Chen 0003, Ziyue Li 0002, Xiang Su 0001 |
MobiSys | 2 |
| 2023 | FloodSFCP: Quality and Latency Balanced Service Function Chain Placement for Remote Sensing in LEO Satellite NetworkabstractPrompted by the significant advancements in image processing technologies and their diverse range of applications, remote sensing satellites are poised for rapid expansion. Nonetheless, offloading the vast amount of remote sensing satellite images to the ground gateway station is inefficient due to the exorbitant costs induced by satellite links, while the limited resources of individual satellites hinder local task processing. With the advancement of the network function virtualization (NFV) technology, a new paradigm for service function chain (SFC) has emerged, which can significantly improve the flexibility and resource utilization of network services and alleviate resource conflicts by dividing large services into smaller ones organized in the form of SFCs. As mega-constellations (e.g., Starlink) developed, the number of low earth orbit (LEO) satellites is increasing. By dividing services into small sub-services and organizing them into SFCs throughout the LEO network, services that cannot be completed by a single satellite can be accomplished through multi-satellite cooperation. However, the quality of the remote sensing service is positively correlated with its latency, and the rapidly changing topology of LEO networks also adds complexity to the SFC placement. Hence, how to select appropriate satellites to place the SFC and modulate service levels, in order to obtain better remote sensing results within an acceptable latency, remains a question. To address these issues, this paper proposes the FloodSFCP, an SFC placement method that aims to increase service quality and decrease latency through offline training and online optimization via deep reinforcement learning, taking into account the variation in LEO network topology. By introducing NoisyNet, Dueling, and N-step learning, we improve the model’s generalization ability and reduce the state space, thus enhancing convergence speed while reducing decision and training time. Experimental results demonstrate that FloodSFCP significantly improves service quality while reducing total decision costs. Ruoyi Zhang, Chao Zhu 0002, Xiao Chen 0002, Qingyuan Gong, Xinlei Xie, Xiangyuan Bu |
SECON | 4 |
| 2023 | Detecting Malicious Accounts in Online Developer Communities Using Deep LearningabstractOnline developer communities like GitHub allow a massive number of developers to collaborate. However, the openness of the communities makes them vulnerable to different types of malicious attacks, since attackers can easily join these communities and interact with legitimate users. In this work, we propose GitSec, a deep learning-based solution for detecting malicious accounts in online developer communities. GitSec distinguishes malicious accounts from legitimate ones based on the account profiles, dynamic activity characteristics, as well as social interactions. First, GitSec introduces two user activity sequences and applies a parallel neural network design with an attention mechanism to process the sequences. Second, GitSec constructs two graphs to represent the interactions between users according to their repository operations. Especially, graph neural networks and structural hole theory are employed to deal with the two constructed graphs. Third, GitSec makes use of the descriptive features to enhance the detection performance. The final judgement is made by a decision maker implemented by a supervised machine learning-based classifier. Based on the real-world data of GitHub users, our comprehensive evaluations show that GitSec achieves a better performance than state-of-the-art solutions, with an AUC value of 0.916. Qingyuan Gong, Jiayun Zhang, Yang Chen 0001, Qi Li 0002, Yu Xiao 0001, Xin Wang 0002, Pan Hui 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | DeepPick: A Deep Learning Approach to Unveil Outstanding Users With Public Attainable FeaturesabstractOutstanding users (OUs) denote the influential, "core" or "bridge" users in the online community. How to accurately detect and rank them is an important problem for third-party online service providers and researchers. Conventional efforts, ranging from early graph-based algorithms to recent machine learning-based approaches, typically rely on an entire network's information or at least ego networks. However, for privacy-conscious users or newly-registered users, such information is not easily accessible. To address this issue, we present DeepPick, a novel framework that considers both the generalization and specialization in the detection task of OUs. For generalization, we introduce deep neural networks to capture nonlinear features. For specialization, we leverage the traditional well-defined metrics to preserve common features. Extensive experiments based on real-world datasets demonstrate that our approach achieves a high efficacy in terms of detection performance against the state-of-the-art. Wanda Li, Qingyuan Gong, Yang Chen 0001, Aaron Yi Ding, Xin Wang 0002, Pan Hui 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Trimming Mobile Applications for Bandwidth-Challenged Networks in Developing RegionsabstractDespite continuous efforts to build and update mobile network infrastructure, mobile devices in developing regions continue to be constrained by limited bandwidth. Unfortunately, this coincides with a period of unprecedented growth in the sizes of mobile applications. Thus it is becoming prohibitively expensive for users in developing regions to download and update mobile apps critical to their economic and educational development. Unchecked, these trends can further contribute to a large and growing global digital divide. Our goal is to better understand the source of this rapid growth in mobile app code size, whether it is reflective of new functionality, and identify steps that can be taken to make existing mobile apps more friendly to bandwidth constrained mobile networks. We hypothesize that much of this growth in mobile apps is due to poor resource/code management, and do not reflect proportional increases in functionality. Our hypothesis is partially validated by mini-programs, apps with extremely small footprints gaining popularity in Chinese mobile platforms. Here, we use functionally equivalent pairs of mini-programs and Android apps to identify potential sources of “bloat,” i.e., inefficient uses of code or resources that contribute to large package sizes. We analyze a large sample of popular Android apps and quantify instances of code and resource bloat. We develop techniques for automated code and resource trimming, and successfully validate them on a large set of Android apps. We hope our results will lead to continued efforts to streamline mobile apps, making them easier to access and maintain for users in developing regions. Qinge Xie, Qingyuan Gong, Xinlei He 0001, Yang Chen 0001, Xin Wang 0002, Haitao Zheng 0001, Ben Y. Zhao |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | Structural Hole Theory in Social Network Analysis: A ReviewabstractSocial networks now connect billions of people around the world, where individuals occupying different positions often represent different social roles and show different characteristics in their behaviors. The structural hole (SH) theory demonstrates that users occupying the bridging positions between different communities have advantages since they control the key information diffusion paths. Users of this type, known as SH spanners, are important when it comes to assimilating social network structures and user behaviors. In this article, we review the use of SHs theory in social network analysis, where SH spanners take advantage of both information and control benefits. We investigate the existing algorithms of SH spanner detection and classify them into information flow-based algorithms and network centrality-based algorithms. For practitioners, we further illustrate the applications of SH theory in various practical scenarios, including enterprise settings, information diffusion in social networks, software development, mobile applications, and machine learning (ML)-based social prediction. Our review provides a comprehensive discussion on the foundation, detection, and practical applications of SHs. The insights can facilitate researchers and service providers to better apply the theory and derive value-added tools with advanced ML techniques. To inspire follow-up research, we identify potential research trends in this area, especially on the dynamics of networks. Zihang Lin, Qingyuan Gong, Yang Chen 0001, Atte Oksanen, Aaron Yi Ding |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2021 | DatingSec: Detecting Malicious Accounts in Dating Apps Using a Content-Based Attention NetworkabstractDating apps have gained tremendous popularity during the past decade. Compared with traditional offline dating means, dating apps ease the process of partner finding significantly. While bringing convenience to hundreds of millions of users, dating apps are vulnerable to become targets of adversaries. In this article, we focus on malicious user detection in dating apps. Existing methods overlooked the signals hidden in the textual information of user interactions, particularly the interplay of temporal-spatial behaviors and textual information, leading to limited detection performance. To tackle this, we propose DatingSec, a novel malicious user detection system for dating apps. Concretely, DatingSec leverages long short-term memory neural networks (LSTM) and an attentive module to capture the interplay of users' temporal-spatial behaviors and user-generated textual content. We evaluate DatingSec on a real-world dataset collected from Momo, a widely used dating app with more than 180 million users. Experimental results show that DatingSec outperforms state-of-the-art methods and achieves an F1-score of 0.857 and AUC of 0.940. Xinlei He 0001, Qingyuan Gong, Yang Chen 0001, Yang Zhang 0016, Xin Wang 0002, Xiaoming Fu 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Cross-site Prediction on Social Influence for Cold-start Users in Online Social NetworksabstractOnline social networks (OSNs) have become a commodity in our daily life. As an important concept in sociology and viral marketing, the study of social influence has received a lot of attentions in academia. Most of the existing proposals work well on dominant OSNs, such as Twitter, since these sites are mature and many users have generated a large amount of data for the calculation of social influence. Unfortunately, cold-start users on emerging OSNs generate much less activity data, which makes it challenging to identify potential influential users among them. In this work, we propose a practical solution to predict whether a cold-start user will become an influential user on an emerging OSN, by opportunistically leveraging the user’s information on dominant OSNs. A supervised machine learning-based approach is adopted, transferring the knowledge of both the descriptive information and dynamic activities on dominant OSNs. Descriptive features are extracted from the public data on a user’s homepage. In particular, to extract useful information from the fine-grained dynamic activities that cannot be represented by the statistical indices, we use deep learning technologies to deal with the sequential activity data. Using the real data of millions of users collected from Twitter (a dominant OSN) and Medium (an emerging OSN), we evaluate the performance of our proposed framework to predict prospective influential users. Our system achieves a high prediction performance based on different social influence definitions. Qingyuan Gong, Yang Chen 0001, Xinlei He 0001, Yu Xiao 0001, Pan Hui 0001, Xin Wang 0002, Xiaoming Fu 0001 |
ACM Trans. Web | 1 |
| 2019 | Detecting Malicious Accounts in Online Developer Communities Using Deep LearningabstractOnline developer communities like GitHub provide services such as distributed version control and task management, which allow a massive number of developers to collaborate online. However, the openness of the communities makes themselves vulnerable to different types of malicious attacks, since the attackers can easily join and interact with legitimate users. In this work, we formulate the malicious account detection problem in online developer communities, and propose GitSec, a deep learning-based solution to detect malicious accounts. GitSec distinguishes malicious accounts from legitimate ones based on the account profiles as well as dynamic activity characteristics. On one hand, GitSec makes use of users' descriptive features from the profiles. On the other hand, GitSec processes users' dynamic behavioral data by constructing two user activity sequences and applying a parallel neural network design to deal with each of them, respectively. An attention mechanism is used to integrate the information generated by the parallel neural networks. The final judgement is made by a decision maker implemented by a supervised machine learning-based classifier. Based on the real-world data of GitHub users, our extensive evaluations show that GitSec is an accurate detection system, with an F1-score of 0.922 and an AUC value of 0.940. Qingyuan Gong, Jiayun Zhang, Yang Chen 0001, Qi Li 0002, Yu Xiao 0001, Xin Wang 0002, Pan Hui 0001 |
CIKM | 1 |
| 2019 | Exploring the power of social hub services
Qingyuan Gong, Yang Chen 0001, Zhichun Guo, Yu Xiao 0001, Fehmi Ben Abdesslem, Xin Wang 0002, Pan Hui 0001 |
World Wide Web | 1 |
| 2018 | A Multi-tab Website Fingerprinting AttackabstractIn a Website Fingerprinting (WF) attack, a local, passive eavesdropper utilizes network flow information to identify which web pages a user is browsing. Previous researchers have extensively demonstrated the feasibility and effectiveness of WF, but only under the strong Single Page Assumption: the network flow extracted by the adversary always belongs to a single page. In other words, the WF classifier will never be asked to classify a network flow corresponding to more than one page, or part of a page. The Single Page Assumption is unrealistic because people often browse with multiple tabs. When this happens, the network flow induced by multiple tabs will overlap, and current WF attacks fail to classify correctly. Yixiao Xu, Tao Wang 0012, Qi Li 0002, Qingyuan Gong, Yang Chen 0001, Yong Jiang 0001 |
ACSAC | 4 |
| 2018 | Deep Learning-Based Malicious Account Detection in the Momo Social NetworkabstractDue to the rapid development of mobile devices and location-based services, location-based social networks (LBSNs) have become very popular in our daily-life. Malicious account detection is very helpful for different kinds of practical applications. In this paper, we explore the malicious account detection problem by introducing a deep learning-based framework. By using the long short-term memory (LSTM) neural network, we are able to build a classifier to achieve the binary classification. By using the real data collected from Momo, a widely used LBSN which has more than 180 million users around the world, we evaluate our framework and the result shows great promise for malicious account detection tasks. Xinlei He 0001, Qingyuan Gong, Yang Chen 0001, Tianyi Wang 0001, Xin Wang 0002 |
ICCCN | 3 |
| 2018 | Understanding Cross-Site Linking in Online Social NetworksabstractAs a result of the blooming of online social networks (OSNs), a user often holds accounts on multiple sites. In this article, we study the emerging “cross-site linking” function available on mainstream OSN services including Foursquare, Quora, and Pinterest. We first conduct a data-driven analysis on crawled profiles and social connections of all 61.39 million Foursquare users to obtain a thorough understanding of this function. Our analysis has shown that the cross-site linking function is adopted by 57.10% of all Foursquare users, and the users who have enabled this function are more active than others. We also find that the enablement of cross-site linking might lead to privacy risks. Based on cross-site links between Foursquare and external OSN sites, we formulate cross-site information aggregation as a problem that uses cross-site links to stitch together site-local information fields for OSN users. Using large datasets collected from Foursquare, Facebook, and Twitter, we demonstrate the usefulness and the challenges of cross-site information aggregation. In addition to the measurements, we carry out a survey collecting detailed user feedback on cross-site linking. This survey studies why people choose to or not to enable cross-site linking, as well as the motivation and concerns of enabling this function. Qingyuan Gong, Yang Chen 0001, Jiyao Hu, Qiang Cao 0005, Pan Hui 0001, Xin Wang 0002 |
ACM Trans. Web | 1 |
| 2015 | Optimal Node Selection for Data Regeneration in Heterogeneous Distributed Storage SystemsabstractDistributed storage systems introduce redundancy to protect data from node failures. After a storage node fails, the lost data should be regenerated at a replacement storage node as soon as possible to maintain the same level of redundancy. Minimizing such a regeneration time is critical to the reliability of distributed storage systems. Existing work commits to reduce the regeneration time by either minimizing the regenerating traffic, or adjusting the regenerating traffic patterns, whereas nodes participating the regeneration are generally assumed to be given beforehand. However, real-world distributed storage systems usually exhibit heterogeneous link capacities, and the regeneration time is highly related to the selection of the participating nodes. In this paper, we consider the minimization of the regeneration time by selecting the participating nodes in heterogeneous networks. We propose optimal node selection algorithms respectively for two cases: 1) the newcomer is not given, 2) both the newcomer and the providers are not given. Analysis shows that the optimal regeneration time can be achieved in each case. We then consider the effect of flexible amount of data blocks from each provider on the regeneration time, and apply this observation to enhance our schemes. Experiment results show that our node selection schemes can significantly reduce the regeneration time, especially in practical networks with heterogeneous link capacities, compared with the scheme based on random node selection. Qingyuan Gong, Dongsheng Wei, Jin Wang 0009, Xin Wang 0002 |
ICPP | 1 |
| 2014 | Decluster: a complex network model-based data center network topology
Xu Zhang 0021, Hai Wang 0007, Qingyuan Gong, Xin Wang 0002 |
J. Supercomput. | 3 |