EDBT 2026 Demo / reviewers in the wild / expert
Shaofeng Hu
dblp:69/8730
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-3056-4440ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pre-trained Behavioral Model for Malicious User Prediction on Social PlatformabstractThe proliferation of malicious users on social platforms poses significant financial and psychological threats, with activities ranging from scams to the dissemination of illicit content. Existing malicious user prediction comprises supervised and self-supervised learning methods. However, the former relies on extensive labeled malicious users for training, while the latter typically focuses on one form of malicious activity and depends heavily on manually crafted rules and features during pre-training. Moreover, existing pre-training methods fail to effectively capture the crucial repetitive and sporadic behavior patterns of malicious users. To address these limitations, we propose a Malicious User Behavior Pre-training framework (MaP) to build pre-trained behavior models. MaP integrates malicious pattern recognition with behavior consistency augmentation and local disruption augmentation strategies for contrastive learning to capture repetitive and sporadic malicious patterns, respectively. We instantiate MaP on a billion-level behavior pre-training scenario within an industry context. Both online and offline evaluations validate the superior performance of MaP in malicious user detection and classification. Wenjie Wang 0007, Shaofeng Hu, Kaishen Ou, Zhenjing Zheng, Fuli Feng |
AAAI | 3 |
| 2025 | An LLM-based Behavior Modeling Framework for Malicious User DetectionabstractMalicious users pose significant threats to social platforms. Extensive efforts have leveraged user behavior sequences to model relationships between various actions and capture behavioral patterns for malicious user detection; however, they rely on behavior IDs, ignoring valuable behavior content such as self-introductions in friend requests, which offer crucial clues for detecting malicious user. We thus propose leveraging Large Language Models (LLMs) to jointly model IDs and content in user behavior sequences. The key to effective malicious user detection is to infer malicious user behavior patterns. However, inferring these patterns from labeled behavior sequences suffers from poor data efficiency and limited generalization, resulting in suboptimal malicious user detection performance. Wenjie Wang 0007, Chongming Gao, Shaofeng Hu, Kaishen Ou, Fuli Feng |
CIKM | 4 |
| 2024 | Crowdsourcing Fraud Detection Over Heterogeneous Temporal MMMA Graph
Zequan Xu, Shaofeng Hu, Jieming Shi 0001, Hui Li 0057 |
DASFAA (7) | 3 |
| 2023 | Self-supervised Graph Representation Learning for Black Market Account DetectionabstractNowadays, Multi-purpose Messaging Mobile App (MMMA) has become increasingly prevalent. MMMAs attract fraudsters and some cybercriminals provide support for frauds via black market accounts (BMAs). Compared to fraudsters, BMAs are not directly involved in frauds and are more difficult to detect. This paper illustrates our BMA detection system SGRL (Self-supervised Graph Representation Learning) used in WeChat, a representative MMMA with over a billion users. We tailor Graph Neural Network and Graph Self-supervised Learning in SGRL for BMA detection. The workflow of SGRL contains a pretraining phase that utilizes structural information, node attribute information and available human knowledge, and a lightweight detection phase. In offline experiments, SGRL outperforms state-of-the-art methods by 16.06%-58.17% on offline evaluation measures. We deploy SGRL in the online environment to detect BMAs on the billion-scale WeChat graph, and it exceeds the alternative by 7.27% on the online evaluation measure. In conclusion, SGRL can alleviate label reliance, generalize well to unseen data, and effectively detect BMAs in WeChat. Zequan Xu, Lianyun Li, Hui Li 0057, Shaofeng Hu, Rongrong Ji |
WSDM | 5 |
| 2022 | Efficiently Answering k-hop Reachability Queries in Large Dynamic Graphs for Fraud Feature ExtractionabstractInstant messaging client (IMC) is now an essential tool for mobile users. In the representative IMC We Chat, cybercriminals deceive frauds, causing financial loss to normal users. Through statistical analysis, we find that certain fraud interactions commonly occur among WeChat users who are not k-hop neighbors. Therefore, efficiently answering whether the distance between two vertices is not longer than k at a certain time point (i.e., k-hop reachability queries) over the dynamic social graph of WeChat becomes a crucial task for fraud feature extraction in the detection system: it can help human experts quickly identify suspicious user interactions and the query results can be further used as the input feature to the downstream machine learning based detection methods. In this paper, we illustrate Bidirectional k-hop Reachability Query Processing over a Dynamic Graph (BREAD) that is used in WeChat for extracting the k-hop reachability feature for fraud detection. BREAD adopts the idea of estimating Personalized PageRank value. It first conducts the backward search from the destination vertex to construct an intermediate vertex set. Then, it performs a certain amount of random walks from the start vertex to see whether they can hit the intermediate vertex set, and the results are returned to answer k-hop reachability queries. We further propose$\text{BREAD}++$that leverages the massive parallel processing power of GPU to achieve a considerable performance gain. Experiments on several large-scale dynamic graph benchmarks and the social graph of WeChat have demonstrated that$\text{BREAD}/\text{BREAD}++$is superior than existing index-free competitors: our methods provide not only fast but also accurate responses and they are of practical value to k-hop reachability feature extraction in the fraud detection system of WeChat. Our implementation is available at https://github.com/XMUDM/BREAD. Zequan Xu, Siqiang Luo, Jieming Shi 0001, Hui Li 0057, Chen Lin 0001, Shaofeng Hu |
MDM | 7 |
| 2022 | Multi-view Heterogeneous Temporal Graph Neural Network for "Click Farming" Detection
Zequan Xu, Shaofeng Hu, Jiguang Qiu, Chen Lin 0001, Hui Li 0057 |
PRICAI (1) | 3 |
| 2021 | On Detecting Growing-Up Behaviors of Malicious Accounts in Privacy-Centric Mobile Social NetworksabstractPrivacy-centric mobile social network (PC-MSN), which allows users to build intimate and private social circles, is an increasingly popular type of online social networks (OSNs). Because of strict usage policy enforced by PC-MSNs (such as restricted account and content access), malicious accounts (or users) have to act like normal accounts to accumulate credentials before committing malicious activities. Therefore, analysis merely relying on static account profile information or social graphs is ineffective to detect such growing-up accounts. Besides, existing behavior-based malicious account detection methods fail to effectively detect growing-up accounts who pretend to be benign and have similar behaviors to benign users during the growing-up stage. Zijie Yang, Binghui Wang, Dong Yuan 0006, Zhuotao Liu, Neil Zhenqiang Gong, Chang Liu 0021, Qi Li 0002, Shaofeng Hu |
ACSAC | 10 |
| 2021 | Unveiling Fake Accounts at the Time of Registration: An Unsupervised ApproachabstractOnline social networks (OSNs) are plagued by fake accounts. Existing fake account detection methods either require a manually labeled training set, which is time-consuming and costly, or rely on rich information of OSN accounts, e.g., content and behaviors, which incurs significant delay in detecting fake accounts. In this work, we propose UFA (Unveiling Fake Accounts) to detect fake accounts immediately after they are registered in an unsupervised fashion. First, through a measurement study on the registration patterns on a real-world registration dataset, we observe that fake accounts tend to cluster on outlier registration patterns, e.g., IP and phone numbers. Then, we design an unsupervised learning algorithm to learn weights for all registration accounts and their features that reveal outlier registration patterns. Next, we construct a registration graph to capture the correlation between registration accounts, and utilize a community detection method to detect fake accounts via analyzing the registration graph structure. We evaluate UFA using real-world WeChat datasets. Our results demonstrate that UFA achieves a precision 94% with a recall ~80%, while a supervised variant requires 600K manual labels to obtain the comparable performance. Moreover, UFA has been deployed by WeChat to detect fake accounts for more than one year. UFA detects 500K fake accounts per day with a precision ~93% on average, via manual verification by the WeChat security team. Binghui Wang, Shaofeng Hu, Zijie Yang, Dong Yuan 0006, Neil Zhenqiang Gong, Qi Li 0002 |
KDD | 4 |
| 2014 | Domain Transfer via Multiple Sources Regularization
Shaofeng Hu, Jiangtao Ren, Changshui Zhang, Chaogui Zhang |
PAKDD (2) | 1 |
| 2010 | Multiple Kernel Learning Improved by MMD
Jiangtao Ren, Zhou Liang, Shaofeng Hu |
ADMA (2) | 3 |