Minnan Luo

dblp:99/10051 · DBLP profile ↗
← Back
36ranked-venue papers in the field
1as first author
27since 2021 · last 2026
0000-0002-0140-7860ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 13Information Retrieval & Web Search · 10Data Mining & Knowledge Discovery · 8Knowledge Engineering, Semantic Web & Information Systems · 4 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the Wild
abstract
The rise of micro-videos has reshaped how misinformation spreads, amplifying its speed, reach, and impact on public trust. Existing benchmarks typically focus on a single deception type, overlooking the diversity of real-world cases that involve multimodal manipulation, AI-generated content, cognitive bias, and out-of-context reuse. Meanwhile, most detection models lack fine-grained attribution, limiting interpretability and practical utility. To address these gaps, we introduce WildFakeBench, a large-scale benchmark of over 10,000 real-world micro-videos covering diverse misinformation types and sources, each annotated with expert-defined attribution labels. Building on this foundation, we develop FakeAgent, a Delphi-inspired multi-agent reasoning framework that integrates multimodal understanding with external evidence for attribution-grounded analysis. FakeAgent jointly analyzes content and retrieved evidence to identify manipulation, recognize cognitive and AI-generated patterns, and detect out-of-context misinformation. Extensive experiments show that FakeAgent consistently outperforms existing MLLMs across all misinformation types, while WildFakeBench provides a realistic and challenging testbed for advancing explainable micro-video misinformation detection. Data and code are available at: https://github.com/Aiyistan/FakeAgent.
Zhi Zeng 0001, Xulang Zhang, Xiangzheng Kong, Herun Wan, Zihan Ma 0001, Minnan Luo
WWW8
2026 Knowledge Graph-Based Debiasing for Trustworthy Recommendation Systems
abstract
These years have witnessed remarkable progress in modeling user behaviour from personalized online services, especially knowledge graph-based recommendation systems. Meanwhile, more studies are focusing on aspects beyond recommendation performance, since such an observational data-driven paradigm is posing threats to both users and society in terms of trustworthiness. In fact, existing problem-oriented solutions still face significant challenges, as almost all of them suffer from the generality limitations to improve their trustworthiness in a uniform fashion. To address these issues, we propose a plug-and-playDebiasing framework forKnowledgeGraph-basedRecommendationSystems, also known as DiKGRS. Specifically, the Knowledge-augmented Pseudo-Samples Generation (KPSG) method, a novel data augmentation perspective, is proposed to explore more auxiliary information beyond observational user behaviors. Furthermore, the Debiasing Value Networks (DVN), is also developed to evaluate the reliability of generated pseudo-samples by modeling both the item popularity and user demographic bias in the platform. Moreover, an adaptive weighting coordination module is performed to coordinate the proposed DiKGRS framework and its backbones. Experimental results on four real-world datasets from different online service personalization scenarios have illustrated that the proposed framework can significantly improve the trustworthiness of existing knowledge graph-based recommendation systems. The code has been released public available at:https://github.com/alipay/A-Knowledge-augmented-Method-DiKGRS.
Youru Li, Xuying Ning, Zhenfeng Zhu, Hanqiu Wang, Zhi Cai, Minnan Luo, Yao Zhao 0001
IEEE Trans. Knowl. Data Eng.6
2026 HCGBot: Learning Homophilous Context Graphs for Twitter Bot Detection
Herun Wan, Minnan Luo, Jihong Wang 0003, Xiaojun Chang
IEEE Trans. Knowl. Data Eng.2
2025 Unveiling the Hidden: Movie Genre and User Bias in Spoiler Detection
Haokai Zhang, Shengtao Zhang, Zijian Cai, Heng Wang 0008, Ruixuan Zhu, Zinan Zeng 0001, Minnan Luo
ECML/PKDD (3)7
2025 Bridging Interests and Truth: Towards Mitigating Fake News with Personalized and Truthful Recommendations
abstract
While the proliferation of fake news poses a significant threat to information integrity, existing efforts to counter it, especially within personalized news recommendation systems, have proven inadequate.Traditional methods, which often rely on classifiers to filter out fake content, are limited by their accuracy and their inability to fully capture the diverse interests of users.To address these challenges, we proposed PRISM-Protection-enhanced Recommendation with Interest-aware Sequential Modeling-a novel framework based on diffusion models.PRISM harnesses the generative and control capabilities of diffusion models to progressively learn the implicit distribution of user interests from their reading history, thereby generating personalized recommendations that align with both their linguistic preferences and interest domains.Furthermore, PRISM incorporates pre-trained authenticity representations as constraints during content generation, ensuring the credibility of the recommended news and effectively curbing the spread of fake news.Comprehensive evaluations from multiple dimensions demonstrate the superiority of our model.
Zihan Ma 0001, Minnan Luo, Yiran Hao, Zhi Zeng 0001, Xiangzheng Kong, Jiahao Wang 0004
SIGIR2
2025 From Predictions to Analyses: Rationale-Augmented Fake News Detection with Large Vision-Language Models
abstract
The rapid development of social media has led to a surge of eye-catching fake news on the Internet, with multimodal news comprising both images and text being particularly prevalent. To address the challenges of Multimodal Fake News Detection (MFND), numerous supervised task-specific Multimodal Small Language Models (MSLMs) have been developed. However, these models lack the breadth of knowledge and the depth of language understanding, which results in unsatisfactory adaptability, generalization, and explainability performance. To address these issues, we attempt to introduce Large Vision-Language Models (LVLMs), aiming to leverage the common sense understanding and logical reasoning abilities of LVLMs for the MFND task. We observed that LVLMs can generate reasonable analyses of news content from specific angles. However, when it comes to synthesizing these analyses for final judgment, their performance declines significantly, failing to meet the accuracy benchmarks set by existing MSLMs detection models. This reflects the need for a more effective way for LVLMs, which have not undergone task-specific training, to utilize their knowledge and capabilities. Based on these findings, we propose the Explainable Adaptive Rationale-Augmented Multimodal (EARAM) framework, which adaptively uses MSLMs to extract useful rationales from the multi-perspective analyses of LVLMs. After making judgments based on these rationales, EARAM then assists LVLMs in generating more reliable explanations. Extensive experiments demonstrate that our model not only achieves state-of-the-art results on widely used datasets but also significantly outperforms other models in terms of generalization and explainability.
Xiaofan Zheng, Zinan Zeng 0001, Heng Wang 0008, Yuyang Bai, Yuhan Liu 0028, Minnan Luo
WWW6
2024 FNDPro: Evaluating the Importance of Propagations during Fake News Spread
Herun Wan, Ningnan Wang, Xiang Zhao 0002, Minnan Luo
DASFAA (6)6
2024 Towards Open-World Cross-Domain Sequential Recommendation: A Model-Agnostic Contrastive Denoising Approach
Wujiang Xu, Xuying Ning, Wenfang Lin, Mingming Ha, Qiongxu Ma, Qianqiao Liang, Xuewen Tao, Linxun Chen, Minnan Luo
ECML/PKDD (1)10
2024 Disentangled Counterfactual Graph Augmentation Framework for Fair Graph Learning with Information Bottleneck
Lijing Zheng, Jihong Wang 0003, Minnan Luo
ECML/PKDD (1)4
2024 LMBot: Distilling Graph Knowledge into Language Model for Graph-less Deployment in Twitter Bot Detection
abstract
As malicious actors employ increasingly advanced and widespread bots to disseminate misinformation and manipulate public opinion, the detection of Twitter bots has become a crucial task. Though graph-based Twitter bot detection methods achieve state-of-the-art performance, we find that their inference depends on the neighbor users multi-hop away from the targets, and fetching neighbors is time-consuming and may introduce sampling bias. At the same time, our experiments reveal that after finetuning on Twitter bot detection task, pretrained language models achieve competitive performance while do not require a graph structure during deployment. Inspired by this finding, we propose a novel bot detection framework LMBot that distills the graph knowledge into language models (LMs) for graph-less deployment in Twitter bot detection to combat data dependency challenge. Moreover, LMBot is compatible with graph-based and graph-less datasets. Specifically, we first represent each user as a textual sequence and feed them into the LM for domain adaptation. For graph-based datasets, the output of LM serves as input features for the GNN, enabling LMBot to optimize for bot detection and distill knowledge back to the LM in an iterative, mutually enhancing process. Armed with the LM, we can perform graph-less inference with graph knowledge, which resolves the graph data dependency and sampling bias issues. For datasets without graph structure, we simply replace the GNN with an MLP, which also shows strong performance. Our experiments demonstrate that LMBot achieves state-of-the-art performance on four Twitter bot detection benchmarks. Extensive studies also show that LMBot is more robust, versatile, and efficient compared to existing graph-based Twitter bot detection methods.
Zijian Cai, Zhaoxuan Tan, Zhenyu Lei 0004, Zifeng Zhu, Hongrui Wang 0004, Minnan Luo
WSDM7
2024 Semi-Supervised Graph Contrastive Learning With Virtual Adversarial Augmentation
abstract
Semi-supervised graph learning aims to improve learning performance by leveraging unlabeled nodes. Typically, it can be approached in two different ways, includingpredictive representation learning(PRL) where unlabeled data provide clues on input distribution andlabel-dependent regularization(LDR) which smooths the output distribution with unlabeled nodes to improve generalization. However, most existing PRL approaches suffer from overfitting in an end-to-end setting or cannot encode task-specific information when used as unsupervised pre-training (i.e., two-stage learning). Meanwhile, LDR strategies often introduce redundant and invalid data perturbations that can slow down and mislead the training. To address all these issues, we propose a general framework SemiGraL for semi-supervised learning on graphs, which bridges and facilitates both PRL and LDR in a single shot. By extending a contrastive learning architecture to the semi-supervised setting, we first develop asemi-supervised contrastive representation learningprocess with virtual adversarial augmentation to map input nodes into a label-preserving representation space while avoiding overfitting. We then introduce amultiview consistency classificationprocess with well-constrained perturbations to achieve adversarially robust classification. Extensive experiments on seven semi-supervised node classification benchmark datasets show that SemiGraL outperforms various baselines while enjoying strong generalization and robustness performance.
Yixiang Dong, Minnan Luo, Jundong Li
IEEE Trans. Knowl. Data Eng.2
2024 Toward Enhanced Robustness in Unsupervised Graph Representation Learning: A Graph Information Bottleneck Perspective
abstract
Recent studies have revealed that GNNs are vulnerable to adversarial attacks. Most existing robust graph learning methods measure model robustness based on label information, rendering them infeasible when label information is not available. A straightforward direction is to employ the widely used Infomax technique from typical Unsupervised Graph Representation Learning (UGRL) to learn robust unsupervised representations. Nonetheless, directly transplanting the Infomax technique from typical UGRL to robust UGRL may involve a biased assumption. In light of the limitation of Infomax, we propose a novel unbiased robust UGRL method calledRobust Graph Information Bottleneck(RGIB), which is grounded in the Information Bottleneck (IB) principle. Our RGIB attempts to learn robust node representations against adversarial perturbations by preserving the original information in the benign graph while eliminating the adversarial information in the adversarial graph. There are mainly two challenges to optimizing RGIB: 1) high complexity of adversarial attack to perturb node features and graph structure jointly in the training procedure; 2) mutual information estimation upon adversarially attacked graphs. To tackle these problems, we further propose an efficient adversarial training strategy with only feature perturbations and an effective mutual information estimator with the subgraph-level summary. Moreover, we theoretically establish a connection between our proposed RGIB and the robustness of downstream classifiers, revealing that RGIB can provide a lower bound on the adversarial risk of downstream classifiers. Extensive experiments over several benchmarks and downstream tasks demonstrate the effectiveness and superiority of our proposed method.
Jihong Wang 0003, Minnan Luo, Jundong Li, Jun Zhou 0011
IEEE Trans. Knowl. Data Eng.2
2023 A Deep Multi-View Framework for Anomaly Detection on Attributed Networks (Extended Abstract)
abstract
Many existing anomaly detection methods on attributed networks do not seriously tackle the inherent multi-view property in attribute space but concatenate multiple views into a single feature vector, which inevitably ignores the incompatibility between heterogeneous views caused by their own statistical properties. In practice, the distinct but complementary information brought by multi-view data promises the potential for more effective anomaly detection than the efforts only based on single-view data. Furthermore, abnormal patterns naturally behave diversely in different views, which coincides with people’s desire to discover specific abnormalities according to their preferences for views (attributes). Most existing methods cannot adapt to people’s requirements as they fail to consider the idiosyncrasy of user preferences. Thus, in this paper, we propose a multi-view framework ALARM to incorporate user preferences into anomaly detection and simultaneously tackle heterogeneous attribute characteristics through multiple graph encoders and a well-designed aggregator that supports self-learning and user-guided learning. Experiments on synthetic and real-world datasets corroborate the desirable performance of ALARM and its effectiveness in supporting user-oriented anomaly detection.
Zhen Peng 0005, Minnan Luo, Jundong Li, Luguo Xue
ICDE2
2023 Empower Post-hoc Graph Explanations with Information Bottleneck: A Pre-training and Fine-tuning Perspective
abstract
Researchers recently investigated to explain Graph Neural Networks (GNNs) on the access to a task-specific GNN, which may hinder their wide applications in practice. Specifically, task-specific explanation methods are incapable of explaining pretrained GNNs whose downstream tasks are usually inaccessible, not to mention giving explanations for the transferable knowledge in pretrained GNNs. Additionally, task-specific methods only consider target models' output in the label space, which are coarse-grained and insufficient to reflect the model's internal logic. To address these limitations, we consider a two-stage explanation strategy, i.e., explainers are first pretrained in a task-agnostic fashion in the representation space and then further fine-tuned in the task-specific label space and representation space jointly if downstream tasks are accessible. The two-stage explanation strategy endows post-hoc graph explanations with the applicability to pretrained GNNs where downstream tasks are inaccessible and the capacity to explain the transferable knowledge in the pretrained GNNs. Moreover, as the two-stage explanation strategy explains the GNNs in the representation space, the fine-grained information in the representation space also empowers the explanations. Furthermore, to achieve a trade-off between the fidelity and intelligibility of explanations, we propose an explanation framework based on the Information Bottleneck principle, named Explainable Graph Information Bottleneck (EGIB). EGIB subsumes the task-specific explanation and task-agnostic explanation into a unified framework. To optimize EGIB objective, we derive a tractable bound and adopt a simple yet effective explanation generation architecture. Based on the unified framework, we further theoretically prove that task-agnostic explanation is a relaxed sufficient condition of task-specific explanation, which indicates the transferability of task-agnostic explanations. Extensive experimental results demonstrate the effectiveness of our proposed explanation method.
Jihong Wang 0003, Minnan Luo, Jundong Li, Yun Lin 0001, Yushun Dong, Jin Song Dong 0001
KDD2
2023 BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific Experts
abstract
Twitter bot detection has become a crucial task in efforts to combat online misinformation, mitigate election interference, and curb malicious propaganda. However, advanced Twitter bots often attempt to mimic the characteristics of genuine users through feature manipulation and disguise themselves to fit in diverse user communities, posing challenges for existing Twitter bot detection models. To this end, we propose BotMoE, a Twitter bot detection framework that jointly utilizes multiple user information modalities (metadata, textual content, network structure) to improve the detection of deceptive bots. Furthermore, BotMoE incorporates a community-aware Mixture-of-Experts (MoE) layer to improve domain generalization and adapt to different Twitter communities. Specifically, BotMoE constructs modal-specific encoders for metadata features, textual content, and graph structure, which jointly model Twitter users from three modal-specific perspectives. We then employ a community-aware MoE layer to automatically assign users to different communities and leverage the corresponding expert networks. Finally, user representations from metadata, text, and graph perspectives are fused with an expert fusion layer, combining all three modalities while measuring the consistency of user information. Extensive experiments demonstrate that BotMoE significantly advances the state-of-the-art on three Twitter bot detection benchmarks. Studies also confirm that BotMoE captures advanced and evasive bots, alleviates the reliance on training data, and better generalizes to new and previously unseen user communities.
Yuhan Liu 0028, Zhaoxuan Tan, Heng Wang 0008, Shangbin Feng, Minnan Luo
SIGIR6
2023 KRACL: Contrastive Learning with Graph Context Modeling for Sparse Knowledge Graph Completion
abstract
Knowledge Graph Embeddings (KGE) aim to map entities and relations to low dimensional spaces and have become the de-facto standard for knowledge graph completion. Most existing KGE methods suffer from the sparsity challenge, where it is harder to predict entities that appear less frequently in knowledge graphs. In this work, we propose a novel framework KRACL1 to alleviate the widespread sparsity in KGs with graph context and contrastive learning. Firstly, we propose the Knowledge Relational Attention Network (KRAT) to leverage the graph context by simultaneously projecting neighboring triples to different latent spaces and jointly aggregating messages with the attention mechanism. KRAT is capable of capturing the subtle semantic information and importance of different context triples as well as leveraging multi-hop information in knowledge graphs. Secondly, we propose the knowledge contrastive loss by combining the contrastive loss with cross entropy loss, which introduces more negative samples and thus enriches the feedback to sparse entities. Our experiments demonstrate that KRACL achieves superior results across various standard knowledge graph benchmarks, especially on WN18RR and NELL-995 which have large numbers of low in-degree entities. Extensive experiments also bear out KRACL’s effectiveness in handling sparse knowledge graphs and robustness against noisy triples.
Zhaoxuan Tan, Zilong Chen, Shangbin Feng, Qingyue Zhang 0003, Jundong Li, Minnan Luo
WWW7
2022 GraTO: Graph Neural Network Framework Tackling Over-smoothing with Neural Architecture Search
abstract
Current Graph Neural Networks (GNNs) suffer from the over-smoothing problem, which results in indistinguishable node representations and low model performance with more GNN layers. Many methods have been put forward to tackle this problem in recent years. However, existing tackling over-smoothing methods emphasize model performance and neglect the over-smoothness of node representations. Additional, different approaches are applied one at a time, while there lacks an overall framework to jointly leverage multiple solutions to the over-smoothing challenge. To solve these problems, we propose GraTO, a framework based on neural architecture search to automatically search for GNNs architecture. GraTO adopts a novel loss function to facilitate striking a balance between model performance and representation smoothness. In addition to existing methods, our search space also includes DropAttribute, a novel scheme for alleviating the over-smoothing challenge, to fully leverage diverse solutions. We conduct extensive experiments on six real-world datasets to evaluate GraTo, which demonstrates that GraTo outperforms baselines in the over-smoothing metrics and achieves competitive performance in accuracy. GraTO is especially effective and robust with increasing numbers of GNN layers. Further experiments bear out the quality of node representations learned with GraTO and the effectiveness of model architecture. We make the code of GraTo available at Github (https://github.com/fxsxjtu/GraTO).
Xinshun Feng, Herun Wan, Shangbin Feng, Hongrui Wang 0004, Jun Zhou 0011, Minnan Luo
CIKM7
2022 Intent Mining: A Social and Semantic Enhanced Topic Model for Operation-Friendly Digital Marketing
abstract
In this paper, we study the digital marketing where marketing officers (MOs) have to commit to creating brand new promotion ads/contents based on understandings of users' needs or preferences. Users' behaviors are typically high dimensional and hard to understand. Therefore, dimension reduction of users' behaviors from high dimensions and explainability are important to help MOs launch operation-friendly marketings. As such, it is natural to exploit topic models to help MOs understand users' intents from users' behaviors (e.g., user-item visits) in case we treat each user as a document and users' behaviors of visiting an item as a word. However, users of low activities and items followed by power law distributions are common in user-item visit data, which pose significant challenges to traditional topic models. We present a social and semantic enhanced topic model (S2TM) for users' intent mining. We optimize the user-intent estimates based on a graph neural network atop of a social network, and optimize the intent-item estimates based on a skip-gram word embedding approach by linking the semantics of items to pre-trained word embeddings. We propose an efficient stochastic vari-ational inference algorithm for the inference of latent variables and learning of parameters. Extensive experiments on real-world data show the effectivenesses of S2TM in terms of perplexities, topic coherence and semantic coherence compared with state-of-the-art topic models. We further show how MOs interact with our operation-friendly intent mining system, and results on real-world marketing campaigns in terms of click-through rate at Alipay.
Weifan Wang 0005, Xiaocheng Cheng, Binbin Hu, Zhiqiang Zhang 0012, Xiaodong Zeng, Jun Zhou 0011, Jinjie Gu, Minnan Luo
ICDE11
2022 Toward Entity Alignment in the Open World: An Unsupervised Approach with Confidence Modeling
abstract
Abstract Entity alignment (EA) aims to discover the equivalent entities in different knowledge graphs (KGs). It is a pivotal step for integrating KGs to increase knowledge coverage and quality. Recent years have witnessed a rapid increase of EA frameworks. However, state-of-the-art solutions tend to rely on labeled data for model training. Additionally, they work under the closed-domain setting and cannot deal with entities that are unmatchable. To address these deficiencies, we offer an unsupervised framework that performs entity alignment in the open world. Specifically, we first mine useful features from the side information of KGs. Then, we devise an unmatchable entity prediction module to filter out unmatchable entities and produce preliminary alignment results. These preliminary results are regarded as the pseudo-labeled data and forwarded to the progressive learning framework to generate structural representations, which are integrated with the side information to provide a more comprehensive view for alignment. Finally, the progressive learning framework gradually improves the quality of structural embeddings and enhances the alignment performance. Furthermore, noticing that the pseudo-labeled data are of various qualities, we introduce the concept of confidence to measure the probability of an entity pair of being true and develop a confidence-based unsupervised EA framework . Our solutions do not require labeled data and can effectively filter out unmatchable entities. Comprehensive experimental evaluations validate the superiority of our proposals .
Xiang Zhao 0002, Weixin Zeng, Jiuyang Tang, Xinyi Li 0001, Minnan Luo
Data Sci. Eng.5
2022 A new self-supervised task on graphs: Geodesic distance prediction
Zhen Peng 0005, Yixiang Dong, Minnan Luo, Xiao-Ming Wu 0003
Inf. Sci.3
2022 LookCom: Learning Optimal Network for Community Detection
abstract
Community detection is one of the fundamental tasks in graph mining, which aims to identify group assignment of nodes in a complex network. Recently, network embedding techniques have demonstrated their strong power in advancing the community detection task and achieve better performance than various traditional methods. Despite their empirical success, most of the existing algorithms directly leverage the observed coarse network structure for community detection. Therefore, they often lead to suboptimal performance as the observed connections fail to capture the essential tie strength information among nodes precisely and account for the impact of noisy links. In this paper, an optimal network structure for community detection is introduced to characterize the fine-grained tie strength information between connected nodes and alleviate the adverse effects of noisy links. To obtain an expressive node representation for community detection, we learn the optimal network structure and network embeddings in a joint framework, instead of using a two-stage approach to derive the node embeddings from the coarse network topology. In particular, we formulate the joint framework as an optimization problem and an alternating optimization algorithm is exploited to solve the proposed optimization problem. Additionally, theoretical analyses regarding the computational complexity and the convergence of the optimization algorithm are also provided. Extensive experiments on both synthetic and real-world networks demonstrate the effectiveness and superiority of the proposed framework.
Yixiang Dong, Minnan Luo, Jundong Li, Deng Cai 0001
IEEE Trans. Knowl. Data Eng.2
2022 A Deep Multi-View Framework for Anomaly Detection on Attributed Networks
abstract
The explosion of modeling complex systems using attributed networks boosts the research on anomaly detection in such networks, which can be applied in various high-impact domains. Many existing attempts, however, do not seriously tackle the inherent multi-view property in attribute space but concatenate multiple views into a single feature vector, which inevitably ignores the incompatibility between heterogeneous views caused by their own statistical properties. Actually, the distinct but complementary information brought by multi-view data promises the potential for more effective anomaly detection than the efforts only based on single-view data. Furthermore, the abnormal patterns naturally behave diversely in different views, which coincides with people’s desire to discover specific abnormality according to their preferences for views (attributes). Most existing methods cannot adapt to people’s requirements as they fail to consider the idiosyncrasy of user preferences. Therefore, we propose a multi-view frameworkAlarmto incorporate user preferences into anomaly detection and simultaneously tackle heterogeneous attribute characteristics through multiple graph encoders and a well-designed aggregator that supports self-learning and user-guided learning. Experiments on synthetic and real-world datasets, e.g., Disney, Books, and Enron, corroborate the improvement ofAlarmin detection accuracy evaluated by the AUC metric and its effectiveness in supporting user-oriented anomaly detection.
Zhen Peng 0005, Minnan Luo, Jundong Li, Luguo Xue
IEEE Trans. Knowl. Data Eng.2
2021 BotRGCN: Twitter bot detection with relational graph convolutional networks
abstract
Twitter bot detection is an important and challenging task. Existing bot detection measures fail to address the challenge of community and disguise, falling short of detecting bots that disguise as genuine users and attack collectively. To address these two challenges of Twitter bot detection, we propose BotRGCN, which is short for Bot detection with Relational Graph Convolutional Networks. BotRGCN addresses the challenge of community by constructing a heterogeneous graph from follow relationships and applies relational graph convolutional networks. Apart from that, BotRGCN makes use of multi-modal user semantic and property information to avoid feature engineering and augment its ability to capture bots with diversified disguise. Extensive experiments demonstrate that BotRGCN outperforms competitive baselines on a comprehensive benchmark TwiBot-20 which provides follow relationships.
Shangbin Feng, Herun Wan, Ningnan Wang, Minnan Luo
ASONAM4
2021 SATAR: A Self-supervised Approach to Twitter Account Representation Learning and its Application in Bot Detection
abstract
Twitter has become a major social media platform since its launching in 2006, while complaints about bot accounts have increased recently. Although extensive research efforts have been made, the state-of-the-art bot detection methods fall short of generalizability and adaptability. Specifically, previous bot detectors leverage only a small fraction of user information and are often trained on datasets that only cover few types of bots. As a result, they fail to generalize to real-world scenarios on the Twittersphere where different types of bots co-exist. Additionally, bots in Twitter are constantly evolving to evade detection. Previous efforts, although effective once in their context, fail to adapt to new generations of Twitter bots. To address the two challenges of Twitter bot detection, we propose SATAR, a self-supervised representation learning framework of Twitter users, and apply it to the task of bot detection. In particular, SATAR generalizes by jointly leveraging the semantics, property and neighborhood information of a specific user. Meanwhile, SATAR adapts by pre-training on a massive number of self-supervised users and fine-tuning on detailed bot detection scenarios. Extensive experiments demonstrate that SATAR outperforms competitive baselines on different bot detection datasets of varying information completeness and collection time. SATAR is also proved to generalize in real-world scenarios and adapt to evolving generations of social media bots.
Shangbin Feng, Herun Wan, Ningnan Wang, Jundong Li, Minnan Luo
CIKM5
2021 TwiBot-20: A Comprehensive Twitter Bot Detection Benchmark
abstract
Twitter has become a vital social media platform while an ample amount of malicious Twitter bots exist and induce undesirable social effects. Successful Twitter bot detection proposals are generally supervised, which rely heavily on large-scale datasets. However, existing benchmarks generally suffer from low levels of user diversity, limited user information and data scarcity. Therefore, these datasets are not sufficient to train and stably benchmark bot detection measures. To alleviate these problems, we present TwiBot-20, a massive Twitter bot detection benchmark, which contains 229,573 users, 33,488,192 tweets, 8,723,736 user property items and 455,958 follow relationships. TwiBot-20 covers diversified bots and genuine users to better represent the real-world Twittersphere. TwiBot-20 also includes three modals of user information to support both binary classification of single users and community-aware approaches. To the best of our knowledge, TwiBot-20 is the largest Twitter bot detection benchmark to date. We reproduce competitive bot detection methods and conduct a thorough evaluation on TwiBot-20 and two other public datasets. Experiment results demonstrate that existing bot detection measures fail to match their previously claimed performance on TwiBot-20, which suggests that Twitter bot detection remains a challenging task and requires further research efforts.
Shangbin Feng, Herun Wan, Ningnan Wang, Jundong Li, Minnan Luo
CIKM5
2021 Towards Entity Alignment in the Open World: An Unsupervised Approach
Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xinyi Li 0001, Minnan Luo
DASFAA (1)5
2021 Self-weighted Robust LDA for Multiclass Classification with Edge Classes
abstract
Linear discriminant analysis (LDA) is a popular technique to learn the most discriminative features for multi-class classification. A vast majority of existing LDA algorithms are prone to be dominated by the class with very large deviation from the others, i.e., edge class, which occurs frequently in multi-class classification. First, the existence of edge classes often makes the total mean biased in the calculation of between-class scatter matrix. Second, the exploitation of ℓ2-norm based between-class distance criterion magnifies the extremely large distance corresponding to edge class. In this regard, a novel self-weighted robust LDA with ℓ2,1-norm based pairwise between-class distance criterion, called SWRLDA, is proposed for multi-class classification especially with edge classes. SWRLDA can automatically avoid the optimal mean calculation and simultaneously learn adaptive weights for each class pair without setting any additional parameter. An efficient re-weighted algorithm is exploited to derive the global optimum of the challenging ℓ2,1-norm maximization problem. The proposed SWRLDA is easy to implement and converges fast in practice. Extensive experiments demonstrate that SWRLDA performs favorably against other compared methods on both synthetic and real-world datasets while presenting superior computational efficiency in comparison with other techniques.
Caixia Yan, Xiaojun Chang, Minnan Luo, Xiaoqin Zhang 0002, Zhihui Li 0001, Feiping Nie 0001
ACM Trans. Intell. Syst. Technol.3
2020 Graph Few-shot Learning with Attribute Matching
abstract
Due to the expensive cost of data annotation, few-shot learning has attracted increasing research interests in recent years. Various meta-learning approaches have been proposed to tackle this problem and have become the de facto practice. However, most of the existing approaches along this line mainly focus on image and text data in the Euclidean domain. However, in many real-world scenarios, a vast amount of data can be represented as attributed networks defined in the non-Euclidean domain, and the few-shot learning studies in such structured data have largely remained nascent. Although some recent studies have tried to combine meta-learning with graph neural networks to enable few-shot learning on attributed networks, they fail to account for the unique properties of attributed networks when creating diverse tasks in the meta-training phase---the feature distributions of different tasks could be quite different as instances (i.e., nodes) do not follow the data i.i.d. assumption on attributed networks. Hence, it may inevitably result in suboptimal performance in the meta-testing phase. To tackle the aforementioned problem, we propose a novel graph meta-learning framework--Attribute Matching Meta-learning Graph Neural Networks (AMM-GNN). Specifically, the proposed AMM-GNN leverages an attribute-level attention mechanism to capture the distinct information of each task and thus learns more effective transferable knowledge for meta-learning. We conduct extensive experiments on real-world datasets under a wide range of settings and the experimental results demonstrate the effectiveness of the proposed AMM-GNN framework.
Ning Wang 0020, Minnan Luo, Kaize Ding, Lingling Zhang 0005, Jundong Li
CIKM2
2020 Cross-Graph Representation Learning for Unsupervised Graph Alignment
Weifan Wang 0004, Minnan Luo, Caixia Yan, Meng Wang 0009, Xiang Zhao 0002
DASFAA (2)2
2020 Unsupervised Hierarchical Feature Selection on Networked Data
Yuzhe Zhang 0003, Chen Chen 0022, Minnan Luo, Jundong Li, Caixia Yan
DASFAA (3)3
2020 Graph Representation Learning via Graphical Mutual Information Maximization
abstract
The richness in the content of various information networks such as social networks and communication networks provides the unprecedented potential for learning high-quality expressive representations without external supervision. This paper investigates how to preserve and extract the abundant information from graph-structured data into embedding space in an unsupervised manner. To this end, we propose a novel concept, Graphical Mutual Information (GMI), to measure the correlation between input graphs and high-level hidden representations. GMI generalizes the idea of conventional mutual information computations from vector space to the graph domain where measuring mutual information from two aspects of node features and topological structure is indispensable. GMI exhibits several benefits: First, it is invariant to the isomorphic transformation of input graphs—an inevitable constraint in many existing graph representation learning algorithms; Besides, it can be efficiently estimated and maximized by current mutual information estimation methods such as MINE; Finally, our theoretical analysis confirms its correctness and rationality. With the aid of GMI, we develop an unsupervised learning model trained by maximizing GMI between the input and output of a graph neural encoder. Considerable experiments on transductive as well as inductive node classification and link prediction demonstrate that our method outperforms state-of-the-art unsupervised counterparts, and even sometimes exceeds the performance of supervised ones.
Zhen Peng 0005, Wenbing Huang 0001, Minnan Luo, Yu Rong 0001, Tingyang Xu, Junzhou Huang
WWW3
2020 Scalable attack on graph data by injecting vicious nodes
Jihong Wang 0003, Minnan Luo, Fnu Suya, Jundong Li, Zijiang Yang 0006
Data Min. Knowl. Discov.2
2020 Dual-stream generative adversarial networks for distributionally robust zero-shot learning
Huan Liu 0012, Lina Yao 0001, Minnan Luo, Hongke Zhao, Yanzhang Lyu
Inf. Sci.4
2019 Heterogeneous Information Network Hashing for Fast Nearest Neighbor Search
Zhen Peng 0005, Minnan Luo, Jundong Li, Chen Chen 0022
DASFAA (1)2
2018 Top-k multi-class SVM using multiple features
Caixia Yan, Minnan Luo, Huan Liu 0012, Zhihui Li 0001
Inf. Sci.2
2012 A new algorithm for testing diagnosability of fuzzy discrete event systems
Minnan Luo, Yongming Li 0001, Fuchun Sun 0001, Huaping Liu 0001
Inf. Sci.1