VLDB 2026 Research / reviewers in the wild / expert
Wei Zhou 0019
dblp:69/5011-19
· DBLP profile ↗
75ranked-venue papers
5as first author
52since 2021 · last 2026
0000-0003-3622-3970ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 16 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLOBA: Rethinking Parameter Conflicts in Model MergingabstractModel merging serves as a training-free technique that combines multiple task-specific models into a unified multi-task model, but parameter conflicts often lead to performance drops. Previous methods flatten weight matrices into one-dimensional vectors, losing the inherent structural information of their row and column spaces. We mathematically prove and experimentally validate that parameter conflicts arise from non-orthogonal components of task vectors, while orthogonal components are conflict-free. Furthermore, we find that non-orthogonal components can contain both harmful conflicts and beneficial synergies. To precisely locate parameter conflicts and extract orthogonal components, we propose GLOBA (GLObal Basis Analysis Framework), which projects task vectors onto a global basis to align them within a unified coordinate system and construct a task interaction matrix. Following energy-based pruning, we divide parameters into five types based on the orthogonal relationships between the row spaces and column spaces of task vectors. Experiments on three fine-tuned models (mathematics, coding, and instruction-following) using LLaMA-2-7B and LLaMA-2-13B demonstrate significant performance gains through selective retention of beneficial parameters and removal of conflicting ones. Wei Zhou 0019 |
AAAI | 3 |
| 2026 | Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous AgentsabstractLarge Language Model (LLM)-based agents are increasingly deployed for complex, toolbased tasks where long-term memory is critical to driving actions.Existing benchmarks, however, primarily test an agent's ability to passively retrieve isolated facts in response to explicit questions.They fail to evaluate the more crucial capability of actively applying memory to execute tasks.To address this gap, we introduce MEM2ACTBENCH, a benchmark for evaluating whether agents can proactively leverage long-term memory to execute tool-based actions by selecting appropriate tools and grounding their parameters.The benchmark simulates persistent assistant usage, where users mention the same topic across long, interrupted interactions and expect previously established preferences and task states to be implicitly applied.We build the dataset with an automated pipeline that merges heterogeneous sources (ToolACE, BFCL, OASST1), resolves conflicts via consistency modeling, and synthesizes 2,029 sessions with 12 user-assistant-tool turns on average.From these memory chains, a reverse-generation method produces 400 tooluse tasks, with human evaluation confirming 91.3% are strongly memory-dependent.Experiments on seven memory frameworks show that current systems remain inadequate at actively utilizing memory for parameter grounding, highlighting the need for more effective approaches to evaluate and improve memory application in task execution. Yiting Shen, Wei Zhou 0019, Songlin Hu 0001 |
ACL (1) | 3 |
| 2026 | Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News DetectionabstractPropagation-based fake news detectors primarily extract structural patterns from news propagation trees via graph neural networks (GNNs), which are crucial for trustworthy information access on social platforms. However, these systems remain vulnerable to adversarial message injection, increasingly enabled by large language models (LLMs). Such attacks pollute both semantic and structural signals, causing GNN-based aggregators to fuse logically conflicting content and yield unreliable representations. To address this, we propose LLM-TKT, a novel framework that distills LLM-based narrative reasoning into lightweight GNNs for robust fake news detection. The framework operates in two stages. First, we construct an offline LLM-driven narrative hub to synthesize global propagation narratives and diagnose local node-level coherence. Second, we design a dual-level narrative alignment to learn the semantic invariance of propagation with the guidance of propagation narratives. It filters unreliable neighbor nodes via local consistency and optimizes graph representations via global anchoring. Experiments on three real-world datasets demonstrate that LLM-TKT significantly outperforms existing methods, particularly in defending against sophisticated LLM-driven injection attacks without incurring runtime LLM inference costs. Mengyang Chen, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
SIGIR | 3 |
| 2025 | An Information-theoretic Multi-task Representation Learning Framework for Natural Language UnderstandingabstractThis paper proposes a new principled multi-task representation learning framework (InfoMTL) to extract noise-invariant sufficient representations for all tasks. It ensures sufficiency of shared representations for all tasks and mitigates the negative effect of redundant features, which can enhance language understanding of pre-trained language models (PLMs) under the multi-task paradigm. Firstly, a shared information maximization principle is proposed to learn more sufficient shared representations for all target tasks. It can avoid the insufficiency issue arising from representation compression in the multi-task paradigm. Secondly, a task-specific information minimization principle is designed to mitigate the negative effect of potential redundant features in the input for each task. It can compress task-irrelevant redundant information and preserve necessary information relevant to the target for multi-task prediction. Experiments on six classification benchmarks show that our method outperforms 12 comparative multi-task methods under the same multi-task settings, especially in data-constrained and noisy scenarios. Extensive experiments demonstrate that the learned representations are more sufficient, data-efficient, and robust. Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
AAAI | 3 |
| 2025 | Enhancing Multi-Hop Fact Verification with Structured Knowledge-Augmented Large Language ModelsabstractThe rapid development of social platforms exacerbates the dissemination of misinformation, which stimulates the research in fact verification. Recent studies tend to leverage semantic features to solve this problem as a single-hop task. However, the process of verifying a claim requires several pieces of evidence with complicated inner logic and relations to verify the given claim in real-world situations. Recent studies attempt to improve both understanding and reasoning abilities to enhance the performance, but they overlook the crucial relations between entities that benefit models to understand better and facilitate the prediction. To emphasize the significance of relations, we resort to Large Language Models (LLMs) considering their excellent understanding ability. Instead of other methods using LLMs as the predictor, we take them as relation extractors, for they do better in understanding rather than reasoning according to the experimental results. Thus, to solve the challenges above, we propose a novel Structured Knowledge-Augmented LLM-based Network (LLM-SKAN) for multi-hop fact verification. Specifically, we utilize an LLM-driven Knowledge Extractor to capture fine-grained information, including entities and their complicated relations. Besides, we leverage a Knowledge-Augmented Relation Graph Fusion module to interact with each node and learn better claim-evidence representations comprehensively. The experimental results on four common-used datasets demonstrate the effectiveness and superiority of our model. Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
AAAI | 3 |
| 2025 | BotSim: LLM-Powered Malicious Social Botnet SimulationabstractSocial media platforms like X(Twitter) and Reddit are vital to global communication. However, advancements in Large Language Model (LLM) technology give rise to social media bots with unprecedented intelligence. These bots adeptly simulate human profiles, conversations, and interactions, disseminating large amounts of false information and posing significant challenges to platform regulation. To better understand and counter these threats, we innovatively design BotSim, a malicious social botnet simulation powered by LLM. BotSim mimics the information dissemination patterns of real-world social networks, creating a virtual environment composed of intelligent agent bots and real human users. In the temporal simulation constructed by BotSim, these advanced agent bots autonomously engage in social interactions such as posting and commenting, effectively modeling scenarios of information flow and user interaction. Building on the BotSim framework, we construct a highly human-like, LLM-driven bot dataset called BotSim-24 and benchmark multiple bot detection strategies against it. The experimental results indicate that detection methods effective on traditional bot datasets perform worse on BotSim-24, highlighting the urgent need for new detection strategies to address the cybersecurity threats posed by these advanced bots. Boyu Qiao, Wei Zhou 0019, Shilong Li 0003, Qianqian Lu, Songlin Hu 0001 |
AAAI | 3 |
| 2025 | Impartial Multi-task Representation Learning via Variance-invariant Probabilistic DecodingabstractMulti-task learning (MTL) enhances efficiency by sharing representations across tasks, but task dissimilarities often cause partial learning, where some tasks dominate while others are neglected.Existing methods mainly focus on balancing loss or gradients but fail to fundamentally address this issue due to the representation discrepancy in latent space.In this paper, we propose variance-invariant probabilistic decoding for multi-task learning (VIP-MTL), a framework that ensures impartial learning by harmonizing representation spaces across tasks.VIP-MTL decodes shared representations into task-specific probabilistic distributions and applies variance normalization to constrain these distributions to a consistent scale.Experiments on two language benchmarks show that VIP-MTL outperforms 12 representative methods under the same multi-task settings, especially in heterogeneous task combinations and dataconstrained scenarios.Further analysis shows that VIP-MTL is robust to sampling distributions, efficient on optimization process, and scale-invariant to task losses.Additionally, the learned task-specific representations are more informative, enhancing the language understanding abilities of pre-trained language models under the multi-task paradigm. Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
ACL (1) | 3 |
| 2025 | Capture the Key in Reasoning to Enhance CoT Distillation GeneralizationabstractAs Large Language Models (LLMs) scale up and gain powerful Chain-of-Thoughts (CoTs) reasoning abilities, practical resource constraints drive efforts to distill these capabilities into more compact Smaller Language Models (SLMs).We find that CoTs consist mainly of simple reasoning forms, with a small proportion (≈ 4.7%) of key reasoning steps that truly impact conclusions.However, previous distillation methods typically involve supervised fine-tuning student SLMs only on correct CoTs data produced by teacher LLMs, resulting in students struggling to learn the key, instead imitating the teacher's reasoning forms and making errors or omissions in reasoning.To address these issues, drawing an analogy to human learning, where analyzing mistakes according to correct solutions often reveals the crucial steps leading to successes or failures, we propose mistakE-Driven key reasonIng step distillaTion (EDIT), a novel method that further aids SLMs learning key reasoning steps rather than mere simple fine-tuning.Firstly, to expose the crucial steps in CoTs, we carefully design specific prompts to generate dual CoTs data with similar reasoning paths but divergent conclusions.Then, we apply the minimum edit distance algorithm on the dual CoTs data to locate these key steps and optimize the likelihood on these tokens.Extensive experiments and analysis validate the effectiveness of EDIT across both in-domain(IND) and out-ofdomain(OOD) benchmark reasoning datasets 1 .* Kun Li is the corresponding author. 1 Code can be found at https://github.com/ C-W-D/EDIT On the floor, I see two magenta scrunchiephone chargers, one gold fidget spinner, three gold booklets, three silver scrunchiephone chargers, two silver booklets, and two magenta booklets.If I remove all the magenta items from the floor, how many scrunchiephone chargers remain on it?Op ons: (A) zero (B) one (C) two (D) three (E) four (F) five (G) six According to this ques on, there are two magenta scrunchiephone chargers and three silver scrunchiephone chargers on the floor.If we remove all the magenta items from the floor, we are le with 2 silver scrunchiephone chargers.Therefore, the answer is (C).According to this ques on, there are two magenta scrunchiephone chargers and three silver scrunchiephone chargers on the floor.If we remove all the magenta items, we are le with only the silver scrunchiephone chargers.Therefore, the number of scrunchiephone chargers that remain on the floor is three.Therefore, the answer is (D).Teacher Ques on Unthinking Student On the floor, I see two magenta scrunchiephone chargers, one gold fidget spinner...According to this ques on, there are two magenta scrunchiephone chargers and three silver scrunchiephone chargers on the floor.If we remove all the magenta items from the floor, we are le with 2 silver scrunchiephone chargers.Therefore, the answer is (C).According to this ques on, there are a total of 5 scrunchiephone chargers on the floor: 2 magentascrunchiephone chargers and 3 silver scrunchiephone chargers.If we remove all the magenta items from the floor, we remove the 2 magenta scrunchiephone chargers.Therefore, the number of scrunchiephone chargers remaining on the floor is 3. Therefore, the answer is (D). Key Steps in Correct CoTs Data Key Steps in Wrong CoTs Data Chengwei Dai, Wei Zhou 0019, Songlin Hu 0001 |
ACL (1) | 3 |
| 2025 | MPPFND: A Dataset and Analysis of Detecting Fake News with Multi-Platform Propagation
Congyuan Zhao, Lingwei Wei, Ziming Qin, Wei Zhou 0019, Yunya Song, Songlin Hu 0001 |
CogSci | 4 |
| 2025 | DORA: Dynamic Optimization Prompt for Continuous Reflection of LLM-based AgentabstractAutonomous agents powered by large language models (LLMs) hold significant potential across various domains. The Reflection framework is designed to help agents learn from past mistakes in complex tasks. While previous research has shown that reflection can enhance performance, our investigation reveals a key limitation: meaningful self-reflection primarily occurs at the beginning of iterations, with subsequent attempts failing to produce further improvements. We term this phenomenon “Early Stop Reflection,” where the reflection process halts prematurely, limiting the agent’s ability to engage in continuous learning. To address this, we propose the DORA method (Dynamic and Optimized Reflection Advice), which generates task-adaptive and diverse reflection advice. DORA introduces an external open-source small language model (SLM) that dynamically generates prompts for the reflection LLM. The SLM uses feedback from the agent and optimizes the prompt generation process through a non-gradient Bayesian Optimization (BO) algorithm, ensuring the reflection process evolves and adapts over time. Our experiments in the MiniWoB++ and Alfworld environments confirm that DORA effectively mitigates the “Early Stop Reflection” issue, enabling agents to maintain iterative improvements and boost performance in long-term, complex tasks. Code are available at https://anonymous.4open.science/r/DORA-44FB/. Tingzhang Zhao, Wei Zhou 0019, Songlin Hu 0001 |
COLING | 3 |
| 2025 | APEE: Assessing the Personality Expressions of LLM-Driven Role Play Agent Beyond Self-PerceptionabstractLarge language models (LLMs) have demonstrated significant progress in role-playing tasks, yet evaluating their ability to simulate personality traits remains a challenge. Traditional psychological questionnaires-based method have been used to assess LLMs' personality traits. However, these approaches have limitations when applied to LLM-driven role-playing agents (RPAs), as they are designed for humans and rely on stable, self-assessed personality traits. To bridge this gap, we extend simple self-perception questionnaires to more objective, real-world evaluations. In this paper, we introduce APEE, a new dataset consisting of 473 instances across three real-world scenario types: practical goal planning, social media behavior, and leaderless group discussions (LGD). In addition to evaluating whether LLMs adhere to predefined character traits, we introduce two key metrics: Stability and Differentiation. These metrics assess how consistently LLMs express personality traits across different scenarios (Stability) and how effectively they differentiate their behavior when assuming multiple roles (Differentiation). We conducted experiments on 339 different roles using 11 advanced LLMs with the APEE dataset. Discuss the impact of factors such as model size and architecture. Code and dataset are available at https://github.com/linkseed18612254945/APEE_Personality. Chenwei Dai, Wei Zhou 0019, Songlin Hu 0001 |
CSCWD | 3 |
| 2025 | GRAgent: A Generative Retrieval Framework for Action Subspace SelectionabstractAs Large Language Models (LLMs) are increasingly deployed as autonomous agents to accomplish complex real-world tasks, they must select appropriate actions from massive action spaces. However, in open-domain settings, LLMs often generate hallucinations when selecting from large action spaces due to their lack of practical operation experience. While existing approaches attempt to reduce action spaces through external planning mechanisms, they heavily rely on substantial training data from executable environments. To address this challenge, we propose a generative retrieval-based methodology for identifying necessary action subspaces. Our approach leverages the autoregressive capabilities of pre-trained models, enabling better utilization of knowledge and superior performance in low-data scenarios, while providing interpretable retrieval results through generated action encodings. We design a two-phase training framework that combines self-learning generative retrieval with targeted optimization. Through comprehensive experiments on multiple benchmark datasets, our approach demonstrates superior retrieval performance while effectively reducing model hallucinations and enhancing plan feasibility. Results show that our method consistently outperforms existing approaches across various metrics. Minxuan Lv, Wei Zhou 0019, Songlin Hu 0001 |
CSCWD | 3 |
| 2025 | DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMsabstractMinxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong, Zijia Lin, Hui Chen, Wei Zhou, Jungong Han, Guiguang Ding, Wenwu Ou, Di Zhang, Kun Gai, Songlin Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Minxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong, Zijia Lin, Hui Chen 0013, Wei Zhou 0019, Jungong Han, Guiguang Ding, Wenwu Ou, Di Zhang 0026, Kun Gai, Songlin Hu 0001 |
EMNLP | 7 |
| 2025 | Identifying Bots on Social Media through Coordinated Group PerceptionabstractIdentifying bots on social media has become a crucial and challenging task for regulating online discourse. Existing detection methods primarily focus on individual account-level information, identifying potential threats by detecting inconsistencies between genuine humans and anomalous bots in personal profiles, textual content, and social relationships. However, these approaches generally overlook the coordinated behavior characteristics inherent in groups of bot accounts. To address this research gap, we propose a novel Bot detection network based on Coordinated Group Perception (BotCGP), which enhances bot identification performance by uncovering the collective coordinated features among bot groups. Specifically, our method jointly models account profiles, textual content, and social relationships using a Student’s t-distribution kernel function and a differentiable modularity function to capture potential coordinated characteristics. Experimental results demonstrate that BotCGP significantly outperforms existing methods in bot detection across three real-world X/Twitter datasets. Our code is available at https://github.com/QQQQQQBY/BotCGP. Boyu Qiao, Wei Zhou 0019, Shilong Li 0003, Qianqian Lu, Songlin Hu 0001 |
ICASSP | 3 |
| 2025 | ComSFI: A Community-Aware Approach to Early Rumor Propagation PredictionabstractInformation propagation prediction in social networks remains challenging due to the complex interactions between community structures. Existing models typically overlook multi-peaked cascade patterns that emerge from community-specific dynamics, especially during critical early propagation stages. To address these limitations, we propose ComSFI (Community-aware Susceptible-Forwarding-Immune), a dynamic prediction framework that captures both intra- and inter-community information diffusion. ComSFI makes three key contributions: (1) community-specific diffusion modeling through individualized SFI modules, (2) cross-community propagation modeling with a learnable threshold mechanism that identifies cascade initiation timing, and (3) integration of user interest profiles to estimate propagation willingness across community boundaries. Experiments on Weibo, Douban, and Memetracker datasets demonstrate that ComSFI outperforms state-of-the-art baselines by 3%-12% in cascade size prediction and temporal pattern accuracy, with particular effectiveness in early-stage rumor propagation prediction—a critical application for timely intervention. Our results establish ComSFI as a versatile framework for analyzing and predicting complex information diffusion patterns in real-world networks. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 2 |
| 2025 | Dispelling the Fake: Social Bot Detection Based on Edge Confidence EvaluationabstractSocial bot detection is essential for maintaining the safety and integrity of online social networks (OSNs). Graph neural networks (GNNs) have emerged as a promising solution. Mainstream GNN-based social bot detection methods learn rich user representations by recursively performing message passing along user-user interaction edges, where users are treated as nodes and their relationships as edges. However, these methods face challenges when detecting advanced bots interacting with genuine accounts. Interaction with real accounts results in the graph structure containing camouflaged and unreliable edges. These unreliable edges interfere with the differentiation between bot and human representations, and the iterative graph encoding process amplifies this unreliability. In this article, we propose a social Bot detection method based on Edge Confidence Evaluation (BECE). Our model incorporates an edge confidence evaluation module that assesses the reliability of the edges and identifies the unreliable edges. Specifically, we design features for edges based on the representation of user nodes and introduce parameterized Gaussian distributions to map the edge embeddings into a latent semantic space. We optimize these embeddings by minimizing Kullback-Leibler (KL) divergence from the standard distribution and evaluate their confidence based on edge representation. Experimental results on three real-world datasets demonstrate that BECE is effective and superior in social bot detection. Additionally, experimental results on six widely used GNN architectures demonstrate that our proposed edge confidence evaluation module can be used as a plug-in to improve detection performance. Boyu Qiao, Wei Zhou 0019, Shilong Li 0003, Songlin Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Structured Probabilistic CodingabstractThis paper presents a new supervised representation learning framework, namely structured probabilistic coding (SPC), to learn compact and informative representations from input related to the target task. SPC is an encoder-only probabilistic coding technology with a structured regularization from the target space. It can enhance the generalization ability of pre-trained language models for better language understanding. Specifically, our probabilistic coding simultaneously performs information encoding and task prediction in one module to more fully utilize the effective information from input data. It uses variational inference in the output space to reduce randomness and uncertainty. Besides, to better control the learning process of probabilistic representations, a structured regularization is proposed to promote uniformity across classes in the latent space. With the regularization term, SPC can preserve the Gaussian structure of the latent code and achieve better coverage of the hidden space with class uniformly. Experimental results on 12 natural language understanding tasks demonstrate that our SPC effectively improves the performance of pre-trained language models for classification and regression. Extensive experiments show that SPC can enhance the generalization capability, robustness to label noise, and clustering quality of output representations. Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
AAAI | 4 |
| 2024 | Representation Learning with Conditional Information Flow MaximizationabstractThis paper proposes an information-theoretic representation learning framework, named conditional information flow maximization, to extract noise-invariant sufficient representations for the input data and target task.It promotes the learned representations have good feature uniformity and sufficient predictive ability, which can enhance the generalization of pre-trained language models (PLMs) for the target task.Firstly, an information flow maximization principle is proposed to learn more sufficient representations for the input and target by simultaneously maximizing both inputrepresentation and representation-label mutual information.Unlike the information bottleneck, we handle the input-representation information in an opposite way to avoid the overcompression issue of latent representations.Besides, to mitigate the negative effect of potential redundant features from the input, we design a conditional information minimization principle to eliminate negative redundant features while preserve noise-invariant features.Experiments on 13 language understanding benchmarks demonstrate that our method effectively improves the performance of PLMs for classification and regression.Extensive experiments show that the learned representations are more sufficient, robust and transferable. Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
ACL (1) | 3 |
| 2024 | Multi-stream Information Fusion Framework for Emotional Support ConversationabstractEmotional support conversation (ESC) task aims to relieve the emotional distress of users who have high-intensity of negative emotions. However, due to the ignorance of emotion intensity modelling which is essential for ESC, previous methods fail to capture the transition of emotion intensity effectively. To this end, we propose a Multi-stream information Fusion Framework (MFF-ESC) to thoroughly fuse three streams (text semantics stream, emotion intensity stream, and feedback stream) for the modelling of emotion intensity, based on a designed multi-stream fusion unit. As the difficulty of modelling subtle transitions of emotion intensity and the strong emotion intensity-feedback correlations, we use the KL divergence between feedback distribution and emotion intensity distribution to further guide the learning of emotion intensities. Experimental results on automatic and human evaluations indicate the effectiveness of our method. Yinan Bao, Dou Hu 0001, Lingwei Wei, Shuchong Wei, Wei Zhou 0019, Songlin Hu 0001 |
LREC/COLING | 5 |
| 2024 | Adaptive Mixture of Domain-aware Experts for Detecting Social BotsabstractSocial bot detection has received widely attention from academic and industrial communities. However, existing bot detection methods are far from perfect. To mimic genuine users on social networks, advanced social bots are often active in multiple domains and have mixed characteristics of multiple domains. It is unreasonable to classify a bot with only one domain. To effectively extract and fuse features from multiple domains, we propose a novel method for Domain-aware Social Bot Detection (DSBD). Specifically, we first use a prompt-based method for zero-shot domain classification to obtain accurate domain distribution for any user. We then aggregate multiple domain expert representations through a domain gate, and finally use the fused representation to classify. Experimental results show that our approach consistently outperforms all baselines and that our fusion strategy perform well in various settings especially zero-shot situation. Qianqian Lu, Shilong Li 0003, Wei Zhou 0019, Liangjun Zang |
CSCWD | 4 |
| 2024 | Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs DistillationabstractLarge language models (LLMs) exhibit enhanced reasoning at larger scales, driving efforts to distill these capabilities into smaller models via teacher-student learning.Previous works simply fine-tune student models on teachers' generated Chain-of-Thoughts (CoTs) data.Although these methods enhance indomain (IND) reasoning performance, they struggle to generalize to out-of-domain (OOD) tasks.We believe that the widespread spurious correlations between questions and answers may lead the model to preset a specific answer which restricts the diversity and generalizability of its reasoning process.In this paper, we propose Cascading Decomposed CoTs Distillation (CasCoD) to address these issues by decomposing the traditional single-step learning process into two cascaded learning steps.Specifically, by restructuring the training objectives-removing the answer from outputs and concatenating the question with the rationale as input-CasCoD's two-step learning process ensures that students focus on learning rationales without interference from the preset answers, thus improving reasoning generalizability.Extensive experiments demonstrate the effectiveness of CasCoD on both IND and OOD benchmark reasoning datasets 1 .* Kun Li is the corresponding author. 1 Code available at https://github.com/C-W-D/CasCoD(a) Answer SFT consistently outperform Std-CoT on OOD tasks.(b) A case of spurious correla on between ques ons and answers.Question: Why did someone bring a swimsuit to a ski resort?Options: (A) To swim in a heated pool.(B) To wear as an underlayer for warmth.(C) To use as a fashion statement.(D) To participate in a polar bear plunge event.Answer: (A) To swim in a heated pool. OOD IND Chengwei Dai, Wei Zhou 0019, Songlin Hu 0001 |
EMNLP | 3 |
| 2024 | Transferring Structure Knowledge: A New Task to Fake News Detection towards Cold-Start PropagationabstractMany fake news detection studies have achieved promising performance by extracting effective semantic and structure features from both content and propagation trees. However, it is challenging to apply them to practical situations, especially when using the trained propagation-based models to detect news with no propagation data. Towards this scenario, we study a new task named cold-start fake news detection, which aims to detect content-only samples with missing propagation. To achieve the task, we design a simple but effective Structure Adversarial Net (SAN) framework to learn transferable features from available propagation to boost the detection of content-only samples. SAN introduces a structure discriminator to estimate dissimilarities among learned features with and without propagation, and further learns structure-invariant features to enhance the generalization of existing propagation-based methods for content-only samples. We conduct qualitative and quantitative experiments on three datasets. Results show the challenge of the new task and the effectiveness of our SAN framework. Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Songlin Hu 0001 |
ICASSP | 3 |
| 2024 | Multi-source Knowledge Enhanced Graph Attention Networks for Multimodal Fact VerificationabstractMultimodal fact verification is an under-explored and emerging field that has gained increasing attention in recent years. The goal is to assess the veracity of claims that involve multiple modalities by analyzing the retrieved evidence. The main challenge in this area is to effectively fuse features from different modalities to learn meaningful multimodal representations. To this end, we propose a novel model named Multi-Source Knowledge-enhanced Graph Attention Network (MultiKE-GAT). MultiKE-GAT introduces external multimodal knowledge from different sources and constructs a heterogeneous graph to capture complex cross-modal and cross-source interactions. We exploit a Knowledge-aware Graph Fusion (KGF) module to learn knowledge-enhanced representations for each claim and evidence and eliminate inconsistencies and noises introduced by redundant entities. Experiments on two public benchmark datasets demonstrate that our model outperforms other comparison methods, showing the effectiveness and superiority of the proposed model. Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
ICME | 3 |
| 2024 | CSFI for Social Media: Understanding and Predicting Cross-Community Information PropagationabstractSocial platform users are intricately interconnected, forming a complex social system. Personalized recommendations help users access information from diverse sources, fostering community communication. Micro-level dynamic analysis technology offers a scientific approach to measuring and predicting information transmission. While fundamental dissemination principles are studied, cross-community communication scenarios are under-researched. This study presents a community-centric communication dynamics model (CSFI) to explain information transmission within and across communities on social media. Tested on Sina Weibo data, the model improves retweet prediction accuracy by 11.3 % over baselines. Accurately predicting information propagation on social media is crucial for public opinion analysis. This research aids in predicting dissemination trajectories, reflecting public sentiment, and guiding targeted information strategies or interventions. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 2 |
| 2024 | Information Diffusion Prediction with Graph Neural Ordinary Differential Equation NetworkabstractInformation diffusion prediction aims to forecast the path of information spreading in social networks by exploiting user correlations or preferences. Recent works focus on characterizing the dynamic of user preferences and propose to capture users' dynamic preferences by discretizing the diffusion process into structure snapshots. Despite their effectiveness, these works simply summarize users' dynamic preferences from partially observed structure snapshots, ignoring the continuous evolution of the preferences. Moreover, discretizing the diffusion process makes these models overlook abundant structure information across different periods, reducing their ability to discover potential participants. To address the above issues, we propose a novel Graph Neural Ordinary Differential Equation Network (GODEN) for information diffusion prediction, which incorporates neural ordinary differential equations (ODE) to model the continuous dynamics of the diffusion process. Specifically, we design two coupled ODE functions on nodes and edges to describe their co-evolution dynamics and infer users' dynamic preferences based on the solution of ODEs. To predict the future infections of the observed cascade, we represent its diffusion pattern in terms of temporal and user contexts and apply a multi-head attention module to attend to different contexts. Experimental results confirm our approach's effectiveness, with our model outperforming the state-of-the-art diffusion prediction models. Wei Zhou 0019, Songlin Hu 0001 |
ACM Multimedia | 2 |
| 2024 | Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue SystemsabstractZhenpeng Su, Xing W, Wei Zhou, Guangyuan Ma, Songlin Hu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhenpeng Su, Xing Wu 0002, Wei Zhou 0019, Guangyuan Ma, Songlin Hu 0001 |
NAACL-HLT | 3 |
| 2024 | Propagation Structure-Semantic Transfer Learning for Robust Fake News Detection
Mengyang Chen, Lingwei Wei, Wei Zhou 0019, Zhou Yan, Songlin Hu 0001 |
ECML/PKDD (7) | 4 |
| 2024 | Knowledge Enhanced Vision and Language Model for Multi-Modal Fake News DetectionabstractThe rapid dissemination of fake news and rumors through the Internet and social media platforms poses significant challenges and raises concerns in the public sphere. Automatic detection of fake news plays a crucial role in mitigating the spread of misinformation. While recent approaches have focused on leveraging neural networks to improve textual and visual representations in multi-modal fake news analysis, they often overlook the potential of incorporating knowledge information to verify facts within news articles. In this paper, we propose a knowledge enhanced vision and language model for multi-modal fake news detection. Our proposed model integrates information from large scale open knowledge graphs to augment its ability to discern the veracity of news content. Unlike previous methods that utilize separate models to extract textual and visual features, we synthesize a unified model capable of extracting both types of features simultaneously. To represent news articles, we introduce a graph structure where nodes encompass entities, relationships extracted from the textual content, and objects depicted in associated images. By utilizing the knowledge graph, we establish meaningful relationships between nodes within the news articles. Experimental evaluations on a real-world multi-modal dataset from Twitter demonstrate significant performance improvement by incorporating knowledge information. Xingyu Gao 0001, Xi Wang 0014, Zhenyu Chen 0003, Wei Zhou 0019, Steven C. H. Hoi |
IEEE Trans. Multim. | 4 |
| 2024 | Modeling the Uncertainty of Information Propagation for Rumor Detection: A Neuro-Fuzzy ApproachabstractAutomatic rumor detection is critical for maintaining a healthy social media environment. The mainstream methods generally learn rich features from information cascades by modeling the cascade as a tree or graph structure where edges are built based on interactions between a tweet and retweets. Some psychology studies have empirically shown that users' various subjective factors always cause the uncertainty of interactions such as differences among interactive behavior activation thresholds or semantic relevancy. However, previous works model interactions by employing a simple fully connected layer on fixed edge weights in the graph and cannot reasonably describe this inherent uncertainty of complex interactions. In this article, inspired by the fuzzy theory, we propose a novel neuro-fuzzy method, fuzzy graph convolutional networks (FGCNs), to sufficiently understand uncertain interactions in the information cascade in a fuzzy perspective. Specifically, a new strategy of graph construction is first designed to convert each information cascade into a heterogeneous graph structure with the consideration of explicit interactive behaviors between a tweet and its retweet, as well as implicit interactive behaviors among retweets, enriching more structural clues in the graph. Then, we improve graph convolutional networks by incorporating edge fuzzification (EF) modules. The EFs adapt edge weights according to predefined membership to enhance message passing in the graph. The proposed model can provide a stronger relational inductive bias for expressing uncertain interactions and capture more discriminative and robust structural features for rumor detection. Extensive experiments demonstrate the effectiveness and superiority of FGCN on both rumor detection and early rumor detection. Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xin Wang 0086, Songlin Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Supervised Adversarial Contrastive Learning for Emotion Recognition in ConversationsabstractExtracting generalized and robust representations is a major challenge in emotion recognition in conversations (ERC).To address this, we propose a supervised adversarial contrastive learning (SACL) framework for learning classspread structured representations in a supervised manner.SACL applies contrast-aware adversarial training to generate worst-case samples and uses joint class-spread contrastive learning to extract structured representations.It can effectively utilize label-level feature consistency and retain fine-grained intra-class features.To avoid the negative impact of adversarial perturbations on context-dependent data, we design a contextual adversarial training (CAT) strategy to learn more diverse features from context and enhance the model's context robustness.Under the framework with CAT, we develop a sequence-based SACL-LSTM to learn label-consistent and context-robust features for ERC.Experiments on three datasets show that SACL-LSTM achieves state-of-the-art performance on ERC.Extended experiments prove the effectiveness of SACL and CAT. Dou Hu 0001, Yinan Bao, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
ACL (1) | 4 |
| 2023 | MeaeQ: Mount Model Extraction Attacks with Efficient QueriesabstractWe study model extraction attacks in natural language processing (NLP) where attackers aim to steal victim models by repeatedly querying the open Application Programming Interfaces (APIs).Recent works focus on limited-query budget settings and adopt random sampling or active learning-based sampling strategies on publicly available, unannotated data sources.However, these methods often result in selected queries that lack task relevance and data diversity, leading to limited success in achieving satisfactory results with low query costs.In this paper, we propose MeaeQ (Model extraction attack with efficient Queries), a straightforward yet effective method to address these issues.Specifically, we initially utilize a zero-shot sequence inference classifier, combined with API service information, to filter task-relevant data from a public text corpus instead of a problem domain-specific dataset.Furthermore, we employ a clustering-based data reduction technique to obtain representative data as queries for the attack.Extensive experiments conducted on four benchmark datasets demonstrate that MeaeQ achieves higher functional similarity to the victim model than baselines while requiring fewer queries.Our code is available at https://github.com/C-W-D/MeaeQ. Chengwei Dai, Minxuan Lv, Wei Zhou 0019 |
EMNLP | 4 |
| 2023 | CT-GAT: Cross-Task Generative Adversarial Attack based on TransferabilityabstractNeural network models are vulnerable to adversarial examples, and adversarial transferability further increases the risk of adversarial attacks.Current methods based on transferability often rely on substitute models, which can be impractical and costly in real-world scenarios due to the unavailability of training data and the victim model's structural details.In this paper, we propose a novel approach that directly constructs adversarial examples by extracting transferable features across various tasks.Our key insight is that adversarial transferability can extend across different tasks.Specifically, we train a sequence-to-sequence generative model named CT-GAT (Cross-Task Generative Adversarial ATtack) using adversarial sample data collected from multiple tasks to acquire universal adversarial features and generate adversarial examples for different tasks.We conduct experiments on ten distinct datasets, and the results demonstrate that our method achieves superior attack performance with small cost.You can get our code and data at: https: //github.com/xiaoxuanNLP/CT-GAT Minxuan Lv, Chengwei Dai, Wei Zhou 0019, Songlin Hu 0001 |
EMNLP | 4 |
| 2023 | Semantic Stage-Wise Learning for Knowledge DistillationabstractKnowledge distillation enhances the performance of the student model by transferring knowledge from the teacher model. Moreover, the attention mechanism has been introduced recently to enable each layer of the student to learn knowledge from all teacher layers, which brings about considerable optimization. However, noted that features from different layers, such as shallow and deep layers, might have a big semantic gap, and compulsively aligning one student layer to all teacher layers would mislead the learning process. To tackle this problem, an effective framework called Semantic Stage-Wise learning for Knowledge Distillation (SSWKD) is presented in this paper. We divide all layers into shallow and deep stages, and only allow feature alignment within the same stage to alleviate semantic mismatch. In addition, with the observation that the performance of deep networks relies more on some key features rather than evenly on all of them, a crucial feature enhancement method based on KL divergence is then proposed for SSWKD, forcing the student to pay more attention to critical features of the teacher. Extensive experiments and visualizations show that our SSWKD outperforms other distillation methods on CIFAR-100 and COCO2017 datasets for image classification, object detection, and instance segmentation tasks. Dongqin Liu, Wei Zhou 0019, Zhaoxing Li, Jiao Dai, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001 |
ICME | 3 |
| 2023 | Social Bot Detection Based on Window StrategyabstractWith social bots evolving continually, the new bots post highly anthropomorphic posts to evade detection. For post content information, existing methods choose to extract features from single posts or use the rough characterization of the pretraining language model for bot detection. However, the evolution of bots has made them better at camouflage in previous posts processing techniques. According to our observation, the purpose of bot posting is to publicize different contents in different periods, and its posting often shows abnormal interest changes, while human posting is to express hobbies and daily life, and its interest changes are relatively stable. Therefore, we extract the interest changes between multiple posts published by users for bot detection. In this paper, we propose a social Bot detection model based on the Window Strategy(BotWS). Specifically, we first employ the window strategy to obtain the user posting windows, each containing multiple posts. Then, we extract the interest changes between the posting windows by multi-head attention mechanism. Finally, we embed the interest changes information into the user representation and construct a heterogeneous graph classification module to classify. We conduct experiments on three challenging datasets. Results indicate that our approach achieves state-of-the-art. Boyu Qiao, Wei Zhou 0019, Zhou Yan, Shilong Li 0003, Songlin Hu 0001 |
ICME | 3 |
| 2023 | Multi-modal Social Bot Detection: Learning Homophilic and Heterophilic Connections AdaptivelyabstractThe detection of social bots has become a critical task in maintaining the integrity of social media. With social bots evolving continually, they primarily evade detection by imitating human features and engaging in interactions with humans. To reduce the impact of social bots imitating human features, also known as feature camouflage, existing methods mainly utilize multi-modal user information for detection, especially GNN-based methods that utilize additional topological structure information. However, these methods ignore relation camouflage, which involves disguising through interactions with humans. We find that relation camouflage results in both homophilic connections formed by nodes of the same type and heterophilic connections formed by nodes of different types in social networks. The existing GNN-based detection methods assume all connections are homophilic while ignoring the difference among neighbors in heterophilic connections, which leads to a poor detection performance for bots with relation camouflage. To address this, we propose a multi-modal social bot detection method with learning homophilic and heterophilic connections adaptively (BothH for short). Specifically, firstly we determine whether each connection is homophilic or heterophilic with the connection classifier, and then we design a novel message propagating strategy that can learn the homophilic and heterophilic connections adaptively. We conduct experiments on the mainstream datasets and the results show that our model is superior to state-of-the-art methods. Shilong Li 0003, Boyu Qiao, Qianqian Lu, Wei Zhou 0019 |
ACM Multimedia | 6 |
| 2023 | HIM: An End-to-End Hierarchical Interaction Model for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) is an emerging task of fine-grained sentiment analysis, which aims to extract aspect terms, associated opinion terms, and sentiment polarities in the form of triplets. Thus, ASTE involves two groups of subtasks: aspect/opinion term extraction and aspect-opinion-pair sentiment classification. Due to the high correlations of subtasks, three categories of joint methods have been proposed, includingend-to-end tagging-based methods,cascaded span-based methods, andsequence-to-sequence generation-based methods. These methods basically learn either a shared feature space or a shared sentence encoder to capture interactions across all subtasks by parameter sharing. However, they fail to learn deep and mutual interactive features for ASTE. In this work, we present a novel tagging scheme to cast ASTE as a unified boundary-words relation classification problem. Subsequently, we propose an end-to-end Hierarchical Interaction Model (HIM), exploiting deep and mutual interactions across subtasks mainly with two interaction modules. The first-level interaction module primarily leverages multi-task learning models to capture implicit subtask interactions. Then, the second-level interaction module, namely Gated Interaction Network (GIN), adopts a novel gated control mechanism and a newly-designed Conditional BiLSTM (Cond-BiLSTM) network to capture explicit subtask interactions. Moreover, to refine the unreliable outputs of the first-level module, we develop a General word-Pair Relationship Learning (G-PRL) component. With the task-shared features as input, G-PRL further facilitates interactions between term extraction and pair classification. We conduct experiments on two benchmarks and achieve promising results. Extensive analyses demonstrate the effectiveness and flexibility of our work. Wei Zhou 0019, Songlin Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Mining Weak Relations Between Reviews for Opinion Spam DetectionabstractOnline reviews play a significant role in purchase decisions of consumers by providing feedback information from buyers of products. In order to mislead consumers, opinion spammers are hired to write fake reviews to promote or demote specific products for illegitimate benefits. Existing methods for spam review detection mainly focused on designing manual features, which highly rely on expert knowledge. Although recent works utilized deep learning methods to automatically learn the semantics of reviews through the inherent user-review-product strong relation, they fail to capture the weak relations between reviews at the content, sentiment and temporal levels, which provides various semantic information to expose fake reviews. Moreover, the imbalanced class distribution in spam detection issues makes this work even more challenging. To address the above problems, we propose a novel Weak-Strong Unified Network (WSUN) for opinion spam detection. Multi-level weak relation graphs are constructed to reveal the abnormal behavioral patterns of spammers, which aggregates the semantics of strong relations by graph convolutional networks and extracts comprehensive review representation by utilizing relation-level attention mechanism. In addition, a graph-based over-sampling method is devised to mitigate the impact of imbalanced class distribution. Extensive experimental results on real-world datasets show that our model is more effective than the state-of-the-art methods. Yingrui Xu, Jingguo Ge, Xiaodan Zhang 0004, Yulei Wu, Honglei Lv, Hongbin Shi, Wei Zhou 0019 |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2023 | Modeling Both Intra- and Inter-Modality Uncertainty for Multimodal Fake News DetectionabstractMultimodal fake news detection has obtained increasing attention recently. Existing works generally encode multimodal contents into a deterministic point in semantic subspaces, and then fuse multimodal features by simple concatenation or attention mechanisms. However, most methods suffer from adapting to noisy multimodal contents since they neglect the robustness of modality-specific features. Besides, as different modalities usually have varying confidence levels, previous attention-based fusion models that learn modality-independent weights based on the input data feature, would limit the optimal integration of multimodal contents. To alleviate the above issues, we propose novel Multimodal Uncertainty Learning Network (MM-ULN) to enhance multimodal fake news detection by modeling both intra- and inter-modality uncertainty. Specifically, we incorporate a novel intra-modality uncertainty learning (EUL) module to better understand noisy multimodal contents. EULs provide feature regularization in a variational way, successfully alleviating the effects of data uncertainty within modalities. We design a new variational attention fusion (VAF) module to adaptively fuse multimodal contents with modality-dependent weights. The VAF module consider the relative confidence between modalities and enables to explore complementary properties for detection. Extensive experiments on two benchmark datasets demonstrate the effectiveness and superiority of MM-ULN on multimodal fake news detection. Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Songlin Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Uncertainty-aware Propagation Structure Reconstruction for Fake News DetectionabstractThe widespread of fake news has detrimental societal effects. Recent works model information propagation as graph structure and aggregate structural features from user interactions for fake news detection. However, they usually neglect a broader propagation uncertainty issue, caused by some missing and unreliable interactions during actual spreading, and suffer from learning accurate and diverse structural properties. In this paper, we propose a novel dual graph-based model, Uncertainty-aware Propagation Structure Reconstruction (UPSR) for improving fake news detection. Specifically, after the original propagation modeling, we introduce propagation structure reconstruction to fully explore latent interactions in the actual propagation. We design a novel Gaussian Propagation Estimation to refine the original deterministic node representation by multiple Gaussian distributions and arise latent interactions with KL divergence between distributions in a multi-facet manner. Extensive experiments on two real-world datasets demonstrate the effectiveness and superiority of our model. Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Songlin Hu 0001 |
COLING | 3 |
| 2022 | A Unified Propagation Forest-based Framework for Fake News DetectionabstractFake news’s quick propagation on social media brings severe social ramifications and economic damage. Previous fake news detection usually learn semantic and structural patterns within a single target propagation tree. However, they are usually limited in narrow signals since they do not consider latent information cross other propagation trees. Motivated by a common phenomenon that most fake news is published around a specific hot event/topic, this paper develops a new concept of propagation forest to naturally combine propagation trees in a semantic-aware clustering. We propose a novel Unified Propagation Forest-based framework (UniPF) to fully explore latent correlations between propagation trees to improve fake news detection. Besides, we design a root-induced training strategy, which encourages representations of propagation trees to be closer to their prototypical root nodes. Extensive experiments on four benchmarks consistently suggest the effectiveness and scalability of UniPF. Lingwei Wei, Dou Hu 0001, Yantong Lai, Wei Zhou 0019, Songlin Hu 0001 |
COLING | 4 |
| 2022 | Cascade-Enhanced Graph Convolutional Network for Information Diffusion Prediction
Lingwei Wei, Chunyuan Yuan, Yinan Bao, Wei Zhou 0019, Xian Zhu, Songlin Hu 0001 |
DASFAA (1) | 5 |
| 2022 | MRCE: A Multi-Representation Collaborative Enhancement Model for Aspect-Opinion Pair Extraction
Dongjun Wei, Wei Zhou 0019, Songlin Hu 0001 |
ICONIP (7) | 5 |
| 2022 | Speaker-Guided Encoder-Decoder Framework for Emotion Recognition in ConversationabstractThe emotion recognition in conversation (ERC) task aims to predict the emotion label of an utterance in a conversation. Since the dependencies between speakers are complex and dynamic, which consist of intra- and inter-speaker dependencies, the modeling of speaker-specific information is a vital role in ERC. Although existing researchers have proposed various methods of speaker interaction modeling, they cannot explore dynamic intra- and inter-speaker dependencies jointly, leading to the insufficient comprehension of context and further hindering emotion prediction. To this end, we design a novel speaker modeling scheme that explores intra- and inter-speaker dependencies jointly in a dynamic manner. Besides, we propose a Speaker-Guided Encoder-Decoder (SGED) framework for ERC, which fully exploits speaker information for the decoding of emotion. We use different existing methods as the conversational context encoder of our framework, showing the high scalability and flexibility of the proposed framework. Experimental results demonstrate the superiority and effectiveness of SGED. Yinan Bao, Qianwen Ma, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
IJCAI | 4 |
| 2022 | A Heterogeneous Propagation Graph Model for Rumor Detection Under the Relationship Among Multiple Propagation Subtrees
Guoyi Li, Yulei Wu, Xiaodan Zhang 0004, Wei Zhou 0019, Honglei Lyu |
ECML/PKDD (2) | 5 |
| 2022 | Deception Detection Towards Multi-turn Question Answering with Context Selector Network
Yinan Bao, Qianwen Ma, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001 |
PRICAI (1) | 5 |
| 2021 | Label-Specific Dual Graph Neural Network for Multi-Label Text ClassificationabstractQianwen Ma, Chunyuan Yuan, Wei Zhou, Songlin Hu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Songlin Hu 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Towards Propagation Uncertainty: Edge-enhanced Bayesian Graph Convolutional Networks for Rumor DetectionabstractLingwei Wei, Dou Hu, Wei Zhou, Zhaojuan Yue, Songlin Hu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Zhaojuan Yue, Songlin Hu 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Entity and Relation Matching Consensus for Entity AlignmentabstractEntity alignment aims to match synonymous entities across different knowledge graphs, which is a fundamental task for knowledge integration. Recently, researchers have devoted to leveraging rich information within relations to enhance entity alignment. They explicitly incorporate relations in entity representation and alignment, demonstrating remarkable results. However, affected by the semantic assumptions from early works, these works represent a relation by combining all the entities it connects, ignoring the semantic independence between entity and relation. Moreover, since these works perform alignment by comparing embedding similarity, they fail to consider a graph level alignment and tend to find local false correspondences. Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Xin Wang 0086, Jizhong Han, Songlin Hu 0001 |
CIKM | 3 |
| 2021 | Topic Sequence Embedding for User Identity Linkage from Heterogeneous Behavior DataabstractIn social media, user identity linkage is a vital information security issue of identifying users’ private information across multiple online social networks. With the popularity of behavior-rich social services, existing methods attempt to align users through encoding behaviors. However, most of the efforts suffer from the high variety and heterogeneity of behavior data across social networks, resulting in a limitation of modeling user intrinsic characteristics. To address the above issues, we focus on keyword-based topics to formulate user’s variety behaviors for user identity linkage. In this paper, a novel Topic Sequence Embedding (TSeqE) method is proposed to embed contextual information of topics to represent users’ intrinsic characteristics for identity linkage. Furthermore, we introduce a domain-adversarial training strategy to tackle the behavior heterogeneity problem. Our experiments on two real-world datasets demonstrate that TSeqE produces a significant improvement compared with several strong baselines. Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Jizhong Han, Songlin Hu 0001 |
ICASSP | 2 |
| 2021 | SRLF: A Stance-aware Reinforcement Learning Framework for Content-based Rumor Detection on Social MediaabstractThe rapid development of social media changes the lifestyle of people and simultaneously provides an ideal place for publishing and disseminating rumors, which severely exacerbates social panic and triggers a crisis of social trust. Early content-based methods focused on finding clues from the text and user profiles for rumor detection. Recent studies combine the stances of users' comments with news content to capture the difference between true and false rumors. Although the user's stance is effective for rumor detection, the manual labeling process is time-consuming and labor-intensive, which limits the application of utilizing it to facilitate rumor detection. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework. Chunyuan Yuan, Wanhui Qian, Qianwen Ma, Wei Zhou 0019, Songlin Hu 0001 |
IJCNN | 4 |
| 2021 | SRLF: A Stance-aware Reinforcement Learning Framework for Content-based Rumor Detection on Social MediaabstractThe rapid development of social media changes the lifestyle of people and simultaneously provides an ideal place for publishing and disseminating rumors, which severely exacerbates social panic and triggers a crisis of social trust. Early content-based methods focused on finding clues from the text and user profiles for rumor detection. Recent studies combine the stances of users' comments with news content to capture the difference between true and false rumors. Although the user's stance is effective for rumor detection, the manual labeling process is time-consuming and labor-intensive, which limits the application of utilizing it to facilitate rumor detection. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework. Chunyuan Yuan, Wanhui Qian, Qianwen Ma, Wei Zhou 0019, Songlin Hu 0001 |
IJCNN | 4 |
| 2021 | PEN4Rec: Preference Evolution Networks for Session-Based Recommendation
Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Xiaoyong Huai, Zhiqi Fang, Songlin Hu 0001 |
KSEM | 3 |
| 2020 | Who Are Controlled by The Same User? Multiple Identities Deception Detection via Social Interaction Activity (Student Abstract)abstractSocial media has become a preferential place for sharing information. However, some users may create multiple accounts and manipulate them to deceive legitimate users. Most previous studies utilize verbal or behavior features based methods to solve this problem, but they are only designed for some particular platforms, leading to low universalness.In this paper, to support multiple platforms, we construct interaction tree for each account based on their social interactions which is common characteristic of social platforms. Then we propose a new method to calculate the social interaction entropy of each account and detect the accounts which are controlled by the same user. Experimental results on two real-world datasets show that the method has robust superiority over state-of-the-art methods. Chunyuan Yuan, Wei Zhou 0019, Jingli Wang, Songlin Hu 0001 |
AAAI | 3 |
| 2020 | Early Detection of Fake News by Utilizing the Credibility of News, Publishers, and Users based on Weakly Supervised LearningabstractThe dissemination of fake news significantly affects personal reputation and public trust.Recently, fake news detection has attracted tremendous attention, and previous studies mainly focused on finding clues from news content or diffusion path.However, the required features of previous models are often unavailable or insufficient in early detection scenarios, resulting in poor performance.Thus, early fake news detection remains a tough challenge.Intuitively, the news from trusted and authoritative sources or shared by many users with a good reputation is more reliable than other news.Using the credibility of publishers and users as prior weakly supervised information, we can quickly locate fake news in massive news and detect them in the early stages of dissemination.In this paper, we propose a novel Structure-aware Multi-head Attention Network (SMAN), which combines the news content, publishing, and reposting relations of publishers and users, to jointly optimize the fake news detection and credibility prediction tasks.In this way, we can explicitly exploit the credibility of publishers and users for early fake news detection.We conducted experiments on three real-world datasets, and the results show that SMAN can detect fake news in 4 hours with an accuracy of over 91%, which is much faster than the state-of-the-art models. Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
COLING | 3 |
| 2020 | RE-GCN: Relation Enhanced Graph Convolutional Network for Entity Alignment in Heterogeneous Knowledge Graphs
Jinzhu Yang, Wei Zhou 0019, Lingwei Wei, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001 |
DASFAA (2) | 2 |
| 2020 | Exploiting Heterogeneous Artist and Listener Preference Graph for Music Genre ClassificationabstractMusic genres are useful for indexing, organizing, searching, and recommending songs and albums. Therefore, the automatic classification of music genres is an essential part of almost all kinds of music applications. Recent works focus on exploiting text, audio, or multi-modal information for genre classification, without considering the influence of the artists' and listeners' preference. However, intuitively, artists have their composing preferences, and listeners also have their music tastes. Both of them provide helpful hints to the music genre from different views, which are crucial to improve classification performance. Chunyuan Yuan, Qianwen Ma, Junyang Chen 0001, Wei Zhou 0019, Xiaodan Zhang 0004, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 4 |
| 2020 | AutoSUM: Automating Feature Extraction and Multi-user Preference Simulation for Entity Summarization
Dongjun Wei, Fuqing Zhu, Liangjun Zang, Wei Zhou 0019, Songlin Hu 0001 |
PAKDD (2) | 5 |
| 2020 | Hierarchical Interaction Networks with Rethinking Mechanism for Document-Level Sentiment Analysis
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xuehai Tang, Xiaodan Zhang 0004, Xin Wang 0086, Jizhong Han, Songlin Hu 0001 |
ECML/PKDD (3) | 3 |
| 2020 | DyHGCN: A Dynamic Heterogeneous Graph Convolutional Network to Learn Users' Dynamic Preferences for Information Diffusion Prediction
Chunyuan Yuan, Wei Zhou 0019, Xiaodan Zhang 0004, Songlin Hu 0001 |
ECML/PKDD (3) | 3 |
| 2020 | Beyond Statistical Relations: Integrating Knowledge Relations into Style Correlations for Multi-Label Music Style ClassificationabstractAutomatically labeling multiple styles for every song is a comprehensive application in all kinds of music websites. Recently, some researches explore review-driven multi-label music style classification and exploit style correlations for this task. However, their methods focus on mining the statistical relations between different music styles and only consider shallow style relations. Moreover, these statistical relations suffer from the underfitting problem because some music styles have little training data. To tackle these problems, we propose a novel knowledge relations integrated framework (KRF) to capture the complete style correlations, which jointly exploits the inherent relations between music styles according to external knowledge and their statistical relations. Based on the two types of relations, we use graph convolutional network to learn the deep correlations between styles automatically. Experimental results show that our framework significantly outperforms the state-of-the-art methods. Further studies demonstrate that our framework can effectively alleviate the underfitting problem and learn meaningful style correlations. Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
WSDM | 3 |
| 2020 | Yet another approach to understanding news event evolution
Shangwen Lv, Longtao Huang, Liangjun Zang, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
World Wide Web | 4 |
| 2019 | A Time-Series Sockpuppet Detection Method for Dynamic Social Relationships
Wei Zhou 0019, Jingli Wang, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001 |
DASFAA (1) | 1 |
| 2019 | Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based ChatbotsabstractChunyuan Yuan, Wei Zhou, Mingming Li, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunyuan Yuan, Wei Zhou 0019, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Jointly Embedding the Local and Global Relations of Heterogeneous Graph for Rumor DetectionabstractThe development of social media has revolutionized the way people communicate, share information and make decisions, but it also provides an ideal platform for publishing and spreading rumors. Existing rumor detection methods focus on finding clues from text content, user profiles, and propagation patterns. However, the local semantic relation and global structural information in the message propagation graph have not been well utilized by previous works. In this paper, we present a novel global-local attention network (GLAN) for rumor detection, which jointly encodes the local semantic and global structural information. We first generate a better integrated representation for each source tweet by fusing the semantic information of related retweets with the attention mechanism. Then, we model the global relationships among all source tweets, retweets, and users as a heterogeneous graph to capture the rich structural information for rumor detection. We conduct experiments on three real-world datasets, and the results demonstrate that GLAN significantly outperforms the state-of-the-art models in both rumor detection and early detection scenarios. Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
ICDM | 3 |
| 2019 | Learning Review Representations from user and Product Level Information for Spam DetectionabstractOpinion spam has become a widespread problem in social media, where hired spammers write deceptive reviews to promote or demote products to mislead the consumers for profit or fame. Existing works mainly focus on manually designing discrete textual or behavior features, which cannot capture complex global semantics of reviews. Although recent works apply deep learning methods to learn review-level semantic features, their models ignore the impact of the user-level and product-level information on learning review semantics and the inherent user-review-product relationship information. In this paper, we propose a Hierarchical Fusion Attention Network (HFAN) to automatically learn the semantics of reviews from user and product level. Specifically, we design a multiattention unit to extract user(product)-related review information. Then, we use orthogonal decomposition and fusion attention to learn a user, review, and product representation from the review information. Finally, we take the review as a relation between user and product entity and apply TransH to jointly encode this relationship into review representation. Experimental results obtained more than 10% absolute precision improvement over the state-of-the-art performances on four real-world datasets, which show the effectiveness and versatility of the model. Chunyuan Yuan, Wei Zhou 0019, Qianwen Ma, Shangwen Lv, Jizhong Han, Songlin Hu 0001 |
ICDM | 2 |
| 2019 | Fusion Convolutional Attention Network for Opinion Spam Detection
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
ICONIP (1) | 4 |
| 2016 | An efficient graph data processing system for large-scale social network service applicationsabstractSummary Trust in social network draws more and more attentions from both the academia and industry fields. Public opinion analysis is a direct way to increase the trust in social network. Because the public opinion analysis can be expressed naturally by the graph algorithm and graph data are the default data organization mechanism used in large‐scale social network service applications, more and more research works apply the graph processing system to deal with the public opinion analysis. As the data volume is growing rapidly, the distributed graph systems are introduced to process the large‐scale public opinion analysis. Most of graph algorithms introduce a large number of data iterations, so the synchronization requirements between successive iterations can severely jeopardize the effectiveness of parallel operations, which makes the data aggregation and analysis operations become slower. In this paper, we propose a large‐scale graph data processing system to address these issues, which includes a graph data processing model, Arbor. Arbor develops a new graph data organization format to represent the social relationship, and the format can not only save storage space but also accelerate graph data processing operations. Furthermore, Arbor substitutes time‐constrained synchronization operations with non‐time‐constrained control message transmissions to increase the degree of parallelism. Based on the system, we put forward two most frequently used graph applications on Arbor: shortest path and PageRank. In order to evaluate the system, we compare Arbor with the other graph processing systems using large‐scale experimental graph data, and the results show that it outperforms the state‐of‐the‐art systems. Copyright © 2014 John Wiley & Sons, Ltd. Wei Zhou 0019, Jizhong Han, Zhiyong Xu 0003 |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | PADM: Page Rank-Based Anomaly Detection Method of Log Sequences by Graph ComputingabstractWith the popularity of various software applications in cloud computing, software exception becomes an important issue. How to detect the exceptions more quickly seems to be crucial for the software service company. To solve the above problem, this paper presents an efficient log anomaly detection method named PADM (Page Rank-based Anomaly Detection Method) based on the graph computing algorithm. In this method, the logs are transformed into a graph to represent the complex relationship between the log records, then we design an extended Page Rank algorithm based on the graph to get the importance score for each log. After that, we compare the scores to that of the training logs to determine whether they are abnormal or not. Finally, we compare PADM with other anomaly detection methods on the real logs, and the results show that it outperforms the currently widely used mechanisms with higher accuracy, lower time complexity and better scalability. Xiaoben Yan, Wei Zhou 0019, Jizhong Han, Ge Fu |
CloudCom | 2 |
| 2014 | Marbor: A novel large-scale graph data storage and processing frameworkabstractIn this paper, we propose Marbor, a novel graph data processing framework to analyze the large-scale data in social network services. It develops an efficient graph organization model to minimize the costs of graph data accesses and reduce the memory consumption. In addition, we present a novel control message method in Marbor to improve the synchronization iterations performance. During the graph data processing, in each iteration, it analyzes the relationships among tasks and forwards the tasks to the next iteration with control messages, so no synchronization operations are used. We compare Marbor with other graph processing methods on several large-scale real world SNS datasets with two widely used applications, and the results show that Marbor outperforms the current mechanisms. Wei Zhou 0019, Jizhong Han, Zhiyong Xu 0003 |
IPCCC | 1 |
| 2014 | Online Anomaly Detection by Improved Grammar Compression of Log SequencesabstractNowadays, log sequences mining techniques are widely used in detecting anomalies for Internet services. The state-of-the-art anomaly detection methods either need significant computational costs, or require specific assumptions that the test logs are holding certain data distribution patterns in order to be effective. Therefore, it is very difficult to achieve real time responses and it greatly reduces the effectiveness of these mechanisms in reality. To address these issues, we propose an innovative anomaly detection strategy called CADM. In CADM, the relative entropy between test logs and normal logs is exploited to discover the anomalous levels. Instead of calculating the relative entropy based on certain predefined data distribution models, our solution inspects the relationship between relative entropy and compression size with an improved grammar-based compression method. No assumptions are needed. In addition, our mechanism has excellent scalability with only O(n) computational complexity. It can generate the detection results on the fly. Experimental analysis with both synthetic and real world logs proves that CADM is superior to the other methods. It can achieve very high anomaly detection accuracy with the minimal computational overhead. It is suitable for log mining tasks and can be applied on a broad variety of application fields. Wei Zhou 0019, Jizhong Han, Dan Meng 0002, Zhiyong Xu 0003 |
SDM | 2 |
| 2013 | HDKV: supporting efficient high-dimensional similarity search in key-value storesabstractSUMMARY Key‐value stores are widely used on large‐scale data management in the cloud environment. However, they can only naturally support key‐based queries, and do not have efficient solutions for value‐based queries. Thus, dealing with high‐dimensional data in key‐value stores is still a big challenge. State‐of‐the‐art solutions apply value‐based tree‐structure indexes to solve this issue. These methods suffer from the curse of dimensionality and cannot achieve satisfactory performance. They also bring serious load unbalancing problem among servers, and result in dramatic system scalability degradation. Meanwhile, similarity search in high‐dimensional data space becomes more and more popular in today's cloud applications. Due to the lack of efficient algorithms for value‐based queries, users have to wait for a long time before the results are returned. To address this issue, we propose a novel approach called high‐dimensional similarity query in key‐value stores (HDKV), which can generate similarity results in a short time and maintain good database scalability. In HDKV, a strict order‐preserving hash function is designed to map nearby objects in the high‐dimensional space onto adjacent keys of a continuous linear space in key‐value stores. With this strategy, many expensive random accesses are replaced with more efficient scan accesses. The experimental evaluation on real world data set shows that compared to the state‐of‐the‐art methods, HDKV can dramatically reduce the search time with little impact on the accuracy. Copyright © 2012 John Wiley & Sons, Ltd. Wei Zhou 0019, Jizhong Han, Jiao Dai, Zhiyong Xu 0003 |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | LogMaster: Mining Event Correlations in Logs of Large-Scale Cluster SystemsabstractThis paper presents a set of innovative algorithms and a system, named Log Master, for mining correlations of events that have multiple attributions, i.e., node ID, application ID, event type, and event severity, in logs of large-scale cloud and HPC systems. Different from traditional transactional data, e.g., supermarket purchases, system logs have their unique characteristics, and hence we propose several innovative approaches to mining their correlations. We parse logs into an n-ary sequence where each event is identified by an informative nine-tuple. We propose a set of enhanced apriori-like algorithms for improving sequence mining efficiency, we propose an innovative abstraction-event correlation graphs (ECGs) to represent event correlations, and present an ECGs-based algorithm for fast predicting events. The experimental results on three logs of production cloud and HPC systems, varying from 433490 entries to 4747963 entries, show that our method can predict failures with a high precision and an acceptable recall rates. Xiaoyu Fu, Jianfeng Zhan, Wei Zhou 0019, Zhen Jia 0001 |
SRDS | 4 |
| 2010 | Accelerating Spatial Data Processing with MapReduceabstractMap Reduce is a key-value based programming model and an associated implementation for processing large data sets. It has been adopted in various scenarios and seems promising. However, when spatial computation is expressed straightforward by this key-value based model, difficulties arise due to unfit features and performance degradation. In this paper, we present methods as follows: 1) a splitting method for balancing workload, 2) pending file structure and redundant data partition dealing with relation between spatial objects, 3) a strip-based two-direction plane sweeping algorithm for computation accelerating. Based on these methods, ANN(All nearest neighbors) query and astronomical cross-certification are developed. Performance evaluation shows that the Map Reduce-based spatial applications outperform the traditional one on DBMS. Jizhong Han, Bibo Tu, Jiao Dai, Wei Zhou 0019 |
ICPADS | 5 |
| 2010 | Online Event Correlations Analysis in System Logs of Large-Scale Cluster Systems
Wei Zhou 0019, Jianfeng Zhan, Dan Meng 0002 |
NPC | 1 |
| 2007 | A layered design methodology of cluster system stackabstractThe application range of cluster has expanded beyond scientific computing, but the present cluster system software fails to provide a flexible architecture to promote code reuse and facilitate building cluster system software for different computing contexts, most of which are developed from scratch case by case, or integrated or packaged with “the best practice”. In this paper, we have proposed a layered design methodology to build cluster system stack with different layers concentrating on different functions, and developed common sets of core service as reusing framework for different computing context. Following this methodology, we have built Phoenix-a complete cluster system stack for both scientific and business computing, which is verified and deployed on Dawning 4000A super computer for scientific computing and other cluster systems for business computing. The qualitative evaluation and our practices show the design methodology of Phoenix has advantages over other methodologies. Jianfeng Zhan, Lei Wang 0004, Bibo Tu, Yu Wen 0001, Yuansheng Chen, Wei Zhou 0019, Dan Meng 0002, Ninghui Sun |
CLUSTER | 7 |