Wenya Wang 0001

dblp:62/7948-1 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0001-5612-7818ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 11 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Skill Path: Unveiling Language Skills from Circuit Graphs
abstract
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffer from atomic ablation, which causes the loss of causal dependencies between connected components. In addition, their discovery process, designed to preserve output faithfulness, inadvertently captures extraneous effects other than an isolated target skill. To alleviate these challenges, we introduce skill paths, which offer a more refined and compact representation by isolating individual skills within a linear chain of components. To enable skill path extracting from circuit graphs, we propose a three-step framework, consisting of decomposition, pruning, and post-hoc causal mediation. In particular, we offer a complete linear decomposition of the transformer model which leads to a disentangled computation graph. After pruning, we further adopt causal analysis techniques, including counterfactuals and interventions, to extract the final skill paths from the circuit graph. To underscore the significance of skill paths, we investigate three generic language skills—Previous Token Skill, Induction Skill, and In-Context Learning Skill—using our framework. Experiments support two crucial properties of these skills, namely stratification and inclusiveness.
Hang Chen 0002, Xinyu Yang 0001, Jiaying Zhu, Wenya Wang 0001
AAAI4
2026 LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit
abstract
Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM.
Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu 0024, Yumeng Shi, Wenya Wang 0001
AAAI10
2026 Causality Matters: How Temporal Information Emerges in Video Language Models
abstract
Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and relationships across time, still remains a core challenge. Prior works emphasize positional encodings (PEs) as a key mechanism for encoding temporal structure. Surprisingly, we find that removing or modifying PEs in video inputs yields minimal degradation in the performance of temporal understanding. In contrast, reversing the frame sequence while preserving the original PEs causes a substantial drop. To explain this behavior, we conduct substantial analysis experiments to trace how temporal information is integrated within the model. We uncover a causal information pathway: temporal cues are progressively synthesized through inter-frame attention, aggregated in the final frame, and subsequently integrated into the query tokens. This emergent mechanism shows that temporal reasoning emerges from inter-visual token interactions under the constraints of causal attention, which implicitly encodes temporal structure. Based on these insights, we propose two efficiency-oriented strategies: staged cross-modal attention and a temporal exit mechanism for early token truncation. Experiments on two benchmarks validate the effectiveness of both approaches.
Yumeng Shi, Quanyu Long, Yin Wu 0001, Wenya Wang 0001
AAAI4
2026 Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding
abstract
Jianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu, Sarthak Ketanbhai Modi, Zi Pong Lim, Yon Shin Teo, Wenya Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu, Sarthak Ketanbhai Modi, Zi Pong Lim, Yon Shin Teo, Wenya Wang 0001
ACL (1)8
2026 Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
abstract
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by hallucinations-generating text that contradicts visual input.Recent studies often attribute these errors to inadequate visual attention.In this work, we analyze the attention mechanisms via the logit lens, uncovering a distinct anomaly we term Vocabulary Hijacking.We discover that specific visual tokens, defined as Inert Tokens, disproportionately attract attention.Crucially, when their intermediate hidden states are projected into the vocabulary space, they consistently decode to a fixed set of unrelated words (termed Hijacking Anchors) across layers, revealing a rigid semantic collapse.Leveraging this semantic rigidity, we propose Hijacking Anchor-Based Identification (HABI), a robust strategy to accurately localize these Inert Tokens.To quantify the impact of this phenomenon, we introduce the Non-Hijacked Visual Attention Ratio (NHAR), a novel metric designed to identify attention heads that remain resilient to hijacking and are critical for factual accuracy.Building on these insights, we propose Hijacking-Aware Visual Attention Enhancement (HAVAE), a trainingfree intervention that selectively strengthens the focus of these identified heads on salient visual content.Extensive experiments across multiple benchmarks demonstrate that HAVAE significantly mitigates hallucinations with no additional computational overhead, while preserving the model's general capabilities.
Yangneng Chen, Weijun Yao, Xilai Ma, Guodong Du 0008, Wenya Wang 0001
ACL (1)6
2026 Programming over Thinking: Efficient and Robust Multi-Constraint Planning
abstract
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints.Existing large language model (LLM) approaches face fundamental limitations in this domain.Pure reasoning paradigms, which rely on long natural language chains, are prone to inconsistency, error accumulation, and prohibitive cost as constraints compound.Conversely, LLMs combined with coding-or solver-based strategies lack flexibility: they often generate problem-specific code from scratch or depend on fixed solvers, failing to capture generalizable logic across diverse problems.To address these challenges, we introduce the Scalable COde Planning Engine (SCOPE), a framework that disentangles query-specific reasoning from generic code execution.By separating reasoning from execution, SCOPE produces solver functions that are consistent, deterministic, and reusable across queries while requiring only minimal changes to input parameters.SCOPE achieves state-of-the-art performance while lowering cost and latency.For example, with GPT-4o, it reaches 93.1% success on TravelPlanner, a 61.6% gain over the best baseline (CoT) while cutting inference cost by 1.4x and time by 4.67x.Code is available at https://github.com/DerrickGXD/SCOPE.
Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang 0001
ACL (1)5
2026 Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification
abstract
Multi-hop claim verification is inherently challenging, requiring multi-step reasoning to construct verification chains while iteratively searching for information to uncover hidden bridging facts. This process is fundamentally interleaved, as effective reasoning relies on dynamically retrieved evidence, while effective search demands reasoning to refine queries based on partial information. To achieve this, we propose Hierarchical Agent Reasoning and Information Search (HARIS), explicitly modeling the coordinated process of reasoning-driven searching and search-informed reasoning. HARIS consists of a high-level reasoning agent that focuses on constructing the main verification chain, generating factual questions when more information is needed, and a low-level search agent that iteratively retrieves more information, refining its search based on intermediate findings. This design allows each agent to specialize in its respective task, enhancing verification accuracy and interpretability. HARIS is trained using reinforcement learning with outcome-based rewards. Experimental results on the EX-FEVER and HOVER benchmarks demonstrate that HARIS achieves strong performance, greatly advancing multi-hop claim verification.
Qisheng Hu, Quanyu Long, Wenya Wang 0001
ACL (1)3
2026 From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
abstract
Ziwei Huang, Ying Shu, Fanghao, Quanyu Long, Wenya Wang, Qiushi Guo, Tiezheng Ge, Leilei Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziwei Huang 0005, Yin Shu, Quanyu Long, Wenya Wang 0001, Qiushi Guo, Tiezheng Ge, Leilei Gan
ACL (1)5
2026 Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios
abstract
Large language models (LLMs) exhibit substantial variability in performance and computational cost across tasks and queries, motivating routing systems that select models to meet user-specific cost--performance trade-offs. However, existing routers generalize poorly in cold-start scenarios where in-domain training data is unavailable. We address this limitation with a multi-level task-profile--guided data synthesis framework that constructs a hierarchical task taxonomy and produces diverse question--answer pairs to approximate the test-time query distribution. Building on this, we introduce TRouter, a task-type--aware router approach that models query-conditioned cost and performance via latent task-type variables, with prior regularization derived from the synthesized task taxonomy. This design enhances TRouter's routing utility under both cold-start and in-domain settings. Across multiple benchmarks, we show that our synthesis framework alleviates cold-start issues and that TRouter delivers effective LLM routing. ©2026 Association for Computational Linguistics
Hui Liu 0036, Kecheng Chen, Jie Liu 0044, Wenya Wang 0001, Haoliang Li
ACL (1)5
2026 Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
abstract
Retrieval-augmented generation (RAG) augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate multimodal LLMs (MLLMs) under multimodal RAG settings, they predominantly retrieve from textual corpora and do not explicitly assess how models exploit visual evidence during answer generation. Consequently, there still lacks benchmark that cleanly isolates and measures the contribution of retrieved images in a visual knowledge-intensive RAG pipeline. We introduce Visual-RAG, a question-answering benchmark that targets visually-grounded, knowledge-intensive questions in a visual evidence-centric manner. Unlike prior work, Visual-RAG requires text-to-image retrieval and the integration of retrieved clue images whose pixel content explicitly encodes the visual knowledge necessary for answer generation. With Visual-RAG, we evaluate five open-source and three proprietary MLLMs and find that current systems still substantially underutilize the visual information available in retrieved images. Despite clear opportunities for multimodal evidence integration, state-of-the-art models struggle to extract and exploit fine-grained visual knowledge, and text-to-image retrieval itself remains challenging even under constrained entity-level corpora. These results underscore the need for improved visual retrieval, grounding, and attribution in multimodal RAG. Visual-RAG is publicly available at: github.com/visual-rag/visual-rag
Yin Wu 0001, Quanyu Long, Jing Li 0034, Jianfei Yu, Wenya Wang 0001
SIGIR5
2026 Modeling Deep Fusion of Intra- and Inter-Modal Incongruity for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection receives increasing attentions due to people's growing interest in posting multimodal information. The key factor of multimodal sarcasm detection is to leverage incongruity information across different modalities. Existing works are mainly based on the late fusion strategy by simply concatenating the intra- and inter-modal incongruity features, which are prone to learning surface patterns. In contrast, this work mainly focuses on modeling the deep fusion of intra- and inter-modal incongruity information. To this end, this work first discusses the incompatibility between the two kinds of incongruity features within existing multimodal frameworks. Under this motivation, we further propose an end-to-end cooperative framework dubbed Cooperative Multimodal Incongruity Learning (CoMIL). Specifically, our approach incorporates a primary module to model the deep fusion of intra- and inter modal incongruity information. To prevent the integrated inter modal visual information from disturbing the modeling of intra text incongruity, CoMIL introduces a cooperative mechanism incorporating a reference module which focuses on token-level correlations as a structural guidance to the primary module. Based on the proposed cooperative mechanism, the intra- and inter-modal incongruity information can be compactly and compatibly integrated into deep features of neural models. Extensive experiments are conducted to validate the effectiveness of our proposed CoMIL approach.
Fengmao Lv, Junlin Fang, Guosheng Lin, Wenya Wang 0001
IEEE Trans. Knowl. Data Eng.4
2025 Re2LLM: Reflective Reinforcement Large Language Model for Session-based Recommendation
abstract
Emerging advancements in large language models (LLMs) show significant potential for enhancing recommendations. However, prompt-based methods often struggle to find ideal prompts without task-specific feedback, while fine-tuning-based methods are hindered by high computational demands and dependence on open-source backbones. To address these challenges, we propose a Reflective Reinforcement Large Language Model (Re2LLM) for session-based recommendation, which refines LLMs to generate and utilize specialized knowledge effectively and efficiently. Specifically, we first devise the Reflective Exploration Module to extract and present knowledge in a form that LLMs can easily process. This module enables LLMs to reflect on their recommendation mistakes and construct a hint knowledge base to rectify them effectively. Next, we design the Reinforcement Utilization Module to train a lightweight retrieval agent that elicits correct LLM reasoning. This module recognizes hints as signals to facilitate LLM recommendations and learns to select appropriate hints from the constructed knowledge base using task-specific feedback efficiently. Lastly, we conduct experiments on real-world datasets and demonstrate the superiority of our Re2LLM over state-of-the-art methods.
Yingpeng Du, Zhu Sun 0001, Haoyan Chua, Kaidong Feng, Wenya Wang 0001, Jie Zhang 0002
AAAI6
2025 Quantifying Semantic Emergence in Language Models
abstract
Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning.Yet, there remains no established metric to quantify this capability.In this work, we introduce a quantitative metric, Information Emergence (IE), designed to measure LLMs' ability to extract semantics from input tokens.We formalize "semantics" as the meaningful information abstracted from a sequence of tokens and quantify this by comparing the entropy reduction observed for a sequence of tokens (macro-level) and individual tokens (micro-level).To achieve this, we design a lightweight estimator to compute the mutual information at each transformer layer, which is agnostic to different tasks and language model architectures.We apply IE in both synthetic in-context learning (ICL) scenarios and natural sentence contexts.Experiments demonstrate informativeness and patterns about semantics.While some of these patterns confirm the conventional prior linguistic knowledge, the rest are relatively unexpected, which may provide new insights.
Hang Chen 0002, Xinyu Yang 0001, Jiaying Zhu, Wenya Wang 0001
ACL (1)4
2025 MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
abstract
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures.However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses.In this paper, we propose the Multi-Turn Safety Alignment (MTSA) framework, to address the challenge of securing LLMs in multi-round interactions.It consists of two stages: In the thought-guided attack learning stage, the redteam model learns about thought-guided multiround jailbreak attacks to generate adversarial prompts.In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction.Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment.Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks.
Weiyang Guo, Jing Li 0034, Wenya Wang 0001, Yu Li 0007, Daojing He, Jun Yu 0002, Min Zhang 0005
ACL (1)3
2025 Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
abstract
Junlin Li, Guodong Du, Jing Li, Sim Kuan Goh, Wenya Wang, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guodong Du 0002, Jing Li 0034, Sim Kuan Goh, Wenya Wang 0001, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang 0005
ACL (1)5
2025 Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning
abstract
Hui Liu, Wenya Wang, Hao Sun, Chris Xing Tian, Chenqi Kong, Xin Dong, Haoliang Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hui Liu 0036, Wenya Wang 0001, Chris Xing Tian, Chenqi Kong, Haoliang Li
ACL (1)2
2025 Exploring Quality and Diversity in Synthetic Data Generation for Argument Mining
abstract
The advancement of Argument Mining (AM) is hindered by a critical bottleneck: the scarcity of structure-annotated datasets, which are expensive to create manually.Inspired by recent successes in synthetic data generation across various NLP tasks, this paper explores methodologies for LLMs to generate synthetic data for AM.We investigate two complementary synthesis perspectives: a quality-oriented synthesis approach, which employs structure-aware paraphrasing to preserve annotation quality, and a diversity-oriented synthesis approach, which generates novel argumentative texts with diverse topics and argument structures.Experiments on three datasets show that augmenting original training data with our synthetic data, particularly when combining both quality-and diversity-oriented instances, significantly enhances the performance of existing AM models, both in full-data and low-resource settings.Moreover, the positive correlation between synthetic data volume and model performance highlights the scalability of our methods.
Jianzhu Bao, Wenya Wang 0001, Yice Zhang, Bojun Jin, Ruifeng Xu 0001
EMNLP4
2025 STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment
abstract
In-Context Learning (ICL) has become a powerful paradigm that enables LLMs to perform a wide range of tasks without task-specific finetuning.However, the effectiveness of ICL heavily depends on the quality of exemplar selection.In particular, for structured prediction tasks such as semantic parsing, existing ICL selection strategies often overlook structural alignment, leading to suboptimal performance and poor generalization.To address this issue, we propose a novel two-stage exemplar selection strategy that achieves a strong balance between efficiency, generalizability, and performance.First, we fine-tune a BERT-based retriever using structure-aware supervision, guiding it to select exemplars that are both semantically relevant and structurally aligned.Then, we enhance the retriever with a plug-in module, which amplifies syntactically meaningful information in the hidden representations.This plug-in is model-agnostic, requires minimal overhead, and can be seamlessly integrated into existing pipelines.Experiments on four benchmarks spanning three semantic parsing tasks demonstrate that our method consistently outperforms existing baselines with multiple recent LLMs as inference-time models 1 .
Qisheng Hu, Jing Li 0049, Wenya Wang 0001
EMNLP4
2025 Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
abstract
Video question answering benefits from the rich information in videos, enabling various applications.However, the large volume of tokens generated from long videos presents challenges to memory efficiency and model performance.To alleviate this, existing works propose to compress video inputs, but often overlook the varying importance of static and dynamic information across different queries, leading to inefficient token usage within limited budgets.We propose a novel token selection strategy, EXPLORE-THEN-SELECT, that adaptively adjusts static and dynamic information based on question requirements.Our framework first explores different token allocations between key frames, which preserve spatial details, and delta frames, which capture temporal changes.Then it employs a query-aware attention-based metric to select the optimal token combination without model updates.Our framework is plug-and-play and can be seamlessly integrated within diverse video language models.Extensive experiments show that our method achieves significant performance improvements (up to 5.8%) on multiple video question answering benchmarks.Our code is available at https://github.com/ANDgate99/Explore- Then-Select.
Yumeng Shi, Quanyu Long, Wenya Wang 0001
EMNLP3
2025 Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated Data
Junlin Fang, Wenya Wang 0001, Fengmao Lv
ACM Multimedia2
2025 Global Question-Aware Multimodal Retrieval-Augmented Generation for Multimedia Multi-Hop Question Answering
abstract
Multimedia Multi-Hop Question Answering (MMQA) is a complex task that requires models to reason over and integrate information from both visual (e.g., images) and textual (e.g., documents) modalities to answer questions that cannot be resolved in a single step. Existing methods for MMQA suffer from two common limitations: (1) insufficient cross-modal information fusion, which restricts interaction between different modalities; and (2) weak global understanding of multi-hop questions, making them vulnerable to distractions from intermediate steps. To address these challenges, we introduce Global question-aware Multimodal Retrieval-Augmented Generation (GMRAG), a framework designed to enhance cross-modal reasoning and improve retrieval precision for multi-hop multimodal questions. It brings two core innovations: (1) restructuring the training data into a global question–aware, evidence-centric retrieval dataset, enabling the retriever to perform richer cross-modal reasoning on multi-hop multimodal questions and to identify globally relevant information more effectively; and (2) employing cross-modal contrastive learning to fine-tune a joint image–text encoder, achieving stronger alignment between holistic multimodal questions and evidence. The retrieved evidence is then directly fed into a Multimodal Large Language Models (MLLMs) to generate the final answer to the original question, eliminating dependence on potentially flawed intermediate answers and thus mitigating error propagation. Experimental results on two public MMQA datasets show that GMRAG consistently outperforms existing RAG methods across various MLLMs.
Zhixiao Shen, Jianfei Yu, Wenya Wang 0001
MMAsia3
2025 Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
abstract
Qisheng Hu, Quanyu Long, Wenya Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qisheng Hu, Quanyu Long, Wenya Wang 0001
NAACL (Long Papers)3
2025 Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
abstract
Circuit discovery has gradually become one of the prominent methods for mechanistic interpretability, and research on circuit completeness has also garnered increasing attention. Methods of circuit discovery that do not guarantee completeness not only result in circuits that are not fixed across different runs but also cause key mechanisms to be omitted. The nature of incompleteness arises from the presence of OR gates within the circuit, which are often only partially detected in standard circuit discovery methods. To this end, we systematically introduce three types of logic gates: AND, OR, and ADDER gates, and decompose the circuit into combinations of these logical gates. Through the concept of these gates, we derive the minimum requirements necessary to achieve faithfulness and completeness. Furthermore, we propose a framework that combines noising-based and denoising-based interventions, which can be easily integrated into existing circuit discovery methods without significantly increasing computational complexity. This framework is capable of fully identifying the logic gates and distinguishing them within the circuit. In addition to the extensive experimental validation of the framework's ability to restore the faithfulness, completeness, and sparsity of circuits, using this framework, we uncover fundamental properties of the three logic gates, such as their proportions and contributions to the output, and explore how they behave among the functionalities of language models.
Hang Chen 0002, Jiaying Zhu, Xinyu Yang 0001, Wenya Wang 0001
NeurIPS4
2025 Damage Analysis via Bidirectional Multi-Task Cascaded Multimodal Fusion
abstract
Damage analysis in social media platforms such as Twitter is a comprehensive problem which involves different subtasks for mining damage-related information from tweets ( e.g., informativeness, humanitarian categories and severity assessment). The comprehensive information obtained by damage analysis enables to identify breaking events around the world in real-time and hence provides aids in emergency responses. Recently, with the rapid development of web technologies, multimodal damage analysis has received increasing attentions due to users' preference of posting multimodal information in social media. Multimodal damage analysis leverages the associated image modality to improve the identification of damage-related information in social media. However, existing works on multimodal damage analysis address each damage-related subtask individually and do not consider their joint training mechanism. In this work, we propose the Bidirectional Multi-task Cascaded multimodal Fusion (BiMCF) approach towards joint multimodal damage analysis. To this end, we introduce the cascaded multimodal fusion framework to separately integrate effective visual and text information for each task, considering that different tasks attend to different information. To exploit the interactions across tasks, bidirectional propagation of the attended image-text interactive information is implemented between tasks, which can lead to enhanced multimodal fusion. Comprehensive experiments are conducted to validate the effectiveness of the proposed approach. Code is available at https://github.com/tiggers23/BiMCF.
Siying Wu, Junfeng Fang, Guowu Yang, Wenya Wang 0001, Fengmao Lv
WWW5
2025 DT4LM: Differential Testing for Reliable Language Model Updates in Classification Tasks
Xinyue Zuo, Yan Xiao 0002, Xiaochun Cao, Wenya Wang 0001, Jin Song Dong 0001
IEEE Trans. Software Eng.4
2024 Training Language Models to Generate Text with Citations via Fine-grained Rewards
abstract
While recent Large Language Models (LLMs) have proven useful in answering user queries, they are prone to hallucination, and their responses often lack credibility due to missing references to reliable sources.An intuitive solution to these issues would be to include in-text citations referring to external documents as evidence.While previous works have directly prompted LLMs to generate in-text citations, their performances are far from satisfactory, especially when it comes to smaller LLMs.In this work, we propose an effective training framework using fine-grained rewards to teach LLMs to generate highly supportive and relevant citations, while ensuring the correctness of their responses.We also conduct a systematic analysis of applying these fine-grained rewards to common LLM training strategies, demonstrating its advantage over conventional practices.We conduct extensive experiments on Question Answering (QA) datasets taken from the ALCE benchmark and validate the model's generalizability using EXPERTQA.On LLaMA-2-7B, the incorporation of fine-grained rewards achieves the best performance among the baselines, even surpassing that of GPT-3.5-turbo.1
Chengyu Huang 0003, Zeqiu Wu, Yushi Hu, Wenya Wang 0001
ACL (1)4
2024 Progressive Multimodal Pivot Learning: Towards Semantic Discordance Understanding as Humans
abstract
Multimodal recognition can achieve enhanced performance by leveraging the complementary information from different modali- ties. However, in real-world scenarios, multimodal samples often express discordant semantic meanings across modalities, lacking evident complementary information. Unlike humans who can easily understand the intrinsic semantic information of these semantically discordant samples, existing multimodal recognition models show poor performance on them. With the motivation of improving the robustness of multimodal recognition models in practical scenar- ios, this work poses a new challenge in multimodal recognition, which is coined as Semantic Discordance Understanding. Unlike ex- isting works only focusing on detecting semantically discordant samples as noisy data, this new challenge requires deep models to follow humans’ ability in understanding the inherent seman- tic meanings of semantically discordant samples. To address this challenge, we further propose the Progressive Multimodal Pivot Learning (PMPL) approach by introducing a learnable pivot mem- ory to explore the inherent semantics meaning hidden under dis- cordant modalities. To this end, our approach inserts Pivot Memory Learning (PML) modules into multiple layers of unimodal foun- dation models to progressively trade-off the conflict information across modalities. By introducing the multimodal pivot learning paradigm for multimodal recognition, the proposed PMPL approach can alleviate the negative effect of semantic discordance caused by the cross-modal information exchange mechanism of existingmultimodal recognition models. Experiments on different bench- marks validate the superiority of our approach. Code is available at https://github.com/tiggers23/PMPL.
Junlin Fang, Wenya Wang 0001, Tianze Luo, Yanyong Huang, Fengmao Lv
CIKM2
2024 Sentiment-oriented Sarcasm Integration for Video Sentiment Analysis Enhancement with Sarcasm Assistance
abstract
Sarcasm is an intricate expression phenomenon and has garnered increasing attentions over the recent years, especially for multimodal contexts such as videos.Nevertheless, despite being a significant aspect of human sentiment, the effect of sarcasm is consistently overlooked in sentiment analysis.Videos with sarcasm often convey sentiments that diverge or even contradict their explicit messages.Prior works mainly concentrate on simply modeling sarcasm and sentiment features by utilizing the Multi-Task Learning (MTL) framework, which we found introduces detrimental interplays between the sarcasm detection task and sentiment analysis task.Therefore, this study explores the effective enhancement of video sentiment analysis through the incorporation of sarcasm information.To this end, we propose the Progressively Sentiment-oriented Sarcasm Refinement and Integration (PS2RI) framework, which focuses on modeling sentiment-oriented sarcasm features to enhance sentiment prediction.Instead of naively combining sarcasm detection and sentiment prediction under an MTL framework, PS2RI iteratively performs the sentiment-oriented sarcasm refinement and sarcasm integration operations within the sentiment recognition framework, in order to progressively learn sarcasm-aware sentiment feature without suffering the detrimental interplays caused by information irrelevant to the sentiment analysis task.Extensive experiments are conducted to validate the effectiveness of our approach.Code is available at https://github.com/tiggers23/PS2RI.
Junlin Fang, Wenya Wang 0001, Guosheng Lin, Fengmao Lv
ACM Multimedia2
2024 Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity Recognition
abstract
As an important task in multimodal information extraction, Multimodal Named Entity Recognition (MNER) has recently attracted considerable attention. One key challenge of MNER lies in the lack of sufficient fine-grained annotated data, especially in low-resource scenarios. Although data augmentation is a widely used technique to tackle the above issue, it is challenging to simultaneously generate synthetic text-image pairs and their corresponding high-quality entity annotations. In this work, we propose a novel Generative Multimodal Data Augmentation (GMDA) framework for MNER, which contains two stages: Multimodal Text Generation and Multimodal Image Generation. Specifically, we first transform each annotated sentence into a linearized labeled sequence, and then train a Label-aware Multimodal Large Language Model (LMLLM) to generate the labeled sequence based on a label-aware prompt and its associated image. We further employ a Stable Diffusion model to generate the synthetic images that are semantically related to these sentences. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed GMDA framework, which consistently boosts the performance of several competitive methods for two subtasks of MNER in both full-supervision and low-resource settings. The low-resource dataset and source code are released at https://github.com/NUSTM/GMDA.
Jianfei Yu, Wenya Wang 0001, Li Yang 0025
ACM Multimedia4
2024 Robust Domain Misinformation Detection via Multi-Modal Feature Alignment
abstract
Social media misinformation harms individuals and societies and is potentialized by fast-growing multi-modal content (i.e., texts and images), which accounts for higher “credibility” than text-only news pieces. Although existing supervised misinformation detection methods have obtained acceptable performances in key setups, they may require large amounts of labeled data from various events, which can be time-consuming and tedious. In turn, directly training a model by leveraging a publicly available dataset may fail to generalize due to domain shifts between the training data (a.k.a. source domains) and the data from target domains. Most prior work on domain shift focuses on a single modality (e.g., text modality) and ignores the scenario where sufficient unlabeled target domain data may not be readily available in an early stage. The lack of data often happens due to the dynamic propagation trend (i.e., the number of posts related to fake news increases slowly before catching the public attention). We propose a novel robust domain and cross-modal approach (RDCM) for multi-modal misinformation detection. It reduces the domain shift by aligning the joint distribution of textual and visual modalities through an inter-domain alignment module and bridges the semantic gap between both modalities through a cross-modality alignment module. We also propose a framework that simultaneously considers application scenarios of domain generalization (in which the target domain data is unavailable) and domain adaptation (in which unlabeled target domain data is available). Evaluation results on two public multi-modal misinformation detection datasets (Pheme and Twitter Datasets) evince the superiority of the proposed model.
Hui Liu 0036, Wenya Wang 0001, Anderson Rocha 0001, Haoliang Li
IEEE Trans. Inf. Forensics Secur.2
2024 Progressive Multigranularity Information Propagation for Coupled Aspect-Opinion Extraction
abstract
Coupled aspect-opinion extraction aims to identify aspect-opinion pairs in the form of (aspect term, opinion term) or triplets in the form of (aspect term, opinion term, sentiment polarity) from user-generated texts. Compared to the traditional aspect-based sentiment prediction or extraction tasks, coupled aspect-opinion extraction needs to associate aspects with their corresponding opinions and organize opinion-related information into structured outputs. The existing works either divide this task into subproblems (i.e., term extraction and relation prediction) or utilize a unified tagging scheme. However, these methods only focus on atomic word-level interactions and ignore the intensive information propagation among different granularities (e.g., words and word pairs). To address this limitation, we propose a progressive multigranularity information propagation network that progressively explores three types of correlations with different granularities. Specifically, our model starts with the most basic word-level correlations by composing all possible word pairs. In the second stage, the pairwise relation information is used to update the word features. The last stage propagates information among word pairs to produce the relation scores. We treat the task as a unified relation prediction problem and construct an end-to-end framework that iteratively conducts the three-stage information propagation to refine the textual representations. Comprehensive experiments on different aspect-based sentiment analysis benchmarks clearly demonstrate the effectiveness of the proposed approach.
Fengmao Lv, Zhihui Fei, Wenya Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Elaboration-Generating Commonsense Question Answering at Scale
abstract
In question answering requiring common sense, language models (e.g., GPT-3) have been used to generate text expressing background knowledge that helps improve performance.Yet the cost of working with such models is very high; in this work, we finetune smaller language models to generate useful intermediate context, referred to here as elaborations.Our framework alternates between updating two language models-an elaboration generator and an answer predictor-allowing each to influence the other.Using less than 0.5% of the parameters of GPT-3, our model outperforms alternatives with similar sizes and closes the gap with GPT-3 on four commonsense question answering benchmarks.Human evaluations show that the quality of the generated elaborations is high. 1
Wenya Wang 0001, Vivek Srikumar, Hannaneh Hajishirzi, Noah A. Smith
ACL (1)1
2023 Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements
abstract
Today's language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors.Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements.To support diverse commonsense domains, VERA is trained on ∼7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives.When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs.We find that VERA excels at filtering machinegenerated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings.
Jiacheng Liu 0010, Wenya Wang 0001, Dianzhuo Wang, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi
EMNLP2
2023 Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context Learning
abstract
Large language models (LLMs) have showcased their capability with few-shot inference known as in-context learning.However, indomain demonstrations are not always readily available in real scenarios, leading to crossdomain in-context learning.Besides, LLMs are still facing challenges in long-tail knowledge in unseen and unfamiliar domains.The above limitations demonstrate the necessity of Unsupervised Domain Adaptation (UDA).In this paper, we study the UDA problem under an in-context learning setting to adapt language models from the source domain to the target domain without any target labels.The core idea is to retrieve a subset of cross-domain elements that are the most similar to the query, and elicit language model to adapt in an in-context manner by learning both target domain distribution and the discriminative task signal simultaneously with the augmented cross-domain incontext examples.We devise different prompting and training strategies, accounting for different LM architectures to learn the target distribution via language modeling.With extensive experiments on Sentiment Analysis (SA) and Named Entity Recognition (NER) tasks, we thoroughly study the effectiveness of ICL for domain transfer and demonstrate significant improvements over baseline models.
Quanyu Long, Wenya Wang 0001, Sinno Jialin Pan
EMNLP2
2022 Deep Inductive Logic Reasoning for Multi-Hop Reading Comprehension
abstract
Multi-hop reading comprehension requires an ability to reason across multiple documents.On the one hand, deep learning approaches only implicitly encode query-related information into distributed embeddings which fail to uncover the discrete relational reasoning process to infer the correct answer.On the other hand, logic-based approaches provide interpretable rules to infer the target answer, but mostly work on structured data where entities and relations are well-defined.In this paper, we propose a deep-learning based inductive logic reasoning method that firstly extracts query-related (candidate-related) information, and then conducts logic reasoning among the filtered information by inducing feasible rules that entail the target relation.The reasoning process is accomplished via attentive memories with novel differentiable logic operators.To demonstrate the effectiveness of our model, we evaluate it on two reading comprehension datasets, namely WikiHop and MedHop.
Wenya Wang 0001, Sinno Jialin Pan
ACL (1)1
2022 Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge Enhancement
abstract
Sarcasm is a linguistic phenomenon indicating a discrepancy between literal meanings and implied intentions.Due to its sophisticated nature, it is usually challenging to be detected from the text itself.As a result, multi-modal sarcasm detection has received more attention in both academia and industries.However, most existing techniques only modeled the atomic-level inconsistencies between the text input and its accompanying image, ignoring more complex compositions for both modalities.Moreover, they neglected the rich information contained in external knowledge, e.g., image captions.In this paper, we propose a novel hierarchical framework for sarcasm detection by exploring both the atomic-level congruity based on multi-head cross attention mechanism and the composition-level congruity based on graph neural networks, where a post with low congruity can be identified as sarcasm.In addition, we exploit the effect of various knowledge resources for sarcasm detection.Evaluation results on a public multi-modal sarcasm detection dataset based on Twitter demonstrate the superiority of our proposed model.
Hui Liu 0036, Wenya Wang 0001, Haoliang Li
EMNLP2
2022 Domain Confused Contrastive Learning for Unsupervised Domain Adaptation
abstract
In this work, we study Unsupervised Domain Adaptation (UDA) in a challenging selfsupervised approach.One of the difficulties is how to learn task discrimination in the absence of target labels.Unlike previous literature which directly aligns cross-domain distributions or leverages reverse gradient, we propose Domain Confused Contrastive Learning (DCCL) to bridge the source and the target domains via domain puzzles, and retain discriminative representations after adaptation.Technically, DCCL searches for a most domainchallenging direction and exquisitely crafts domain confused augmentations as positive pairs, then it contrastively encourages the model to pull representations towards the other domain, thus learning more stable and effective domain invariances.We also investigate whether contrastive learning necessarily helps with UDA when performing other data augmentations.Extensive experiments demonstrate that DCCL significantly outperforms baselines.
Quanyu Long, Tianze Luo, Wenya Wang 0001, Sinno Jialin Pan
NAACL-HLT3
2022 Weakly Supervised Domain Adaptation for Aspect Extraction via Multilevel Interaction Transfer
abstract
Fine-grained aspect term extraction is an essential subtask in aspect-based opinion analysis. It aims to identify the aspect terms (also known as opinion targets) of a product or service in each sentence. To learn a good aspect extraction model, an expensive annotation process is usually involved to acquire sufficient token-level labels for each domain, which is not realistic. To address this limitation, some previous works propose domain adaptation strategies to transfer knowledge from a sufficiently labeled source domain to unlabeled target domains. However, due to both the difficulty of fine-grained prediction problems and the large domain gap between different domains, the performance is still far from satisfactory. In this work, we conduct a pioneer study on leveraging sentence-level aspect category labels that can be usually available in commercial services, such as review sites or social media to promote token-level transfer for extraction purpose. Specifically, the aspect category information can be used to construct pivot knowledge for transfer with the assumption that the interactions between the sentence-level aspect category and the token-level aspect terms are invariant across domains. To this end, we propose a novel multilevel reconstruction mechanism that aligns both the fine- and coarse-grained information in multiple levels of abstractions. Comprehensive experiments over several benchmark data sets clearly demonstrate that our approach can fully utilize the sentence-level aspect category labels to improve cross-domain aspect term extraction with a large performance gain.
Wenya Wang 0001, Fengmao Lv
IEEE Trans. Neural Networks Learn. Syst.2
2021 Variational Deep Logic Network for Joint Inference of Entities and Relations
abstract
Abstract Currently, deep learning models have been widely adopted and achieved promising results on various application domains. Despite their intriguing performance, most deep learning models function as black boxes, lacking explicit reasoning capabilities and explanations, which are usually essential for complex problems. Take joint inference in information extraction as an example. This task requires the identification of multiple structured knowledge from texts, which is inter-correlated, including entities, events, and the relationships between them. Various deep neural networks have been proposed to jointly perform entity extraction and relation prediction, which only propagate information implicitly via representation learning. However, they fail to encode the intensive correlations between entity types and relations to enforce their coexistence. On the other hand, some approaches adopt rules to explicitly constrain certain relational facts, although the separation of rules with representation learning usually restrains the approaches with error propagation. Moreover, the predefined rules are inflexible and might result in negative effects when data is noisy. To address these limitations, we propose a variational deep logic network that incorporates both representation learning and relational reasoning via the variational EM algorithm. The model consists of a deep neural network to learn high-level features with implicit interactions via the self-attention mechanism and a relational logic network to explicitly exploit target interactions. These two components are trained interactively to bring the best of both worlds. We conduct extensive experiments ranging from fine-grained sentiment terms extraction, end-to-end relation prediction, to end-to-end event extraction to demonstrate the effectiveness of our proposed method.
Wenya Wang 0001, Sinno Jialin Pan
Comput. Linguistics1
2020 Integrating Deep Learning with Logic Fusion for Information Extraction
abstract
Information extraction (IE) aims to produce structured information from an input text, e.g., Named Entity Recognition and Relation Extraction. Various attempts have been proposed for IE via feature engineering or deep learning. However, most of them fail to associate the complex relationships inherent in the task itself, which has proven to be especially crucial. For example, the relation between 2 entities is highly dependent on their entity types. These dependencies can be regarded as complex constraints that can be efficiently expressed as logical rules. To combine such logic reasoning capabilities with learning capabilities of deep neural networks, we propose to integrate logical knowledge in the form of first-order logic into a deep learning system, which can be trained jointly in an end-to-end manner. The integrated framework is able to enhance neural outputs with knowledge regularization via logic rules, and at the same time update the weights of logic rules to comply with the characteristics of the training data. We demonstrate the effectiveness and generalization of the proposed model on multiple IE tasks.
Wenya Wang 0001, Sinno Jialin Pan
AAAI1
2020 Deep Weighted MaxSAT for Aspect-based Opinion Extraction
abstract
Though deep learning has achieved significant success in various NLP tasks, most deep learning models lack the capability of encoding explicit domain knowledge to model complex causal relationships among different types of variables.On the other hand, logic rules offer a compact expression to represent the causal relationships to guide the training process.Logic programs can be cast as a satisfiability problem which aims to find truth assignments to logic variables by maximizing the number of satisfiable clauses (MaxSAT).We adopt the MaxSAT semantics to model logic inference process and smoothly incorporate a weighted version of MaxSAT that connects deep neural networks and a graphical model in a joint framework.The joint model feeds deep learning outputs to a weighted MaxSAT layer to rectify the erroneous predictions and can be trained via end-to-end gradient descent.Our proposed model associates the benefits of highlevel feature learning, knowledge reasoning, and structured learning with observable performance gain for the task of aspect-based opinion extraction.
Meixi Wu, Wenya Wang 0001, Sinno Jialin Pan
EMNLP (1)2
2019 Deep Neural Network Quantization via Layer-Wise Optimization Using Limited Training Data
abstract
The advancement of deep models poses great challenges to real-world deployment because of the limited computational ability and storage space on edge devices. To solve this problem, existing works have made progress to prune or quantize deep models. However, most existing methods rely heavily on a supervised training process to achieve satisfactory performance, acquiring large amount of labeled training data, which may not be practical for real deployment. In this paper, we propose a novel layer-wise quantization method for deep neural networks, which only requires limited training data (1% of original dataset). Specifically, we formulate parameters quantization for each layer as a discrete optimization problem, and solve it using Alternative Direction Method of Multipliers (ADMM), which gives an efficient closed-form solution. We prove that the final performance drop after quantization is bounded by a linear combination of the reconstructed errors caused at each layer. Based on the proved theorem, we propose an algorithm to quantize a deep neural network layer by layer with an additional weights update step to minimize the final error. Extensive experiments on benchmark deep models are conducted to demonstrate the effectiveness of our proposed method using 1% of CIFAR10 and ImageNet datasets. Codes are available in: https://github.com/csyhhu/L-DNQ
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
AAAI2
2019 Transferable Interactive Memory Network for Domain Adaptation in Fine-Grained Opinion Extraction
abstract
In fine-grained opinion mining, aspect and opinion terms extraction has become a fundamental task that provides key information for user-generated texts. Despite its importance, a lack of annotated resources in many domains impede the ability to train a precise model. Very few attempts have applied unsupervised domain adaptation methods to transfer fine-grained knowledge (in the word level) from some labeled source domain(s) to any unlabeled target domain. Existing methods depend on the construction of “pivot” knowledge, e.g., common opinion terms or syntactic relations between aspect and opinion words. In this work, we propose an interactive memory network that consists of local and global memory units. The model could exploit both local and global memory interactions to capture intra-correlations among aspect words or opinion words themselves, as well as the interconnections between aspect and opinion words. The source space and the target space are aligned through these domaininvariant interactions by incorporating an auxiliary task and domain adversarial networks. The proposed model does not require any external resources and demonstrates promising results on 3 benchmark datasets.
Wenya Wang 0001, Sinno Jialin Pan
AAAI1
2019 Cooperative Pruning in Cross-Domain Deep Neural Network Compression
abstract
The advancement of deep models poses great challenges to real-world deployment because of the limited computational ability and storage space on edge devices. To solve this problem, existing works have made progress to compress deep models by pruning or quantization. However, most existing methods rely on a large amount of training data and a pre-trained model in the same domain. When only limited in-domain training data is available, these methods fail to perform well. This prompts the idea of transferring knowledge from a resource-rich source domain to a target domain with limited data to perform model compression. In this paper, we propose a method to perform cross-domain pruning by cooperatively training in both domains: taking advantage of data and a pre-trained model from the source domain to assist pruning in the target domain. Specifically, source and target pruned models are trained simultaneously and interactively, with source information transferred through the construction of a cooperative pruning mask. Our method significantly improves pruning quality in the target domain, and shed light to model compression in the cross-domain setting.
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
IJCAI2
2019 MetaQuant: Learning to Quantize by Learning to Penetrate Non-differentiable Quantization
abstract
Tremendous amount of parameters make deep neural networks impractical to be deployed for edge-device-based real-world applications due to the limit of computational power and storage space. Existing studies have made progress on learning quantized deep models to reduce model size and energy consumption, i.e. converting full-precision weights ($r$'s) into discrete values ($q$'s) in a supervised training manner. However, the training process for quantization is non-differentiable, which leads to either infinite or zero gradients ($g_r$) w.r.t. $r$. To address this problem, most training-based quantization methods use the gradient w.r.t. $q$ ($g_q$) with clipping to approximate $g_r$ by Straight-Through-Estimator (STE) or manually design their computation. However, these methods only heuristically make training-based quantization applicable, without further analysis on how the approximated gradients can assist training of a quantized network. In this paper, we propose to learn $g_r$ by a neural network. Specifically, a meta network is trained using $g_q$ and $r$ as inputs, and outputs $g_r$ for subsequent weight updates. The meta network is updated together with the original quantized network. Our proposed method alleviates the problem of non-differentiability, and can be trained in an end-to-end manner. Extensive experiments are conducted with CIFAR10/100 and ImageNet on various deep networks to demonstrate the advantage of our proposed method in terms of a faster convergence rate and better performance. Codes are released at: \texttt{https://github.com/csyhhu/MetaQuant}
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
NeurIPS2
2019 Syntactically Meaningful and Transferable Recursive Neural Networks for Aspect and Opinion Extraction
abstract
In fine-grained opinion mining, extracting aspect terms (a.k.a. opinion targets) and opinion terms (a.k.a. opinion expressions) from user-generated texts is the most fundamental task in order to generate structured opinion summarization. Existing studies have shown that the syntactic relations between aspect and opinion words play an important role for aspect and opinion terms extraction. However, most of the works either relied on predefined rules or separated relation mining with feature learning. Moreover, these works only focused on single-domain extraction, which failed to adapt well to other domains of interest where only unlabeled data are available. In real-world scenarios, annotated resources are extremely scarce for many domains, motivating knowledge transfer strategies from labeled source domain(s) to any unlabeled target domain. We observe that syntactic relations among target words to be extracted are not only crucial for single-domain extraction, but also serve as invariant “pivot” information to bridge the gap between different domains. In this article, we explore the constructions of recursive neural networks based on the dependency tree of each sentence for associating syntactic structure with feature learning. Furthermore, we construct transferable recursive neural networks to automatically learn the domain-invariant fine-grained interactions among aspect words and opinion words. The transferability is built on an auxiliary task and a conditional domain adversarial network to reduce domain distribution difference in the hidden spaces effectively in word level through syntactic relations. Specifically, the auxiliary task builds structural correspondences across domains by predicting the dependency relation for each path of the dependency tree in the recursive neural network. The conditional domain adversarial network helps to learn domain-invariant hidden representation for each word conditioned on the syntactic structure. In the end, we integrate the recursive neural network with a sequence labeling classifier on top that models contextual influence in the final predictions. Extensive experiments and analysis are conducted to demonstrate the effectiveness of the proposed model and each component on three benchmark data sets.
Wenya Wang 0001, Sinno Jialin Pan
Comput. Linguistics1
2018 Recursive Neural Structural Correspondence Network for Cross-domain Aspect and Opinion Co-Extraction
abstract
Fine-grained opinion analysis aims to extract aspect and opinion terms from each sentence for opinion summarization.Supervised learning methods have proven to be effective for this task.However, in many domains, the lack of labeled data hinders the learning of a precise extraction model.In this case, unsupervised domain adaptation methods are desired to transfer knowledge from the source domain to any unlabeled target domain.In this paper, we develop a novel recursive neural network that could reduce domain shift effectively in word level through syntactic relations.We treat these relations as invariant "pivot information" across domains to build structural correspondences and generate an auxiliary task to predict the relation between any two adjacent words in the dependency tree.In the end, we demonstrate state-ofthe-art results on three benchmark datasets.
Wenya Wang 0001, Sinno Jialin Pan
ACL (1)1
2018 Transition-based Adversarial Network for Cross-lingual Aspect Extraction
abstract
In fine-grained opinion mining, the task of aspect extraction involves the identification of explicit product features in customer reviews. This task has been widely studied in some major languages, e.g., English, but was seldom addressed in other minor languages due to the lack of annotated corpus. To solve it, we develop a novel deep model to transfer knowledge from a source language with labeled training data to a target language without any annotations. Different from cross-lingual sentiment classification, aspect extraction across languages requires more fine-grained adaptation. To this end, we utilize transition-based mechanism that reads a word each time and forms a series of configurations that represent the status of the whole sentence. We represent each configuration as a continuous feature vector and align these representations from different languages into a shared space through an adversarial network. In addition, syntactic structures are also integrated into the deep model to achieve more syntactically-sensitive adaptations. The proposed method is end-to-end and achieves state-of-the-art performance on English, French and Spanish restaurant review datasets.
Wenya Wang 0001, Sinno Jialin Pan
IJCAI1
2018 Memory networks for fine-grained opinion mining
Wenya Wang 0001, Sinno Jialin Pan, Daniel Dahlmeier
Artif. Intell.1
2017 Coupled Multi-Layer Attentions for Co-Extraction of Aspect and Opinion Terms
abstract
The task of aspect and opinion terms co-extraction aims to explicitly extract aspect terms describing features of an entity and opinion terms expressing emotions from user-generated texts. To achieve this task, one effective approach is to exploit relations between aspect terms and opinion terms by parsing syntactic structure for each sentence. However, this approach requires expensive effort for parsing and highly depends on the quality of the parsing results. In this paper, we offer a novel deep learning model, named coupled multi-layer attentions. The proposed model provides an end-to-end solution and does not require any parsers or other linguistic resources for preprocessing. Specifically, the proposed model is a multi-layer attention network, where each layer consists of a couple of attentions with tensor operators. One attention is for extracting aspect terms, while the other is for extracting opinion terms. They are learned interactively to dually propagate information between aspect terms and opinion terms. Through multiple layers, the model can further exploit indirect relations between terms for more precise information extraction. Experimental results on three benchmark datasets in SemEval Challenge 2014 and 2015 show that our model achieves state-of-the-art performances compared with several baselines.
Wenya Wang 0001, Sinno Jialin Pan, Daniel Dahlmeier, Xiaokui Xiao
AAAI1
2016 Recursive Neural Conditional Random Fields for Aspect-based Sentiment Analysis
abstract
In aspect-based sentiment analysis, extracting aspect terms along with the opinions being expressed from user-generated content is one of the most important subtasks.Previous studies have shown that exploiting connections between aspect and opinion terms is promising for this task.In this paper, we propose a novel joint model that integrates recursive neural networks and conditional random fields into a unified framework for explicit aspect and opinion terms co-extraction.The proposed model learns high-level discriminative features and double propagates information between aspect and opinion terms, simultaneously.Moreover, it is flexible to incorporate hand-crafted features into the proposed model to further boost its information extraction performance.Experimental results on the dataset from SemEval Challenge 2014 task 4 show the superiority of our proposed model over several baseline methods as well as the winning systems of the challenge.
Wenya Wang 0001, Sinno Jialin Pan, Daniel Dahlmeier, Xiaokui Xiao
EMNLP1