VLDB 2026 Research / reviewers in the wild / expert
Rui Mao 0010
dblp:324/1909
· DBLP profile ↗
63ranked-venue papers
11as first author
60since 2021 · last 2026
0000-0002-1082-8755ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 11 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLER: Improving Multimodal Financial Reasoning by Cross-MLLM Error ReflectionabstractRecent advances in Multimodal Large Language Models (MLLMs) have enabled joint reasoning over financial textual and visual inputs. However, they still struggle with financial terminology, logical consistency, and numerical computations. Moreover, while commercial large models perform well on reasoning tasks, their high inference costs limit their scalable usage in real world financial applications. We thus propose a cost-effective framework, CLER, that combines contrastive retrieval with step-wise reflection to improve reasoning performance. Also, the reasoning cost is only generated in the test stage when using commercial large models. CLER leverages FinErrorSet, a dataset of 8,000+ mistake correction pairs from diverse open-source MLLMs. A fine grained retriever is trained to identify structurally relevant errors for self-correction through individual reflection. Experiments on three benchmarks show that CLER consistently outperforms other baselines. To our knowledge, CLER is the first framework to use cross-model errors for financial reasoning. Shuangyan Deng, Zhongsheng Wang, Rui Mao 0010, Ciprian Doru Giurcaneanu, Jiamou Liu |
AAAI | 3 |
| 2026 | RefleXNet: Targeted Self-Reflection for Accurate Chest X-ray ReportingabstractAutomated interpretation and reporting of chest X-rays (CXRs) hold significant promise in reducing diagnostic errors and supporting radiologists under heavy clinical workloads. However, existing methods typically rely on global visual features and token-level supervision, limiting their sensitivity to subtle abnormalities and reducing their clinical reliability. To address these challenges, we present Reflective X-ray Network (RefleXNet), which systematically integrates multi-scale visual feature fusion and anatomical relational reasoning with a targeted self-reflective learning strategy. RefleXNet first constructs multi-scale visual representations and captures anatomical context through graph-based relational modeling. Building upon these representations, we introduce a targeted self-reflection strategy that uses clinically guided feedback from generated reports to selectively refine abnormality predictions and their associated region-level visual features. Extensive experiments on MIMIC-CXR demonstrate that RefleXNet consistently outperforms state-of-the-art baselines across clinical factual correctness metrics. Notably, our compact 3B-parameter model surpasses several recent models with over twice the parameter count. Additionally, RefleXNet exhibits strong generalization performance in zero-shot evaluations on IU-Xray compared with leading multimodal language models, highlighting its robustness and clinical effectiveness. Xin Mei, Rui Mao 0010, Xiaoyan Cai, Libin Yang, Erik Cambria |
AAAI | 2 |
| 2026 | Are Language Models Any Good at Density Modeling?abstractLarge Language Models (LLMs) surprised the world with their ability to mimic humans in writing and are starting to be used as simulations of human writers for various kinds of linguistic analyses. However, these analyses rest on the belief that LLMs are good density models that accurately capture the underlying probability distribution of the language. In this paper, we question this basic assumption and try to evaluate language models on their density modelling capabilities. Since a ground truth does not exist for the probability distribution of any natural language, we come up with a synthetic language made up of decimal numbers written in words in English. We train language models from scratch on various probability distributions over this synthetic language and compare the distributions learned by the models with the original distributions. Experiments show that language models can learn underlying probability distributions across a wide range of cases, but they fail when those distributions depend on deep semantic properties of numbers that cannot be inferred from syntactic patterns. Additionally, we observed a strong bias in the models towards numbers that frequently occur as substrings within other numbers. This suggests that such a bias possibly exists in real-world natural language models as well, and negatively impacts downstream tasks and analyses that rely on model-generated probabilities. Sriram Ranga, Sai Shashank Bedampeta, Rui Mao 0010, Anupam Chattopadhyay |
AAAI | 3 |
| 2026 | LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models
Tiesunlong Shen, Rui Mao 0010, Jin Wang 0008, Heming Sun, Jian Zhang 0087, Xuejie Zhang 0002, Erik Cambria |
AAAI | 2 |
| 2026 | MAPS: Multi-Agent Personality Shaping for Collaborative ReasoningabstractCollaborative reasoning with multiple agents offers the potential for more robust and diverse problem-solving. However, existing approaches often suffer from homogeneous agent behaviors and lack of reflective and rethinking capabilities. We propose Multi-Agent Personality Shaping ((MAPS), a novel framework that enhances reasoning through agent diversity and internal critique. Inspired by the Big Five personality theory, MAPS assigns distinct personality traits to individual agents, shaping their reasoning styles and promoting heterogeneous collaboration. To enable deeper and more adaptive reasoning, MAPS introduces a Critic agent that reflects on intermediate outputs, revisits flawed steps, and guides iterative refinement. This integration of personality-driven agent design and structured collaboration improves both reasoning depth and flexibility. Empirical evaluations across three benchmarks demonstrate the strong performance of MAPS, with further analysis confirming its generalizability across different large language models and validating the benefits of multi-agent collaboration. Jian Zhang 0087, Zhangqi Wang, Fangzhi Xu, Qika Lin, Lingling Zhang 0005, Rui Mao 0010, Erik Cambria, Jun Liu 0002 |
AAAI | 7 |
| 2026 | The Evolution of Philosophy: A Metaphorical Cognition Perspective
Rui Mao 0010, Dapeng Chen, Xulang Zhang, Erik Cambria |
LREC | 1 |
| 2026 | Mapping Liberty Metaphors across Cultures and Time
Sidney Suen, Rui Mao 0010, Kenneth Kwok, Erik Cambria |
LREC | 2 |
| 2026 | Evaluating Multimodal Large Language Model Narrative Interpretation through the Lens of Appraisal Theory
Jayant Teotia, Xulang Zhang, Rui Mao 0010, Erik Cambria |
LREC | 4 |
| 2026 | Appraisal Theory-Informed Emotion Prediction
Jayant Teotia, Rui Mao 0010, Wandeep Kaur Ratan Singh, Sabrina Binti Tiun, Erik Cambria |
LREC | 3 |
| 2026 | Implicit Bias in Peer Review: Through the Lens of Language Abstraction
Xulang Zhang, Rui Mao 0010, Erik Cambria |
LREC | 2 |
| 2026 | Towards smart city supervision: A detection pipeline for illegal buildings
Shudong Zhang, Rui Mao 0010 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Multi-granularity semantic extraction and multi-task fusion for Chinese medical entity normalization
Kai He 0001, Rui Mao 0010, Mengling Feng |
Expert Syst. Appl. | 3 |
| 2026 | Language models for environmental, social, and governance analysis: A review
Kelvin Du, Rui Mao 0010, Frank Z. Xing, Gianmarco Mengaldo, Erik Cambria |
Inf. Process. Manag. | 2 |
| 2026 | AOSNet-Sec: Aperture-orientation-spectrum fusion with statistical Markov repair for trustworthy super-resolution
Zhengnan Yin, Luwei Xiao, Xuan Feng 0002, Yiwei Chen 0002, Xianxun Zhu, Cai Luo, Faten S. Alamri, Rui Mao 0010, Erik Cambria |
Pattern Recognit. | 8 |
| 2026 | SenticNet 9: Generative Commonsense for Emotion AI via Conceptual Primitive Discovery and Time Shift MechanismabstractLarge language models (LLMs) generate fluent, context-rich text but suffer from hallucinations and limited interpretability. We introduce SenticNet 9, a neurosymbolic framework that automates commonsense reasoning while preserving transparency. It leverages conceptual primitive discovery (CPD) to learn foundational concepts and a time shift mechanism (TSM) to iteratively refine them through temporal feedback. This combination yields a scalable, cognitively inspired architecture that merges symbolic interpretability with LLM generalization. Experiments show SenticNet 9 outperforming embeddings, transformers, and state-of-the-art LLMs across tasks, delivering higher accuracy without sacrificing explainability. Erik Cambria, Rui Mao 0010, Xulang Zhang, Luwei Xiao, Tiesunlong Shen, Avinash Anand |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Towards Robust ESG Analysis Against Greenwashing Risks: Aspect-Action Analysis with Cross-Category GeneralizationabstractSustainability reports are key for evaluating companies' environmental, social and governance (ESG) performance.To analyze these reports, NLP approaches can efficiently extract ESG insights at scale.However, even the most advanced NLP methods lack robustness against ESG content that is greenwashed -i.e.sustainability claims that are misleading, exaggerated, and fabricated.Accordingly, existing NLP approaches often extract insights that reflect misleading or exaggerated sustainability claims rather than objective ESG performance.To tackle this issue, we introduce A3CG -Aspect-Action Analysis with Cross-Category Generalization, as a novel dataset to improve the robustness of ESG analysis amid the prevalence of greenwashing.By explicitly linking sustainability aspects with their associated actions, A3CG facilitates a more fine-grained and transparent evaluation of sustainability claims, ensuring that insights are grounded in verifiable actions rather than vague or misleading rhetoric.Additionally, A3CG emphasizes cross-category generalization.This ensures robust model performance in aspectaction analysis even when companies change their reports to selectively favor certain sustainability areas.Through experiments on A3CG, we analyze state-of-the-art supervised models and LLMs, uncovering their limitations and outlining key directions for future research. Keane Ong, Rui Mao 0010, Deeksha Varshney, Erik Cambria, Gianmarco Mengaldo |
ACL (1) | 2 |
| 2025 | AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeabstractXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang |
ACL (1) | 8 |
| 2025 | QAEval: Mixture of Evaluators for Question-Answering Task EvaluationabstractQuestion answering (QA) tasks serve as a key benchmark for evaluating generation systems.Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with openended and unstructured responses.LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to instructions, robustness issues, and high computational costs.To overcome these challenges, we introduce QAEval, a hybrid framework combining rule-based reliability with LLM-based adaptability.QAEval utilizes two high-quality datasets: QAExtract for short-answer extraction and QAScore for scoring model training.By integrating a Mixture of Evaluators model with Dynamic Load Balancing Optimization, QAEval enables accurate, cost-effective QA evaluation.Experimental results show it outperforms models like GPT-4o and Claude-3, achieving 92.3% accuracy with only 0.6B parameters. Tan Yue, Rui Mao 0010, Xuzhao Shi, Shuo Zhan, Zuhao Yang, Dongyan Zhao 0001 |
ACL (1) | 2 |
| 2025 | Neuro-Symbolic AI in HealthcareabstractMedical AI has achieved strong predictive performance, yet most systems remain limited by shallow reasoning, poor transparency, and weak generalisation in safety-critical settings. Neurosymbolic AI offers a path beyond these constraints by combining neural models' ability to learn from complex clinical data with the explicit structure, logic, and domain knowledge of symbolic methods. This article examines how neurosymbolic approaches can address core challenges in healthcare AI through five key areas: hybrid reasoning that unifies learning and logic; symbol grounding that links internal representations to clinically meaningful concepts; clinical interpretability that exposes reasoning steps; human-integrated decision-making that keeps clinicians in control; and knowledge-driven diagnosis that incorporates guidelines, ontologies, and causal understanding. Together, these elements outline how neurosymbolic AI can support systems that are not only accurate but also transparent, clinically aligned, and robust in complex or data-sparse scenarios. Advancing this paradigm will require collaboration across AI research, clinical practice, and knowledge engineering, as well as governance mechanisms that ensure fairness and accountability. Neurosymbolic AI thus represents a promising direction for building trustworthy, knowledge-rich intelligence in healthcare. Jialun Wu, Xin Mei, Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Rui Mao 0010 |
BIBM | 7 |
| 2025 | Bridging Cognitive Divide: Uncovering Cognitive Disparities among Diabetic Patients via MetaphorabstractDiabetes self-management continues to exhibit wide interpersonal variability, a phenomenon insufficiently addressed by standardized Diabetes Self-Management Education (DSME) programs. The limitations of existing approaches stem largely from their implicit assumption of homogeneous cognitive structures, an assumption inconsistent with theoretical insights from conceptual metaphor theory and cognitive science indicating that individuals rely on divergent conceptual mappings to make sense of chronic illness. This study addresses this gap by conducting a large-scale computational analysis of 39,880 spontaneous patient self-disclosures from diabetes-related subreddits using MetaPro to uncover the underlying cognitive frameworks of diabetic patients. Comparative analysis demonstrated significant differences in cognitive conceptualizations between type 1 (T1D) and type 2 (T2D) diabetes patients. T1D discourse is dominated by an “action-monitoring” schema, whereas T2D discourse emphasizes a “process-mishap-component” schema. Gender comparisons further distinguish a male “metric-status-trajectory” profile from a female “burden-care-intuition” profile. These findings illuminate the deep cognitive heterogeneity underlying diabetes self-management and underscore the need for personalized, cognitively informed intervention strategies that attend to demographic and clinical distinctions. Wang Zhao 0002, Rui Mao 0010, Shuangyan Deng, Erik Cambria |
BIBM | 2 |
| 2025 | Decoding Metaphors and Brain Signals in Naturalistic Contexts: An Empirical Study based on EEG and MetaPro
Rui Mao 0010, Tao Wang 0049, Erik Cambria |
CogSci | 1 |
| 2025 | Cognitive Mechanisms in Loan Marketing: Insights from Concept Mappings
Rui Mao 0010, Tao Wang 0049, Xulang Zhang, Xinlong Li, Erik Cambria |
CogSci | 1 |
| 2025 | Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual GenerationabstractCounterfactual reasoning typically involves considering alternatives to actual events.While often applied to understand past events, a distinct form-forward counterfactual reasoningfocuses on anticipating plausible future developments.This type of reasoning is invaluable in dynamic financial markets, where anticipating market developments can powerfully unveil potential risks and opportunities for stakeholders, guiding their decision-making.However, performing this at scale is challenging due to the cognitive demands involved, underscoring the need for automated solutions.LLMs offer promise, but remain unexplored for this application.To address this gap, we introduce a novel benchmark, FIN-FORCE-FINancial FORward Counterfactual Evaluation.By curating financial news headlines and providing structured evaluation, FIN-FORCE supports LLM based forward counterfactual generation.This paves the way for scalable and automated solutions for exploring and anticipating future market developments, thereby providing structured insights for decision-making.Through experiments on FIN-FORCE, we evaluate state-of-theart LLMs and counterfactual generation methods, analyzing their limitations and proposing insights for future research.We release the benchmark, supplementary data and all experimental codes at the following link Keane Ong, Rui Mao 0010, Deeksha Varshney, Paul Pu Liang, Erik Cambria, Gianmarco Mengaldo |
EMNLP | 2 |
| 2025 | F2TEval: Human-Aligned Multi-Dimensional Evaluation for Figure-to-Text TaskabstractFigure -to-Text (F2T) tasks aim to convert structured figure information into natural language text, serving as a bridge between visual perception and language understanding.However, existing evaluation methods remain limited: 1) Reference-based methods can only capture shallow semantic similarities and rely on costly labeled reference text; 2) Reference-free methods depend on multimodal large language models, which suffer from low efficiency and instruction sensitivity; 3) Existing methods provide only sample-level evaluations, lacking interpretability and alignment with expert-level multi-dimensional evaluation criteria.Accordingly, we propose F2TEval, a five-dimensional reference-free evaluation method aligned with expert criteria, covering faithfulness, completeness, conciseness, logicality, and analysis, to support fine-grained evaluation.We design a lightweight mixture-of-experts model that incorporates independent scoring heads and applies the Hilbert-Schmidt Independence Criterion to optimize the disentanglement of scoring representations across dimensions.Furthermore, we construct F2TBenchmark, a humanannotated benchmark dataset covering 21 chart types and 35 application domains, to support research on F2T evaluation.Experimental results demonstrate our model's superior performance and efficiency, outperforming Gemini-2.0and Claude-3.5 with only 0.9B parameters. Tan Yue, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001 |
EMNLP | 2 |
| 2025 | Can a Large Language Model be a Gaslighter?abstractLarge language models (LLMs) have gained human trust due to their capabilities and helpfulness. However, this in turn may allow LLMs to affect users' mindsets by manipulating language. It is termed as gaslighting, a psychological effect.
In this work, we aim to investigate the vulnerability of LLMs under prompt-based and fine-tuning-based gaslighting attacks. Therefore, we propose a two-stage framework DeepCoG designed to: 1) elicit gaslighting plans from LLMs with the proposed DeepGaslighting prompting template, and 2) acquire gaslighting conversations from LLMs through our Chain-of-Gaslighting method. The gaslighting conversation dataset along with a corresponding safe dataset is applied to fine-tuning-based attacks on open-source LLMs and anti-gaslighting safety alignment on these LLMs. Experiments demonstrate that both prompt-based and fine-tuning-based attacks transform three open-source LLMs into gaslighters. In contrast, we advanced three safety alignment strategies to strengthen~(by $12.05\%$) the safety guardrail of LLMs. Our safety alignment strategies have minimal impacts on the utility of LLMs. Empirical studies indicate that an LLM may be a potential gaslighter, even if it passed the harmfulness test on general dangerous queries. Wei Li 0076, Ruixi Lin, Rui Mao 0010 |
ICLR | 5 |
| 2025 | AnaFig: A Human-Aligned Dataset for Scientific Figure AnalysisabstractScientific Figure Analysis (SFA) aims to derive analytical insights from figures while incorporating background instructions. Unlike conventional tasks such as figure captioning or description generation, which focus on extracting surface-level information from the sole visual modality, SFA requires an intelligent system to summarize key patterns, infer implications, and contextualize scientific findings from visual and textual inputs. It demands not only visual recognition but also the integration of scientific knowledge, multimodal understanding, and contextual reasoning. In this work, we introduce an SFA dataset, AnaFig, comprising 2,000 high-quality samples across 56 domains. All samples are evaluated by using human-aligned five-dimensional scoring criteria, resulting 10,000 human-annotated score labels. The AnaFig dataset facilitates the assessment of three critical capabilities of multimodal large language models (MLLMs): adherence to complex instructions, multimodal perception, and analytical summarization. By building a new benchmark with widely used MLLMs, this study contributes to scientific knowledge discovery and reasoning, fostering the alignment of MLLMs and human experts in scientific analysis. Tan Yue, Xuzhao Shi, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001 |
ACM Multimedia | 3 |
| 2025 | The Plagiarism Singularity ConjectureabstractSriram Ranga, Rui Mao, Erik Cambria, Anupam Chattopadhyay. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sriram Ranga, Rui Mao 0010, Erik Cambria, Anupam Chattopadhyay |
NAACL (Long Papers) | 2 |
| 2025 | Flow-guided Direct Preference Optimization for Knowledge Graph Reasoning with TreesabstractRecent advancements in knowledge graph question answering (KGQA) have shown promise, yet existing methods often fail to align with human reasoning patterns that involve continuous reflection and refinement. This paper proposes FD-PORT (flow-guided direct preference optimization for knowledge graph reasoning with trees), a novel approach that combines Monte Carlo Tree Search (MCTS) with flow-guided direct preference optimization (FDPO) for KGQA tasks. MCTS simulates human-like reasoning by systematically exploring multiple inference paths in knowledge graphs, while FDPO transforms the search feedback into fine-grained training signals through flow balance conditions. Unlike traditional methods focusing on end-to-end training or sequence-level preferences, FD-PORT establishes flow consistency between any states along the reasoning chain, enabling robust multi-hop reasoning that adapts to local decisions and long-range dependencies. Experimental results on three benchmark datasets demonstrate that FD-PORT significantly outperforms state-of-the-art methods, achieving up to 50.6% improvements over GPT-4 on complex multi-hop reasoning tasks with a smaller open-source language model. The framework is advanced in maintaining diverse reasoning paths while ensuring answer quality, closely mirroring human problem-solving strategies. Tiesunlong Shen, Rui Mao 0010, Jin Wang 0008, Xuejie Zhang 0002, Erik Cambria |
SIGIR | 2 |
| 2025 | CSA-RSIC: Cross-Modal Semantic Alignment for Remote Sensing Image CaptioningabstractRemote sensing image captioning (RSIC) is an important task in environmental monitoring and disaster assessment. However, existing methods are constrained by redundant feature interference, insufficient multi-scale feature integration, and cross-modal semantic gaps, leading to limited performance in scenarios requiring fine-grained descriptions and semantic integrity, such as disaster assessment and emergency response. In this letter, we propose a Cross-Modal Semantic Alignment Model for Remote Sensing Image Captioning (CSA-RSIC), addressing these challenges with three innovations. First, we designed an Adaptive Feature Selection Module (AFSM) that generates channel weights through dual pooling. The AFSM dynamically weights the most informative features at each scale to improve caption accuracy. Second, we propose a Cross-Scale Feature Aggregation Module (CFAM) that constructs a hierarchical feature pyramid by aligning multi-scale resolutions and performs attention-guided fusion with enhanced weighting via AFSM, ensuring the effective integration of fine-grained and global semantic information. Finally, a novel loss function that combines contrastive learning and consistency loss is proposed to enhance the semantic alignment between visual and textual features. Experiments on three datasets show the advancement of CSA-RSIC over strong baselines, indicating its effectiveness in enhancing both semantic completeness and accuracy. Kangda Cheng, Rui Mao 0010, Zhilu Wu, Erik Cambria |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Public Opinion Crisis Management via Social Media MiningabstractEnterprises suffer tremendous economic and reputational losses in public opinion crises. However, providing specific guidance for enterprises to respond effectively to public opinion crises is challenging. This study proposes a netizen-centered approach to identify crisis response opportunities for enterprises by mining social media data. First, we identify the topics discussed by netizens with the Biterm Topic Model, thereby evaluating their importance levels. Then, the negativeness levels of crisis topics are quantified by sentiment analysis. Lastly, crisis response opportunities are quantitatively identified and prioritized by applying an opportunity algorithm that simultaneously considers each crisis topic's importance and negativeness levels. Experimental results demonstrate the validity and superiority of our approach. Moreover, we demonstrate that a very negative and less important topic in the growth stage will likely become very negative and important at the maturity of the crisis. This finding supports decision-making in response to critical topics at the early stage of a crisis, preventing the deterioration of the public opinion crisis. Rui Mao 0010, Peng Wu 0032, Erik Cambria |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Exploring Cognitive and Aesthetic Causality for Multimodal Aspect-Based Sentiment AnalysisabstractMultimodal aspect-based sentiment classification (MASC) is an emerging task due to an increase in user-generated multimodal content on social platforms, aimed at predicting sentiment polarity toward specific aspect targets (i.e., entities or attributes explicitly mentioned in text-image pairs). Despite extensive efforts and significant achievements in existing MASC, substantial gaps remain in understanding fine-grained visual content and the cognitive rationales derived from semantic content and impressions (cognitive interpretations of emotions evoked by image content). In this study, we present Chimera: acognitive and aesthetic sentiment causality understanding framework to derive fine-grained holistic features of aspects and infer the fundamental drivers of sentiment expression from both semantic perspectives and affective-cognitive resonance (the synergistic effect between emotional responses and cognitive interpretations). The framework aligns visual patches with words, extracts coarse and fine-grained visual features, translates them into textual descriptions, and uses LLM-generated sentimental causes and impressions to boost sensitivity to affective cues. Experiments on MASC datasets show the model's effectiveness and greater flexibility compared to LLMs like GPT-4o. We have publicly released the complete implementation and dataset athttps://github.com/Xillv/Chimera Luwei Xiao, Rui Mao 0010, Shuai Zhao 0007, Qika Lin, Yanhao Jia, Liang He 0001, Erik Cambria |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | A Discourse Structure- and Interlocutor-Guided Network for Dialogue Act Recognition and Sentiment ClassificationabstractDialogue act recognition and sentiment classification (DAR and DSC) are closely related tasks in dialogue systems, both benefiting from joint modeling of their interdependencies. Recent advancements have improved performance on these tasks by integrating them, yet many approaches oversimplify dialogues by treating them as monologues and assuming uniform influence across all utterances. This neglects the inherent structural and interactive nature of dialogues. To address these issues, we propose a novelDiscourseStructure- andInterlocutor-Guided (DSIG) network that fuses dialogue act recognition and sentiment classification. Our network synergizes structural dialogue relationships with interlocutor identity information, enabling effective modeling of utterance flow and cross-task interactions. Specifically, we utilize a shared encoder that functions as a dialogue discourse parser to dynamically construct utterance connections, thereby integrating dialogue structural relationships. In addition, we embed interlocutor affiliations into the fusion process during encoding, semantic modeling, and decoding, enhancing dialogue understanding and dual-task reasoning. The core component is a collaborative updating graph interaction layer, which filters redundant connections and introduces interlocutor nodes for exclusive information exchange. Experimental results on two benchmark datasets demonstrate DSIG achieves state-of-the-art performance by effectively fusing dialogue structure and interlocutor information, improving F1 scores by 6.9% and 2.1% for DSC and DAR, respectively, on Mastodon, and by 13.7% and 4.7% on DailyDialog. Model variants trained independently show promise for extension to other dialogue-related tasks. Yunhe Xie, Rui Mao 0010, Wei Li 0076, Atika Qazi, Erik Cambria |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | PGIF: A Personality-Guided Iterative Feedback Graph Network for Multimodal Conversational Emotion RecognitionabstractMultimodal emotion recognition in conversation (MERC) aims to identify emotions in target utterances from multimodal records, drawing significant attention for its value in conversational artificial intelligence. While early research focuses on exploring conversational context, recent efforts emphasize integrating multimodal cues. Existing methods focus on modeling the impact of conversational context on emotion recognition while neglecting the role of the speaker’s personality factors. Furthermore, these approaches often suffer from inefficiencies in information transfer due to full-utterance connectivity and fail to leverage multiple fusion modes for complementary benefits. To address these issues, we propose a personality-guided iterative feedback graph network (PGIF) for MERC. PGIF incorporates personality information as a specialized modality to enhance the feature space for emotional inference. We utilize a graph network to model information flow, integrating interlocutor-aware contextual information by considering interlocutor dependencies between utterances. Additionally, we employ a dialogue discourse parser to directly model semantic relationships between utterances. Our iterative feedback fusion mechanism explicitly simulates emotional interactions between feature-level and decision-level modalities, improving inference without ground truth labels through iterative refinement. PGIF demonstrates improvements of 1.94% and 1.42% over state-of-the-art methods on the IEMOCAP and MELD datasets, respectively. Ablation studies further validate the effectiveness of PGIF’s mechanisms, while its manipulation of input features and global fusion strategies ensures compatibility with existing approaches. Yunhe Xie, Rui Mao 0010 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | RTL Agent: An Agent-Based Approach for Functionally Correct HDL Generation via LLMsabstractLLMs as code generators have undergone rapid progress over the past couple of years. However, the models on their own provide no guarantees for the functional correctness of the generated code. Functional tests can not only be used by designers to assess the functional correctness of code, but also to guide them towards the solution of the problem. The same can be applied to LLMs performing automatic code generation through the use of the Reflexion technique. Reflexion is an agent-based workflow where the model generating code iterates over a loop of code generation, getting feedback from the test bench, reflecting on the feedback, and making appropriate changes to the code. The technique is known to drastically improve the performance of LLMs on software code generation. In this work, we adopt the technique for hardware description language (HDL) code generation as RTL Agent - an implementation of the workflow for Verilog generation with Reflexion. We compare multiple LLMs with standard inference vs with Reflexion on the VerilogEval benchmark. Within 5 iterations of the feedback loop, we observe a relative improvement of 33.2% for the low-performance model Llama3 and an average of 17.7% for the high-performance models GPT-4o and GPT-4o-mini in their Pass@1 performance. We present a cost analysis of the technique to enable cost-performance trade-offs. Sriram Ranga, Rui Mao 0010, Debjyoti Bhattacharjee, Erik Cambria, Anupam Chattopadhyay |
ATS | 2 |
| 2024 | Explainable Stock Price Movement Prediction using Contrastive LearningabstractPredicting stock price movements is a high-stakes task that demands explainability for human decision-makers. A key shortcoming in current methods is treating sub-predictions independently, without learning from accumulated experiences. We propose a novel triplet network for contrastive learning to enhance the explainability of stock movement prediction by considering instances of "integrated textual information and quantitative indicators". We refer to the target past-l-day tweet-price time series as the "anchor instance". Each anchor instance is paired with a "positive instance" characterized by highly correlated return trends yet significant differences across the entire feature space, and a "negative instance" that exhibits similar return trends along with high proximity in the feature space. The model is designed with the objective of (1) minimizing the cross entropy loss between input logits and target, (2) minimizing the distance between the anchor instances and positive instances, and (3) maximizing the distance between the anchor instances and negative instances. Our framework's effectiveness is demonstrated through extensive testing, showing superior performance on stock prediction benchmarks. Kelvin Du, Rui Mao 0010, Frank Z. Xing, Erik Cambria |
CIKM | 2 |
| 2024 | Unveiling Diplomatic Narratives: Analyzing United Nations Security Council Debates Through Metaphorical Cognition
Rui Mao 0010, Qian Liu 0012, Amir Hussain 0001, Erik Cambria |
CogSci | 1 |
| 2024 | GPTEval: A Survey on Assessments of ChatGPT and GPT-4abstractThe emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong curiosity among scholars about its performance in different domains. There have been many studies evaluating the ability of ChatGPT and GPT-4 in different tasks and disciplines. However, a comprehensive review summarizing the collective assessment findings is lacking. The objective of this survey is to thoroughly analyze prior assessments of ChatGPT and GPT-4, focusing on its language and reasoning abilities, scientific knowledge, and ethical considerations. Furthermore, an examination of the existing evaluation methods is conducted, offering several recommendations for future research. Rui Mao 0010, Guanyi Chen, Xulang Zhang, Frank Guerin, Erik Cambria |
LREC/COLING | 1 |
| 2024 | SarcNet: A Multilingual Multimodal Sarcasm Detection DatasetabstractSarcasm poses a challenge in linguistic analysis due to its implicit nature, involving an intended meaning that contradicts the literal expression. The advent of social networks has propelled the utilization of multimodal data to enhance sarcasm detection performance. In prior multimodal sarcasm detection datasets, a single label is assigned to a multimodal instance. Subsequent experiments often highlight the superiority of multimodal models by demonstrating their improvements compared to unimodal models based on these unified labels across multiple modalities. However, our investigation revealed that numerous instances of sarcasm cannot be identified using a single modality. Humans employ the conflict between a statement and factual information as a cue to detect sarcasm, and these cues can stem from different modalities. Then, a unified label for a multimodal instance may be not suitable for the associated text or image. In this work, we introduce SarcNet, a multilingual and multimodal sarcasm detection dataset in English and Chinese, consisting of 3,335 image-text pair samples. We provide annotations for sarcasm in visual, textual, and multimodal data, respectively, resulting in over 10,000 labeled instances. The separated annotation schema for unimodal and multimodal data facilitates a more accurate and reasonable assessment of unimodal and multimodal models. Tan Yue, Xuzhao Shi, Rui Mao 0010, Zonghai Hu, Erik Cambria |
LREC/COLING | 3 |
| 2024 | Understanding Public Perception Towards Weather Disasters Through the Lens of Metaphor
Rui Mao 0010, Qika Lin, Qiawen Liu, Gianmarco Mengaldo, Erik Cambria |
IJCAI | 1 |
| 2024 | A Dynamic Dual-Graph Neural Network for Stock Price Movement PredictionabstractThe prediction of stock price movements is challenging due to the inherently dynamic and complex characteristics of financial markets. A current research gap is the lack of exploration into the complex interrelationships inherent in stock price dynamics, often analyzing predictions in isolation with an implicit presumption that solely the historical data of a given stock influences its future trend. However, stock prices are impacted by a diverse array of driving factors that extend beyond the traditionally examined historical prices, encompassing influences such as inter-stock correlations. In this paper, we present a predictive approach using a dynamic dual-graph neural network. The network combines textual data and quantitative metrics to capture multiple dynamic relationships. Specifically, We have developed a price relationship graph (PRG) and a semantic relationship graph (SRG), which are later integrated using a graph attention neural network. The effectiveness of our neural architecture is validated through extensive testing on two benchmark datasets for stock movement prediction, illustrating its superior performance compared to other graph-based networks for stock market prediction. Kelvin Du, Rui Mao 0010, Frank Z. Xing, Erik Cambria |
IJCNN | 2 |
| 2024 | Multilingual Emotion Recognition: Discovering the Variations of Lexical Semantics between LanguagesabstractThe task of multilingual emotion recognition holds significant importance in cross-cultural communication and data mining. While prior research has concentrated on enhancing classification accuracy using state-of-the-art techniques, it has often overlooked a crucial linguistic aspect—the semantic disparities across different languages. This study aims to address this gap by introducing a novel method to identify lexical semantic variations in diverse languages. The detected semantic variation features are subsequently injected into a multilingual emotion recognition model to enhance its performance within a target language. Notably, existing multilingual pre-trained language models are likely biased toward English word meanings, leading to inaccurate emotion predictions in other languages due to the misinterpretation of semantics. Our proposed semantic variation injection method tackles this limitation, resulting in improved accuracy. These findings contribute to the ongoing development of robust and culturally sensitive emotion recognition systems, offering valuable insights for both the linguistics and computational linguistics communities engaged in multilingual research. Xulang Zhang, Rui Mao 0010, Erik Cambria |
IJCNN | 2 |
| 2024 | Medical Report Generation via Multimodal Spatio-Temporal FusionabstractMedical report generation aims at automating the synthesis of accurate and comprehensive diagnostic reports from radiological images. The task can significantly enhance clinical decision-making and alleviate the workload on radiologists. Existing works normally generate reports from single chest radiographs, although historical examination data also serve as crucial references for radiologists in real-world clinical settings. To address this constraint, we introduce a novel framework that mimics the workflow of radiologists. This framework compares past and present patient images to monitor disease progression and incorporates prior diagnostic reports as references for generating current personalized reports. We tackle the textual diversity challenge in cross-modal tasks by promoting style-agnostic discrete report representation learning and token generation. Furthermore, we propose a novel spatio-temporal fusion method with multi-granularities to fuse textual and visual features by disentangling the differences between current and historical data. We also tackle token generation biases, which arise from long-tail frequency distributions, proposing a novel feature normalization technique. This technique ensures unbiased generation for tokens, whether they are frequent or infrequent, enabling the robustness of report generation for rare diseases. Experimental results on the two public datasets demonstrate that our proposed model outperforms state-of-the-art baselines. Xin Mei, Rui Mao 0010, Xiaoyan Cai, Libin Yang, Erik Cambria |
ACM Multimedia | 2 |
| 2024 | A Wide Evaluation of ChatGPT on Affective Computing TasksabstractWith the rise of foundation models, a new artificial intelligence paradigm has emerged, by simply using general purpose foundation models with prompting to solve problems instead of training a separate machine learning model for each problem. Such models have been shown to have emergent properties of solving problems that they were not initially trained on. The studies for the effectiveness of such models are still quite limited. In this work, we widely study the capabilities of the ChatGPT models, namely GPT-4 and GPT-3.5, on 13 affective computing problems, namely aspect extraction, aspect polarity classification, opinion extraction, sentiment analysis, sentiment intensity ranking, emotions intensity ranking, suicide tendency detection, toxicity detection, well-being assessment, engagement measurement, personality assessment, sarcasm detection, and subjectivity detection. We introduce a framework to evaluate the ChatGPT models on regression-based problems, such as intensity ranking problems, by modelling them as pairwise ranking classification. We compare ChatGPT against more traditional NLP methods, such as end-to-end recurrent neural networks and transformers. The results demonstrate the emergent abilities of the ChatGPT models on a wide range of affective computing problems, where GPT-3.5 and especially GPT-4 have shown strong performance on many problems, particularly the ones related to sentiment, emotions, or toxicity. The ChatGPT models fell short for problems with implicit signals, such as engagement measurement and subjectivity detection. Mostafa M. Amin, Rui Mao 0010, Erik Cambria, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Template-Free Prompting for Few-Shot Named Entity Recognition via Semantic-Enhanced Contrastive LearningabstractPrompt tuning has achieved great success in various sentence-level classification tasks by using elaborated label word mappings and prompt templates. However, for solving token-level classification tasks, e.g., named entity recognition (NER), previous research, which utilizes N-gram traversal for prompting all spans with all possible entity types, is time-consuming. To this end, we propose a novel prompt-based contrastive learning method for few-shot NER without template construction and label word mappings. First, we leverage external knowledge to initialize semantic anchors for each entity type. These anchors are simply appended with input sentence embeddings as template-free prompts (TFPs). Then, the prompts and sentence embeddings are in-context optimized with our proposed semantic-enhanced contrastive loss. Our proposed loss function enables contrastive learning in few-shot scenarios without requiring a significant number of negative samples. Moreover, it effectively addresses the issue of conventional contrastive learning, where negative instances with similar semantics are erroneously pushed apart in natural language processing (NLP)-related tasks. We examine our method in label extension (LE), domain-adaption (DA), and low-resource generalization evaluation tasks with six public datasets and different settings, achieving state-of-the-art (SOTA) results in most cases. Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | SKIER: A Symbolic Knowledge Integrated Model for Conversational Emotion RecognitionabstractEmotion recognition in conversation (ERC) has received increasing attention from the research community. However, the ERC task is challenging, largely due to the complex and unstructured properties of multi-party conversations. Besides, the majority of daily dialogues take place in a specific context or circumstance, which requires rich external knowledge to understand the background of a certain dialogue. In this paper, we address these challenges by explicitly modeling the discourse relations between utterances and incorporating symbolic knowledge into multi-party conversations. We first introduce a dialogue parsing algorithm into ERC and further improve the algorithm through a transfer learning method. Moreover, we leverage different symbolic knowledge graph relations to learn knowledge-enhanced features for the ERC task. Extensive experiments on three benchmarks demonstrate that both dialogue structure graphs and symbolic knowledge are beneficial to the model performance on the task. Additionally, experimental results indicate that the proposed model surpasses baseline models on several indices. Wei Li 0076, Rui Mao 0010, Erik Cambria |
AAAI | 3 |
| 2023 | TECHS: Temporal Logical Graph Networks for Explainable Extrapolation ReasoningabstractExtrapolation reasoning on temporal knowledge graphs (TKGs) aims to forecast future facts based on past counterparts.There are two main challenges: (1) incorporating the complex information, including structural dependencies, temporal dynamics, and hidden logical rules;(2) implementing differentiable logical rule learning and reasoning for explainability.To this end, we propose an explainable extrapolation reasoning framework TEemporal logiCal grapH networkS (TECHS), which mainly contains a temporal graph encoder and a logical decoder.The former employs a graph convolutional network with temporal encoding and heterogeneous attention to embed topological structures and temporal dynamics.The latter integrates propositional reasoning and first-order reasoning by introducing a reasoning graph that iteratively expands to find the answer.A forward message-passing mechanism is also proposed to update node representations, and their propositional and first-order attention scores.Experimental results demonstrate that it outperforms state-of-the-art baselines. Qika Lin, Jun Liu 0002, Rui Mao 0010, Fangzhi Xu, Erik Cambria |
ACL (1) | 3 |
| 2023 | Finding the Pillars of Strength for Multi-Head AttentionabstractRecent studies have revealed some issues of Multi-Head Attention (MHA), e.g., redundancy and over-parameterization.Specifically, the heads of MHA were originally designed to attend to information from different representation subspaces, whereas prior studies found that some attention heads likely learn similar features and can be pruned without harming performance.Inspired by the minimumredundancy feature selection, we assume that focusing on the most representative and distinctive features with minimum resources can mitigate the above issues and lead to more effective and efficient MHAs.In particular, we propose Grouped Head Attention, trained with a self-supervised group constraint that group attention heads, where each group focuses on an essential but distinctive feature subset.We additionally propose a Voting-to-Stay procedure to remove redundant heads, thus achieving a transformer with lighter weights.Moreover, our method achieves significant performance gains on three well-established tasks while considerably compressing parameters. Jinjie Ni, Rui Mao 0010, Zonglin Yang 0001, Han Lei, Erik Cambria |
ACL (1) | 2 |
| 2023 | PAED: Zero-Shot Persona Attribute Extraction in DialoguesabstractPersona attribute extraction is critical for personalized human-computer interaction.Dialogue is an important medium that communicates and delivers persona information.Although there is a public dataset for triplet-based persona attribute extraction from conversations, its automatically generated labels present many issues, including unspecific relations and inconsistent annotations.We fix such issues by leveraging more reliable text-label matching criteria to generate high-quality data for persona attribute extraction.We also propose a contrastive learning-and generation-based model with a novel hard negative sampling strategy for generalized zero-shot persona attribute extraction.We benchmark our model with stateof-the-art baselines on our dataset and a public dataset, showing outstanding accuracy gains.Our sampling strategy also exceeds others by a large margin in persona attribute extraction. Wei Li 0076, Rui Mao 0010, Vlad Pandelea, Erik Cambria |
ACL (1) | 3 |
| 2023 | Discovering the Cognition behind Language: Financial Metaphor Analysis with MetaProabstractMetaphors frequently appear in financial news headlines due to their ability to effectively convey complex financial concepts and market trends in a concise and memorable manner. Cognitive scientists have found that metaphors serve as the reflections of human cognition by means of concept mappings. In this work, we aim to analyze the metaphorical expressions and associated cognitive patterns employed by financial analysts in the headlines of financial analysis reports. Such an examination would enhance our comprehension of the cognitive state of financial analysts regarding various financial trends. We employ the latest computational metaphor processing tool, MetaPro to achieve this target by mining metaphors and cognitive patterns from 1,407,328 financial analyst report headlines, spanning the period from 14 February 2009 to 11 June 2020. We analyze the mined concept mappings by different time periods, and market movements, and deliver novel findings in these two dimensions. Rui Mao 0010, Kelvin Du, Erik Cambria |
ICDM | 1 |
| 2023 | A Multi-task Learning Model for Gold-two-mention Co-reference ResolutionabstractThe task of resolving repeated objects in natural languages is known as co-reference resolution. It is an important part of modern natural language processing and semantic cognition as these implicit relationships are particularly difficult in natural language understanding in downstream tasks. Mention identification and mention linking are the two sub-tasks in the general co-reference resolution research community. Gold-two-mention style co-reference resolution is a special type of co-reference resolution that focuses on linking the ambiguous pronoun to one of the two candidate antecedents. In this paper, we proposed a joint learning model that learns mention identification and mention linking tasks together, because we find that the learning of mention identification can provide supportive dependent information for the learning of mention linking. As far as we know, we propose the first model that introduces a multi-task learning framework to the gold-two-mention co-reference resolution task. We find that our proposed model outperforms state-of-the-art baselines and a single-task learning model on three gold-two-mention co-reference resolution datasets. By comparing the errors made by either the single-task learning model or the multi-task learning model, our error analysis also yields interesting findings about in which way our multi-task learning model makes fewer resolution errors. Ruicheng Liu, Guanyi Chen, Rui Mao 0010, Erik Cambria |
IJCNN | 3 |
| 2023 | Virtual prompt pre-training for prototype-based few-shot relation extraction
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria |
Expert Syst. Appl. | 3 |
| 2023 | Semantic matching in machine reading comprehension: An empirical study
Qian Liu 0012, Rui Mao 0010, Xiubo Geng, Erik Cambria |
Inf. Process. Manag. | 2 |
| 2023 | Meta-Based Self-Training and Re-Weighting for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) means to identify fine-grained aspects, opinions, and sentiment polarities. Recent ABSA research focuses on utilizing multi-task learning (MTL) to achieve less computational costs and better performance. However, there are certain limits in MTL-based ABSA. For example, unbalanced labels and sub-task learning difficulties may result in the biases that some labels and sub-tasks are overfitting, while the others are underfitting. To address these issues, inspired by neuro-symbolic learning systems, we propose a meta-based self-training method with a meta-weighter (MSM). We believe that a generalizable model can be achieved by appropriate symbolic representation selection (in-domain knowledge) and effective learning control (regulation) in a neural system. Thus, MSM trains a teacher model to generate in-domain knowledge (e.g., unlabeled data selection and pseudo-label generation), where the generated pseudo-labels are used by a student model for supervised learning. Then, the meta-weighter of MSM is jointly trained with the student model to provide each instance with sub-task-specific weights to coordinate their convergence rates, balancing class labels, and alleviating noise impacts introduced from self-training. The following experiments indicate that MSM can utilize 50% labeled data to achieve comparable results to state-of-arts models in ABSA and outperform them with all labeled data. Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | The Biases of Pre-Trained Language Models: An Empirical Study on Prompt-Based Sentiment Analysis and Emotion DetectionabstractThanks to the breakthrough of large-scale pre-trained language model (PLM) technology, prompt-based classification tasks, e.g., sentiment analysis and emotion detection, have raised increasing attention. Such tasks are formalized as masked language prediction tasks which are in line with the pre-training objects of most language models. Thus, one can use a PLM to infer the masked words in a downstream task, then obtaining label predictions with manually defined label-word mapping templates. Prompt-based affective computing takes the advantages of both neural network modeling and explainable symbolic representations. However, there still remain many unclear issues related to the mechanisms of PLMs and prompt-based classification. We conduct a systematic empirical study on prompt-based sentiment analysis and emotion detection to study the biases of PLMs towards affective computing. We find that PLMs are biased in sentiment analysis and emotion detection tasks with respect to the number of label classes, emotional label-word selections, prompt templates and positions, and the word forms of emotion lexicons. Rui Mao 0010, Qian Liu 0012, Kai He 0001, Wei Li 0076, Erik Cambria |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Explainable Metaphor Identification Inspired by Conceptual Metaphor TheoryabstractMetaphor is not only a linguistic phenomenon but also reflects the concept projection between source and target domains in human cognition. Previous sequence tagging-based metaphor identification methods could not model the concept projection, resulting in a limitation that the outputs of these models are unexplainable in the predictions of the metaphoricity labels. In this work, we propose the first explainable metaphor identification model, inspired by Conceptual Metaphor Theory. The model is based on statistic learning, a lexical resource, and a novel reward mechanism. Our model can identify the metaphoricity on the word-pair level, and explain the predicted metaphoricity labels via learned concept mappings. The use of the reward mechanism allows the model to learn the optimal concept mappings without knowing their true labels. Our method is also applicable for the concepts that are out of training domains by using the lexical resource. The automatically generated concept mappings demonstrate the implicit human thoughts in metaphoric expressions. Our experiments show the effectiveness of the proposed model in metaphor identification, and concept mapping tasks, respectively. Mengshi Ge, Rui Mao 0010, Erik Cambria |
AAAI | 2 |
| 2022 | Hierarchical Attention Network for Explainable Depression Detection on Twitter Aided by Metaphor Concept MappingsabstractAutomatic depression detection on Twitter can help individuals privately and conveniently understand their mental health status in the early stages before seeing mental health professionals. Most existing black-box-like deep learning methods for depression detection largely focused on improving classification performance. However, explaining model decisions is imperative in health research because decision-making can often be high-stakes and life-and-death. Reliable automatic diagnosis of mental health problems including depression should be supported by credible explanations justifying models’ predictions. In this work, we propose a novel explainable model for depression detection on Twitter. It comprises a novel encoder combining hierarchical attention mechanisms and feed-forward neural networks. To support psycholinguistic studies, our model leverages metaphorical concept mappings as input. Thus, it not only detects depressed individuals, but also identifies features of such users’ tweets and associated metaphor concept mappings. Sooji Han, Rui Mao 0010, Erik Cambria |
COLING | 2 |
| 2022 | COPNER: Contrastive Learning with Prompt Guiding for Few-shot Named Entity RecognitionabstractDistance metric learning has become a popular solution for few-shot Named Entity Recognition (NER). The typical setup aims to learn a similarity metric for measuring the semantic similarity between test samples and referents, where each referent represents an entity class. The effect of this setup may, however, be compromised for two reasons. First, there is typically a limited optimization exerted on the representations of entity tokens after initing by pre-trained language models. Second, the referents may be far from representing corresponding entity classes due to the label scarcity in the few-shot setting. To address these challenges, we propose a novel approach named COntrastive learning with Prompt guiding for few-shot NER (COPNER). We introduce a novel prompt composed of class-specific words to COPNER to serve as 1) supervision signals for conducting contrastive learning to optimize token representations; 2) metric referents for distance-metric inference on test samples. Experimental results demonstrate that COPNER outperforms state-of-the-art models with a significant margin in most cases. Moreover, COPNER shows great potential in the zero-shot setting. Kai He 0001, Xianli Zhang, Tieliang Gong, Rui Mao 0010, Chen Li 0011 |
COLING | 6 |
| 2022 | Saving Earth One Tweet at a Time through the Lens of Artificial IntelligenceabstractThe impacts of climate change and global warming have become more visible globally because of the increasing frequency of extreme weather events, abnormal heatwaves, and other climate crises. Besides the traditional survey method, it is beneficial to automatically distillate climate change opinions from social platforms to measure public reactions quickly. We investigate how to organize climate change opinions on Twitter into meaningful categories to support perspective summarizing tasks. We find that merely using the available taxonomy for this task is ineffective; hence we must consider the entire text content. We recommend five high-level categories (Root cause, Impact, Mitigation, Politics or Policy, Others) and assemble ClimateTweets, a dataset with category and polarity labels. In addition, we construct category classification and polarity detection tasks with a range of opinion mining baselines. The experimental results show that both tasks are challenging for existing models. We release the ClimateTweets dataset to facilitate investigation in public opinion mining using text content and artificial intelligent methods. We hope this study could pave the way for future studies in the climate change domain. Cuc Duong, Qian Liu 0012, Rui Mao 0010, Erik Cambria |
IJCNN | 3 |
| 2022 | JCBIE: a joint continual learning neural network for biomedical information extractionabstractExtracting knowledge from heterogeneous data sources is fundamental for the construction of structured biomedical knowledge graphs (BKGs), where entities and relations are represented as nodes and edges in the graphs, respectively. Previous biomedical knowledge extraction methods simply considered limited entity types and relations by using a task-specific training set, which is insufficient for large-scale BKGs development and downstream task applications in different scenarios. To alleviate this issue, we propose a joint continual learning biomedical information extraction (JCBIE) network to extract entities and relations from different biomedical information datasets. By empirically studying different joint learning and continual learning strategies, the proposed JCBIE can learn and expand different types of entities and relations from different datasets. JCBIE uses two separated encoders in joint-feature extraction, hence can effectively avoid the feature confusion problem comparing with using one hard-parameter sharing encoder. Specifically, it allows us to adopt entity augmented inputs to establish the interaction between named entity recognition and relation extraction. Finally, a novel evaluation mechanism is proposed for measuring cross-corpus generalization errors, which was ignored by traditional evaluation methods. Our empirical studies show that JCBIE achieves promising performance when continual learning strategy is adopted with multiple corpora. Kai He 0001, Rui Mao 0010, Tieliang Gong, Erik Cambria, Chen Li 0011 |
BMC Bioinform. | 2 |
| 2021 | Bridging Towers of Multi-task Learning with a Gating Mechanism for Aspect-based Sentiment Analysis and Sequential Metaphor IdentificationabstractMulti-task learning (MTL) has been widely applied in Natural Language Processing. A major task and its associated auxiliary tasks share the same encoder; hence, an MTL encoder can learn the sharing abstract information between the major and auxiliary tasks. Task-specific towers are then employed upon the sharing encoder to learn task-specific information. Previous works demonstrated that exchanging information between task-specific towers yielded extra gains. This is known as soft-parameter sharing MTL. In this paper, we propose a novel gating mechanism for the bridging of MTL towers. Our method is evaluated based on aspect-based sentiment analysis and sequential metaphor identification tasks. The experiments demonstrate that our method can yield better performance than the baselines on both tasks. Based on the same Transformer backbone, we compare our gating mechanism with other information transformation mechanisms, e.g., cross-stitch, attention and vanilla gating. The experiments show that our method also surpasses these baselines. Rui Mao 0010, Xiao Li 0041 |
AAAI | 1 |
| 2019 | End-to-End Sequential Metaphor Identification Inspired by Linguistic TheoriesabstractEnd-to-end training with Deep Neural Networks (DNN) is a currently popular method for metaphor identification.However, standard sequence tagging models do not explicitly take advantage of linguistic theories of metaphor identification.We experiment with two DNN models which are inspired by two human metaphor identification procedures.By testing on three public datasets, we find that our models achieve state-of-the-art performance in end-to-end metaphor identification. Rui Mao 0010, Chenghua Lin 0002, Frank Guerin |
ACL (1) | 1 |
| 2019 | A Stable Variational Autoencoder for Text ModellingabstractVariational Autoencoder (VAE) is a powerful method for learning representations of highdimensional data.However, VAEs can suffer from an issue known as latent variable collapse (or KL loss vanishing), where the posterior collapses to the prior and the model will ignore the latent codes in generative tasks.Such an issue is particularly prevalent when employing VAE-RNN architectures for text modelling (Bowman et al., 2016).In this paper, we present a simple architecture called holistic regularisation VAE (HR-VAE), which can effectively avoid latent variable collapse.Compared to existing VAE-RNN architectures, we show that our model can achieve much more stable training process and can generate text with significantly better quality. Ruizhe Li 0001, Xiao Li 0041, Chenghua Lin 0002, Matthew Collinson, Rui Mao 0010 |
INLG | 5 |
| 2018 | Word Embedding and WordNet Based Metaphor Identification and InterpretationabstractMetaphoric expressions are widespread in natural language, posing a significant challenge for various natural language processing tasks such as Machine Translation.Current word embedding based metaphor identification models cannot identify the exact metaphorical words within a sentence.In this paper, we propose an unsupervised learning method that identifies and interprets metaphors at word-level without any preprocessing, outperforming strong baselines in the metaphor identification task.Our model extends to interpret the identified metaphors, paraphrasing them into their literal counterparts, so that they can be better translated by machines.We evaluated this with two popular translation systems for English to Chinese, showing that our model improved the systems significantly. Rui Mao 0010, Chenghua Lin 0002, Frank Guerin |
ACL (1) | 1 |