EDBT 2026 Demo / reviewers in the wild / expert
Jianfei Yu
dblp:30/9823
· DBLP profile ↗
48ranked-venue papers
14as first author
27since 2021 · last 2026
0000-0001-8380-0609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 14 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive QueriesabstractRetrieval-augmented generation (RAG) augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate multimodal LLMs (MLLMs) under multimodal RAG settings, they predominantly retrieve from textual corpora and do not explicitly assess how models exploit visual evidence during answer generation. Consequently, there still lacks benchmark that cleanly isolates and measures the contribution of retrieved images in a visual knowledge-intensive RAG pipeline. We introduce Visual-RAG, a question-answering benchmark that targets visually-grounded, knowledge-intensive questions in a visual evidence-centric manner. Unlike prior work, Visual-RAG requires text-to-image retrieval and the integration of retrieved clue images whose pixel content explicitly encodes the visual knowledge necessary for answer generation. With Visual-RAG, we evaluate five open-source and three proprietary MLLMs and find that current systems still substantially underutilize the visual information available in retrieved images. Despite clear opportunities for multimodal evidence integration, state-of-the-art models struggle to extract and exploit fine-grained visual knowledge, and text-to-image retrieval itself remains challenging even under constrained entity-level corpora. These results underscore the need for improved visual retrieval, grounding, and attribution in multimodal RAG. Visual-RAG is publicly available at: github.com/visual-rag/visual-rag Yin Wu 0001, Quanyu Long, Jing Li 0034, Jianfei Yu, Wenya Wang 0001 |
SIGIR | 4 |
| 2026 | Dual-Channel Retrieval-Augmented In-Context Learning for Comparative Opinion Mining
Wenjing Gui, Zinong Yang, Jianfei Yu |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based SegmentationabstractGrounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging attributes: 1) The tenuous correlation between images and text on social media contributes to a notable proportion of named entities being ungroundable. 2) There exists a distinction between coarse-grained noun phrases used in similar tasks (e.g., phrase localization) and fine-grained named entities. In this paper, we propose RiVEG, a unified framework that reformulates GMNERinto a joint modeling paradigm spanning MNER,VE, and VGperspectives by leveraging large language models (LLMs) as connecting bridges. This reformulation brings two benefits: 1) It enables us to optimize the MNER module for optimal MNER performance and eliminates the need to pre-extract region features using object detection methods, thus naturally addressing the two major limitations of existing GMNER methods. 2) The introduction of Entity Expansion Expression module and Visual Entailment (VE) module unifies Visual Grounding (VG) and Entity Grounding (EG). This endows the proposed framework with unlimited data and model scalability. Furthermore, to address the potential ambiguity stemming from the coarse-grained bounding box output in GMNER, we further construct the new Segmented Multimodal Named Entity Recognition (SMNER) task and corresponding Twitter-SMNER dataset aimed at generating fine-grained segmentation masks, and experimentally demonstrate the feasibility and effectiveness of using box prompt-based Segment Anything Model (SAM) to empower any GMNER model with the ability to accomplish the SMNER task. Extensive experiments demonstrate that RiVEG significantly outperforms SoTA methods on four datasets across the MNER, GMNER, and SMNER tasks. Datasets and Code will be released athttps://github.com/JinYuanLi0012/RiVEG. Jianfei Yu, Di Sun 0001, Gang Pan 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | Language Models over Large-Scale Knowledge Base: on Capacity, Flexibility and Reasoning for New FactsabstractAdvancements in language models (LMs) have sparked interest in exploring their potential as knowledge bases (KBs) due to their high capability for storing huge amounts of factual knowledge and semantic understanding. However, existing studies face challenges in quantifying the extent of large-scale knowledge packed into LMs and lack systematic studies on LMs’ structured reasoning capabilities over the infused knowledge. Addressing these gaps, our research investigates whether LMs can effectively act as large-scale KBs after training over an expansive set of world knowledge triplets via addressing the following three crucial questions: (1) How do LMs of different sizes perform at storing world knowledge of different frequencies in a large-scale KB? (2) How flexible are these LMs in recalling the stored knowledge when prompted with natural language queries? (3) After training on the abundant world knowledge, can LMs additionally gain the ability to reason over such information to infer new facts? Our findings indicate that while medium-scaled LMs hold promise as world knowledge bases capable of storing and responding with flexibility, enhancements in their reasoning capabilities are necessary to fully realize their potential. Qiyuan He, Yizhong Wang, Jianfei Yu |
COLING | 3 |
| 2025 | Global Question-Aware Multimodal Retrieval-Augmented Generation for Multimedia Multi-Hop Question AnsweringabstractMultimedia Multi-Hop Question Answering (MMQA) is a complex task that requires models to reason over and integrate information from both visual (e.g., images) and textual (e.g., documents) modalities to answer questions that cannot be resolved in a single step. Existing methods for MMQA suffer from two common limitations: (1) insufficient cross-modal information fusion, which restricts interaction between different modalities; and (2) weak global understanding of multi-hop questions, making them vulnerable to distractions from intermediate steps. To address these challenges, we introduce Global question-aware Multimodal Retrieval-Augmented Generation (GMRAG), a framework designed to enhance cross-modal reasoning and improve retrieval precision for multi-hop multimodal questions. It brings two core innovations: (1) restructuring the training data into a global question–aware, evidence-centric retrieval dataset, enabling the retriever to perform richer cross-modal reasoning on multi-hop multimodal questions and to identify globally relevant information more effectively; and (2) employing cross-modal contrastive learning to fine-tune a joint image–text encoder, achieving stronger alignment between holistic multimodal questions and evidence. The retrieved evidence is then directly fed into a Multimodal Large Language Models (MLLMs) to generate the final answer to the original question, eliminating dependence on potentially flawed intermediate answers and thus mitigating error propagation. Experimental results on two public MMQA datasets show that GMRAG consistently outperforms existing RAG methods across various MLLMs. Zhixiao Shen, Jianfei Yu, Wenya Wang 0001 |
MMAsia | 2 |
| 2025 | An entity-Aware MLLM with conditional image generation for multimodal entity-Category-Sentiment triple extraction
Li Yang 0025, Yaming Zhang, Jin-Cheon Na, Jianfei Yu |
Knowl. Based Syst. | 4 |
| 2025 | From Extraction to Generation: Multimodal Emotion-Cause Pair Generation in ConversationsabstractAs an important task in emotion analysis, Multimodal Emotion-Cause Pair Extraction in conversations (MECPE) aims to extract all the emotion-cause utterance pairs from a conversation. However, there are two shortcomings in the MECPE task: 1) it ignores emotion utterances whose causes cannot be located in the conversation but require contextualized inference; 2) it fails to locate the exact causes that occur in vision or audio modalities beyond text. To address these issues, in this paper, we introduce a new task named Multimodal Emotion-Cause Pair Generation in Conversations (MECPG), which aims to identify the emotion utterances with their emotion categories and generate their corresponding causes in a conversation. To tackle the MECPG task, we construct a dataset based on a benchmark corpus for MECPE. We further propose a generative framework named MONICA, which jointly performs emotion recognition and emotion cause generation with a sequence-to-sequence model. Experiments on our annotated dataset show the superiority of MONICA over several competitive systems. Our dataset and source codes will be publicly released. Heqing Ma, Jianfei Yu, Fanfan Wang, Hanyu Cao |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | A Joint Coreference-Aware Approach to Document-Level Target Sentiment AnalysisabstractMost existing work on aspect-based sentiment analysis (ABSA) focuses on the sentence level, while research at the document level has not received enough attention.Compared to sentence-level ABSA, the document-level ABSA is not only more practical but also requires holistic document-level understanding capabilities such as coreference resolution.To investigate the impact of coreference information on document-level ABSA, we conduct a three-stage research for the document-level target sentiment analysis (DTSA) task: 1) exploring the effectiveness of coreference information for the DTSA task; 2) reducing the reliance on manually annotated coreference information; 3) alleviating the evaluation bias caused by missing the coreference information of opinion targets.Specifically, we first manually annotate the coreferential opinion targets and propose a multi-task learning framework to model the DTSA task and the coreference resolution task jointly.Then we annotate the coreference information with ChatGPT for joint training.Finally, to address the issue of missing coreference targets, we modify the metric from strict matching to a loose matching method based on the clusters of targets.The experimental results demonstrate our framework's effectiveness and reflect the feasibility of using ChatGPT-annotated coreferential entities and the applicability of the modified metric.Our source code is publicly released at https://github.com/NUSTM/DTSA- Coref. Hongjie Cai, Heqing Ma, Jianfei Yu |
ACL (1) | 3 |
| 2024 | Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity RecognitionabstractAs an important task in multimodal information extraction, Multimodal Named Entity Recognition (MNER) has recently attracted considerable attention. One key challenge of MNER lies in the lack of sufficient fine-grained annotated data, especially in low-resource scenarios. Although data augmentation is a widely used technique to tackle the above issue, it is challenging to simultaneously generate synthetic text-image pairs and their corresponding high-quality entity annotations. In this work, we propose a novel Generative Multimodal Data Augmentation (GMDA) framework for MNER, which contains two stages: Multimodal Text Generation and Multimodal Image Generation. Specifically, we first transform each annotated sentence into a linearized labeled sequence, and then train a Label-aware Multimodal Large Language Model (LMLLM) to generate the labeled sequence based on a label-aware prompt and its associated image. We further employ a Stable Diffusion model to generate the synthetic images that are semantically related to these sentences. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed GMDA framework, which consistently boosts the performance of several competitive methods for two subtasks of MNER in both full-supervision and low-resource settings. The low-resource dataset and source code are released at https://github.com/NUSTM/GMDA. Jianfei Yu, Wenya Wang 0001, Li Yang 0025 |
ACM Multimedia | 2 |
| 2024 | Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in ConversationsabstractEmotion cause analysis has attracted increasing attention in recent years. However, the integration of multimodal information with emotion causes remains underexplored. Existing studies merely extract utterances from conversations as cause evidence, which is too coarse-grained to locate the exact causes from other modalities, especially those that may be reflected only in a specific video frame of an utterance. To address these limitations, we introduce a new task named Multimodal Emotion Cause Generation in Conversations (MECGC), which aims to generate an abstractive summary clearly and intuitively describing the causes that trigger the given emotion based on the multimodal context of conversations. We accordingly construct a dataset named ECGF that contains 1,374 conversations and 7,690 emotion instances from TV series. We further develop a generative framework that first generates emotion-cause aware video captions (Observe) and then facilitates the generation of emotion causes (Generate). The captioning model is trained with examples synthesized by a Multimodal Large Language Model (MLLM). Experimental results demonstrate the effectiveness of our framework and the significance of multimodal information for emotion cause analysis. Our dataset and source codes are available at https://github.com/NUSTM/MECGC. Fanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu |
ACM Multimedia | 4 |
| 2024 | A Unimodal Valence-Arousal Driven Contrastive Learning Framework for Multimodal Multi-Label Emotion Recognition
Wenjie Zheng 0005, Jianfei Yu |
ACM Multimedia | 2 |
| 2024 | An empirical study of Multimodal Entity-Based Sentiment Analysis with ChatGPT: Improving in-context learning via entity-aware contrastive learning
Li Yang 0025, Zengzhi Wang, Jin-Cheon Na, Jianfei Yu |
Inf. Process. Manag. | 5 |
| 2024 | Unified ABSA via Annotation-Decoupled Multi-Task Instruction TuningabstractAspect-Based Sentiment Analysis (ABSA) aims to provide fine-grained aspect-level sentiment information. Different ABSA tasks are designed for different real-world applications. However, application scenarios of ABSA tasks are often diverse, typically requiring training separate systems for each task on the task-specific labeled data and making separate predictions. Secondly, different tasks often contain shared sentiment elements. Training task-specific models either fail to exploit the shared knowledge among multiple ABSA tasks effectively or neglect the complementarity between tasks. Thirdly, despite the existence of the compound ABSA task such as quadruple extraction and triple extraction, it is not easy to obtain satisfactory performance due to the coupling of multiple elements. To tackle these issues, we present UNIFIEDABSA, a general-purpose ABSA framework based on multi-task instruction tuning, aiming at “one-model-for-all-tasks”. We also introduce a new annotation-decoupled multi-task learning mechanism that only depends on annotation on the compound task rather than all tasks. This mechanism not only fully utilizes the existing annotations from the compound task, but also alleviates the complicated coupling relationship among multiple elements, making the learning more effective. Extensive experiments show that UNIFIEDABSA, can consistently outperform dedicated models in fully-supervised and low-resource settings for almost all 11 ABSA tasks. We also conduct further experiments to demonstrate the general applicability of our framework. Zengzhi Wang, Jianfei Yu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Grounded Multimodal Named Entity Recognition on Social MediaabstractIn recent years, Multimodal Named Entity Recognition (MNER) on social media has attracted considerable attention.However, existing MNER studies only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction and insufficient for entity disambiguation.To solve these issues, in this work, we introduce a Grounded Multimodal Named Entity Recognition (GM-NER) task.Given a text-image social post, GMNER aims to identify the named entities in text, their entity types, and their bounding box groundings in image (i.e., visual regions).To tackle the GMNER task, we construct a Twitter dataset based on two existing MNER datasets.Moreover, we extend four well-known MNER methods to establish a number of baseline systems and further propose a Hierarchical Index generation framework named H-Index, which generates the entity-type-region triples in a hierarchical manner with a sequence-tosequence model.Experiment results on our annotated dataset demonstrate the superiority of our H-Index framework over baseline systems on the GMNER task.Our dataset annotation and source code are publicly released at https://github.com/NUSTM/GMNER. Jianfei Yu, Jieming Wang |
ACL (1) | 1 |
| 2023 | Cross-Domain Data Augmentation with Domain-Adaptive Language Modeling for Aspect-Based Sentiment AnalysisabstractCross-domain Aspect-Based Sentiment Analysis (ABSA) aims to leverage the useful knowledge from a source domain to identify aspectsentiment pairs in sentences from a target domain.To tackle the task, several recent works explore a new unsupervised domain adaptation framework, i.e., Cross-Domain Data Augmentation (CDDA), aiming to directly generate much labeled target-domain data based on the labeled source-domain data.However, these CDDA methods still suffer from several issues: 1) preserving many source-specific attributes such as syntactic structures; 2) lack of fluency and coherence; 3) limiting the diversity of generated data.To address these issues, we propose a new cross-domain Data Augmentation approach based on Domain-Adaptive Language Modeling named DA 2 LM, which contains three stages: 1) assigning pseudo labels to unlabeled target-domain data; 2) unifying the process of token generation and labeling with a Domain-Adaptive Language Model (DALM) to learn the shared context and annotation across domains; 3) using the trained DALM to generate labeled target-domain data.Experiments show that DA 2 LM consistently outperforms previous feature adaptation and CDDA methods on both ABSA and Aspect Extraction tasks. Jianfei Yu, Qiankun Zhao |
ACL (1) | 1 |
| 2023 | A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party ConversationsabstractMultimodal Emotion Recognition in Multiparty Conversations (MERMC) has recently attracted considerable attention.Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information.Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC.However, given an utterance, the face sequence extracted by previous methods may contain multiple people's faces, which will inevitably introduce noise to the emotion prediction of the real speaker.To tackle this issue, we propose a two-stage framework named Facial expressionaware Multimodal Multi-Task learning (Fa-cialMMT).Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching.With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning.Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. Wenjie Zheng 0005, Jianfei Yu, Shijin Wang 0001 |
ACL (1) | 2 |
| 2023 | Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkabstractMultimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a pair of text and image. However, most previous MNER works focus on extracting entities in the form of text but failing to ground text symbols to their corresponding visual objects. Moreover, existing MNER studies primarily classify entities into four coarse-grained entity types, which are often insufficient to map them to their real-world referents. To solve these limitations, we introduce a task named Fine-grained Multimodal Named Entity Recognition and Grounding (FMNERG) in this paper, which aims to simultaneously extract named entities in text, their fine-grained entity types, and their grounded visual objects in image. Moreover, we construct a Twitter dataset for the FMNERG task, and further propose a T5-based multImodal GEneration fRamework (TIGER), which formulates FMNERG as a generation problem by converting all the entity-type-object triples into a target sequence and adapts a pre-trained sequence-to-sequence model T5 to directly generate the target sequence from an image-text input pair. Experimental results demonstrate that TIGER performs significantly better than a number of baseline systems on the annotated Twitter dataset. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/FMNERG. Jieming Wang, Jianfei Yu, Li Yang 0025 |
ACM Multimedia | 3 |
| 2023 | Generating paraphrase sentences for multimodal entity-category-sentiment triple extraction
Li Yang 0025, Jieming Wang, Jin-Cheon Na, Jianfei Yu |
Knowl. Based Syst. | 4 |
| 2023 | Multimodal Emotion-Cause Pair Extraction in ConversationsabstractConversation is an important form of human communication and contains a large number of emotions. It is interesting to discover emotions and their causes in conversations. Conversation in its natural form is multimodal. Many studies have been carried out on multimodal emotion recognition in conversations, yet there is still a lack of work on multimodal emotion cause analysis. In this article, we introduce a new task named Multimodal Emotion-Cause Pair Extraction in Conversations, aiming to jointly extract emotions and the corresponding causes from conversations reflected in multiple modalities (i.e., text, audio and video). We accordingly construct a multimodal conversational emotion cause dataset, Emotion-Cause-in-Friends, which contains 9,794 multimodal emotion-cause pairs among 13,619 utterances in theFriendssitcom. We benchmark the task by establishing two baseline systems including a heuristic approach considering inherent patterns in the location of causes and emotions and a deep learning approach that incorporates multimodal features for emotion-cause pair extraction, and conduct the human performance test for comparison. Furthermore, we investigate the effect of multimodal information, explore the potential of incorporating commonsense knowledge, and perform the task under both Static and Real-time settings. Fanfan Wang, Zixiang Ding, Jianfei Yu |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Hierarchical Interactive Multimodal Transformer for Aspect-Based Multimodal Sentiment AnalysisabstractAspect-based multimodal sentiment analysis (ABMSA) aims to determine the sentiment polarities of each aspect or entity mentioned in a multimodal post or review. Previous studies to ABMSA can be summarized into two subtasks: aspect-term based multimodal sentiment classification (ATMSC) and aspect-category based multimodal sentiment classification (ACMSC). However, these existing studies have three shortcomings: (1) ignoring the object-level semantics in images; (2) primarily focusing on aspect-text and aspect-image interactions; (3) failing to consider the semantic gap between text and image representations. To tackle these issues, we propose a general Hierarchical Interactive Multimodal Transformer (HIMT) model for ABMSA. Specifically, we extract salient features with semantic concepts from images via an object detection method, and then propose a hierarchical interaction module to first model the aspect-text and aspect-image interactions, followed by capturing the text-image interactions. Moreover, an auxiliary reconstruction module is devised to largely eliminate the semantic gap between text and image representations. Experimental results show that our HIMT model significantly outperforms state-of-the-art methods on two benchmarks for ATMSC and one benchmark for ACMSC. Jianfei Yu |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisabstractAs an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years.However, previous approaches either (i) use separately pre-trained visual and textual models, which ignore the crossmodal alignment or (ii) use vision-language models pre-trained with general pre-training tasks, which are inadequate to identify finegrained aspects, opinions, and their alignments across modalities.To tackle these limitations, we propose a task-specific Vision-Language Pre-training framework for MABSA (VLP-MABSA), which is a unified multimodal encoder-decoder architecture for all the pretraining and downstream tasks.We further design three types of task-specific pre-training tasks from the language, vision, and multimodal modalities, respectively.Experimental results show that our approach generally outperforms the state-of-the-art approaches on three MABSA subtasks.Further analysis demonstrates the effectiveness of each pretraining task.The source code is publicly released at https://github.com/NUSTM/ VLP-MABSA. Yan Ling, Jianfei Yu |
ACL (1) | 2 |
| 2022 | Targeted Multimodal Sentiment Classification based on Coarse-to-Fine Grained Image-Target MatchingabstractTargeted Multimodal Sentiment Classification (TMSC) aims to identify the sentiment polarities over each target mentioned in a pair of sentence and image. Existing methods to TMSC failed to explicitly capture both coarse-grained and fine-grained image-target matching, including 1) the relevance between the image and the target and 2) the alignment between visual objects and the target. To tackle this issue, we propose a new multi-task learning architecture named coarse-to-fine grained Image-Target Matching network (ITM), which jointly performs image-target relevance classification, object-target alignment, and targeted sentiment classification. We further construct an Image-Target Matching dataset by manually annotating the image-target relevance and the visual object aligned with the input target. Experiments on two benchmark TMSC datasets show that our model consistently outperforms the baselines, achieves state-of-the-art results, and presents interpretable visualizations. Jianfei Yu, Jieming Wang |
IJCAI | 1 |
| 2022 | Generative Cross-Domain Data Augmentation for Aspect and Opinion Co-ExtractionabstractAs a fundamental task in opinion mining, aspect and opinion co-extraction aims to identify the aspect terms and opinion terms in reviews.However, due to the lack of fine-grained annotated resources, it is hard to train a robust model for many domains.To alleviate this issue, unsupervised domain adaptation is proposed to transfer knowledge from a labeled source domain to an unlabeled target domain.In this paper, we propose a new Generative Cross-Domain Data Augmentation framework for unsupervised domain adaptation.The proposed framework is aimed to generate targetdomain data with fine-grained annotation by exploiting the labeled data in the source domain.Specifically, we remove the domain-specific segments in a source-domain labeled sentence, and then use this as input to a pre-trained sequence-to-sequence model BART to simultaneously generate a target-domain sentence and predict the corresponding label for each word.Experimental results on three datasets demonstrate that our approach is more effective than previous domain adaptation methods.The source code is publicly released at https://github.com/NUSTM/GCDDA. Jianfei Yu |
NAACL-HLT | 2 |
| 2022 | Cross-Modal Multitask Transformer for End-to-End Multimodal Aspect-Based Sentiment Analysis
Li Yang 0025, Jin-Cheon Na, Jianfei Yu |
Inf. Process. Manag. | 3 |
| 2021 | Aspect-Category-Opinion-Sentiment Quadruple Extraction with Implicit Aspects and OpinionsabstractHongjie Cai, Rui Xia, Jianfei Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hongjie Cai, Jianfei Yu |
ACL/IJCNLP (1) | 3 |
| 2021 | Reinforced Counterfactual Data Augmentation for Dual Sentiment ClassificationabstractData augmentation and adversarial perturbation approaches have recently achieved promising results in solving the over-fitting problem in many natural language processing (NLP) tasks including sentiment classification.However, existing studies aimed to improve the generalization ability by augmenting the training data with synonymous examples or adding random noises to word embeddings, which cannot address the spurious association problem.In this work, we propose an end-toend reinforcement learning framework, which jointly performs counterfactual data generation and dual sentiment classification.Our approach has three characteristics: 1) the generator automatically generates massive and diverse antonymous sentences; 2) the discriminator contains a original-side sentiment predictor and an antonymous-side sentiment predictor, which jointly evaluate the quality of the generated sample and help the generator iteratively generate higher-quality antonymous samples; 3) the discriminator is directly used as the final sentiment classifier without the need to build an extra one.Extensive experiments show that our approach outperforms strong data augmentation baselines on several benchmark sentiment classification datasets.Further analysis confirms our approach's advantages in generating more diverse training samples and solving the spurious association problem in sentiment classification. Jianfei Yu |
EMNLP (1) | 3 |
| 2021 | Comparative Opinion Quintuple Extraction from Product ReviewsabstractAs an important task in opinion mining, comparative opinion mining aims to identify comparative sentences from product reviews, extract the comparative elements, and obtain the corresponding comparative opinion tuples.However, most previous studies simply regarded comparative tuple extraction as comparative element extraction, which ignored the fact that many comparative sentences may contain multiple comparisons.The comparative opinion tuples defined in these studies also failed to explicitly provide comparative preferences.To address these issues, in this work we first introduce a new Comparative Opinion Quintuple Extraction (COQE) task, to identify comparative sentences from product reviews and extract all comparative opinion quintuples (Subject, Object, Comparative Aspect, Comparative Opinion, Comparative Preference).Secondly, based on the existing comparative opinion mining corpora, we make supplementary annotations and construct three datasets for the COQE task.Finally, we benchmark the COQE task by proposing a new multi-stage neural network approach which significantly outperforms the baseline systems extended from previous comparative opinion mining methods.The datasets and source code are publicly released at https://github. com/NUSTM/COQE. Jianfei Yu |
EMNLP (1) | 3 |
| 2020 | ECPE-2D: Emotion-Cause Pair Extraction based on Joint Two-Dimensional Representation, Interaction and PredictionabstractIn recent years, a new interesting task, called emotion-cause pair extraction (ECPE), has emerged in the area of text emotion analysis.It aims at extracting the potential pairs of emotions and their corresponding causes in a document.To solve this task, the existing research employed a two-step framework, which first extracts individual emotion set and cause set, and then pair the corresponding emotions and causes.However, such a pipeline of two steps contains some inherent flaws: 1) the modeling does not aim at extracting the final emotion-cause pair directly; 2) the errors from the first step will affect the performance of the second step.To address these shortcomings, in this paper we propose a new end-toend approach, called ECPE-Two-Dimensional (ECPE-2D), to represent the emotion-cause pairs by a 2D representation scheme.A 2D transformer module and two variants, windowconstrained and cross-road 2D transformers, are further proposed to model the interactions of different emotion-cause pairs.The 2D representation, interaction, and prediction are integrated into a joint framework.In addition to the advantages of joint modeling, the experimental results on the benchmark emotion cause corpus show that our approach improves the F1 score of the state-of-the-art from 61.28% to 68.89%. Zixiang Ding, Jianfei Yu |
ACL | 3 |
| 2020 | Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerabstractIn this paper, we study Multimodal Named Entity Recognition (MNER) for social media posts.Existing approaches for MNER mainly suffer from two drawbacks: (1) despite generating word-aware visual representations, their word representations are insensitive to the visual context; (2) most of them ignore the bias brought by the visual context.To tackle the first issue, we propose a multimodal interaction module to obtain both image-aware word representations and word-aware visual representations.To alleviate the visual bias, we further propose to leverage purely text-based entity span detection as an auxiliary module, and design a Unified Multimodal Transformer to guide the final predictions with the entity span predictions.Experiments show that our unified approach achieves the new state-of-the-art performance on two benchmark datasets. Jianfei Yu, Jing Jiang 0001 |
ACL | 1 |
| 2020 | Aspect-Category based Sentiment Analysis with Hierarchical Graph Convolutional NetworkabstractMost of the aspect based sentiment analysis research aims at identifying the sentiment polarities toward some explicit aspect terms while ignores implicit aspects in text.To capture both explicit and implicit aspects, we focus on aspect-category based sentiment analysis, which involves joint aspect category detection and category-oriented sentiment classification.However, currently only a few simple studies have focused on this problem.The shortcomings in the way they defined the task make their approaches difficult to effectively learn the inner-relations between categories and the inter-relations between categories and sentiments.In this work, we re-formalize the task as a category-sentiment hierarchy prediction problem, which contains a hierarchy output structure to first identify multiple aspect categories in a review sentence, and jointly predict the sentiment for each of the identified categories.Specifically, we propose a Hierarchical Graph Convolutional Network (Hier-GCN), where a lower-level GCN is to model the inner-relations among multiple categories, and the higher-level GCN is to capture the inter-relations between aspect categories and sentiments.Extensive evaluations demonstrate that our hierarchy output structure is superior over existing ones, and the Hier-GCN model consistently achieves the best results on four benchmarks. Hongjie Cai, Yaofeng Tu, Xiangsheng Zhou, Jianfei Yu |
COLING | 4 |
| 2020 | End-to-End Emotion-Cause Pair Extraction based on Sliding Window Multi-Label LearningabstractEmotion-cause pair extraction (ECPE) is a new task that aims to extract the potential pairs of emotions and their corresponding causes in a document. The existing methods first perform emotion extraction and cause extraction independently, and then perform emotion-cause pairing and filtering. However, the above methods ignore the fact that the cause and the emotion it triggers are inseparable, and the extraction of the cause without specifying the emotion is pathological, which greatly limits the performance of the above methods in the first step. To tackle these shortcomings, we propose two joint frameworks for ECPE: 1) multi-label learning for the extraction of the cause clauses corresponding to the specified emotion clause (CMLL) and 2) multi-label learning for the extraction of the emotion clauses corresponding to the specified cause clause (EMLL). The window of multi-label learning is centered on the specified emotion clause or cause clause and slides as their positions move. Finally, CMLL and EMLL are integrated to obtain the final result. We evaluate our model on a benchmark emotion cause corpus, the results show that our approach achieves the best performance among all compared systems on the ECPE task. Zixiang Ding, Jianfei Yu |
EMNLP (1) | 3 |
| 2020 | Unified Feature and Instance Based Domain Adaptation for Aspect-Based Sentiment AnalysisabstractThe supervised models for aspect-based sentiment analysis (ABSA) rely heavily on labeled data.However, fine-grained labeled data are scarce for the ABSA task.To alleviate the dependence on labeled data, prior works mainly focused on feature-based adaptation, which used the domain-shared knowledge to construct auxiliary tasks or domain adversarial learning to bridge the gap between domains, while ignored the attribute of instance-based adaptation.To resolve this limitation, we propose an end-to-end framework to jointly perform feature and instance based adaptation for the ABSA task in this paper.Based on BERT, we learn domain-invariant feature representations by using part-of-speech features and syntactic dependency relations to construct auxiliary tasks, and jointly perform word-level instance weighting in the framework of sequence labeling.Experiment results on four benchmarks show that the proposed method can achieve significant improvements in comparison with the state-of-the-arts in both tasks of cross-domain End2End ABSA and crossdomain aspect extraction. Chenggong Gong, Jianfei Yu |
EMNLP (1) | 2 |
| 2020 | A State-independent and Time-evolving Network for Early Rumor Detection in Social MediaabstractIn this paper, we study automatic rumor detection for in social media at the event level where an event consists of a sequence of posts organized according to the posting time.It is common that the state of an event is dynamically evolving.However, most of the existing methods to this task ignored this problem, and established a global representation based on all the posts in the event's life cycle.Such coarse-grained methods failed to capture the event's unique features in different states.To address this limitation, we propose a state-independent and time-evolving Network (STN) for rumor detection based on fine-grained event state detection and segmentation.Given an event composed of a sequence of posts, STN first predicts the corresponding sequence of states and segments the event into several state-independent sub-events.For each sub-event, STN independently trains an encoder to learn the feature representation for that sub-event and incrementally fuses the representation of the current sub-event with previous ones for rumor prediction.This framework can more accurately learn the representation of an event in the initial stage and enable early rumor detection.Experiments on two benchmark datasets show that STN can significantly improve the rumor detection accuracy in comparison with some strong baseline systems.We also design a new evaluation metric to measure the performance of early rumor detection, under which STN shows a higher advantage in comparison. Kaizhou Xuan, Jianfei Yu |
EMNLP (1) | 3 |
| 2020 | Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media ConversationsabstractThe prevalent use of social media enables rapid spread of rumors on a massive scale, which leads to the emerging need of automatic rumor verification (RV). A number of previous studies focus on leveraging stance classification to enhance RV with multi-task learning (MTL) methods. However, most of these methods failed to employ pre-trained contextualized embeddings such as BERT, and did not exploit inter-task dependencies by using predicted stance labels to improve the RV task. Therefore, in this paper, to extend BERT to obtain thread representations, we first propose a Hierarchical Transformer, which divides each long thread into shorter subthreads, and employs BERT to separately represent each subthread, followed by a global Transformer layer to encode all the subthreads. We further propose a Coupled Transformer Module to capture the inter-task interactions and a Post-Level Attention layer to use the predicted stance labels for RV, respectively. Experiments on two benchmark datasets show the superiority of our Coupled Hierarchical Transformer model over existing MTL approaches. Jianfei Yu, Jing Jiang 0001, Ling Min Serena Khoo, Hai Leong Chieu |
EMNLP (1) | 1 |
| 2020 | Hierarchical Multimodal Transformer with Localness and Speaker Aware Attention for Emotion Recognition in Conversations
Jianfei Yu, Zixiang Ding, Xiangsheng Zhou, Yaofeng Tu |
NLPCC (2) | 2 |
| 2020 | Entity-Sensitive Attention and Fusion Network for Entity-Level Multimodal Sentiment ClassificationabstractEntity-level (aka target-dependent) sentiment analysis of social media posts has recently attracted increasing attention, and its goal is to predict the sentiment orientations over individual target entities mentioned in users' posts. Most existing approaches to this task primarily rely on the textual content, but fail to consider the other important data sources (e.g., images, videos, and user profiles), which can potentially enhance these text-based approaches. Motivated by the observation, we study entity-level multimodal sentiment classification in this article, and aim to explore the usefulness of images for entity-level sentiment detection in social media posts. Specifically, we propose an Entity-Sensitive Attention and Fusion Network (ESAFN) for this task. First, to capture the intra-modality dynamics, ESAFN leverages an effective attention mechanism to generate entity-sensitive textual representations, followed by aggregating them with a textual fusion layer. Next, ESAFN learns the entity-sensitive visual representation with an entity-oriented visual attention mechanism, followed by a gated mechanism to eliminate the noisy visual context. Moreover, to capture the inter-modality dynamics, ESAFN further fuses the textual and visual representations with a bilinear interaction layer. To evaluate the effectiveness of ESAFN, we manually annotate the sentiment orientation over each given entity based on two recently released multimodal NER datasets, and show that ESAFN can significantly outperform several highly competitive unimodal and multimodal methods. Jianfei Yu, Jing Jiang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Aspect and Opinion Aware Abstractive Review Summarization with Reinforced Hard Typed DecoderabstractIn this paper, we study abstractive review summarization. Observing that review summaries often consist of aspect words, opinion words and context words, we propose a two-stage reinforcement learning approach, which first predicts the output word type from the three types, and then leverages the predicted word type to generate the final word distribution. Experimental results on two Amazon product review datasets demonstrate that our method can consistently outperform several strong baseline approaches based on ROUGE scores. Yufei Tian, Jianfei Yu, Jing Jiang 0001 |
CIKM | 2 |
| 2019 | Adapting BERT for Target-Oriented Multimodal Sentiment ClassificationabstractAs an important task in Sentiment Analysis, Target-oriented Sentiment Classification (TSC) aims to identify sentiment polarities over each opinion target in a sentence. However, existing approaches to this task primarily rely on the textual content, but ignoring the other increasingly popular multimodal data sources (e.g., images), which can enhance the robustness of these text-based models. Motivated by this observation and inspired by the recently proposed BERT architecture, we study Target-oriented Multimodal Sentiment Classification (TMSC) and propose a multimodal BERT architecture. To model intra-modality dynamics, we first apply BERT to obtain target-sensitive textual representations. We then borrow the idea from self-attention and design a target attention mechanism to perform target-image matching to derive target-sensitive visual representations. To model inter-modality dynamics, we further propose to stack a set of self-attention layers to capture multimodal interactions. Experimental results show that our model can outperform several highly competitive approaches for TSC and TMSC. Jianfei Yu, Jing Jiang 0001 |
IJCAI | 1 |
| 2019 | Global Inference for Aspect and Opinion Terms Co-Extraction Based on Multi-Task Neural NetworksabstractExtracting aspect terms and opinion terms are two fundamental tasks in opinion mining. The recent success of deep learning has inspired various neural network architectures, which have been shown to achieve highly competitive performance in these two tasks. However, most existing methods fail to explicitly consider the syntactic relations among aspect terms and opinion terms, which may lead to the inconsistencies between the model predictions and the syntactic constraints. To this end, we first apply a multi-task learning framework to implicitly capture the relations between the two tasks, and then propose a global inference method by explicitly modelling several syntactic constraints among aspect term extraction and opinion term extraction to uncover their intra-task and inter-task relationship, which seeks an optimal solution over the neural predictions for both tasks. Extensive evaluations on three benchmark datasets demonstrate that our global inference approach is able to bring consistent improvements over several base models in different scenarios. Jianfei Yu, Jing Jiang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Improving Multi-label Emotion Classification via Sentiment Classification with Dual Attention Transfer NetworkabstractIn this paper, we target at improving the performance of multi-label emotion classification with the help of sentiment classification.Specifically, we propose a new transfer learning architecture to divide the sentence representation into two different feature spaces, which are expected to respectively capture the general sentiment words and the other important emotion-specific words via a dual attention mechanism.Extensive experimental results demonstrate that our transfer learning approach can outperform several strong baselines and achieve the state-of-the-art performance on two benchmark datasets. Jianfei Yu, Luís Marujo, Jing Jiang 0001, Pradeep Karuturi, William Brendel |
EMNLP | 1 |
| 2018 | Densely Connected Bidirectional LSTM with Applications to Sentence Classification
Zixiang Ding, Jianfei Yu, Xiang Li 0041, Jian Yang 0003 |
NLPCC (2) | 3 |
| 2018 | Modelling Domain Relationships for Transfer Learning on Retrieval-based Question Answering Systems in E-commerceabstractNowadays, it is a heated topic for many industries to build automatic question-answering (QA) systems. A key solution to these QA systems is to retrieve from a QA knowledge base the most similar question of a given question, which can be reformulated as a paraphrase identification (PI) or a natural language inference (NLI) problem. However, most existing models for PI and NLI have at least two problems: They rely on a large amount of labeled data, which is not always available in real scenarios, and they may not be efficient for industrial applications. In this paper, we study transfer learning for the PI and NLI problems, aiming to propose a general framework, which can effectively and efficiently adapt the shared knowledge learned from a resource-rich source domain to a resource-poor target domain. Specifically, since most existing transfer learning methods only focus on learning a shared feature space across domains while ignoring the relationship between the source and target domains, we propose to simultaneously learn shared representations and domain relationships in a unified framework. Furthermore, we propose an efficient and effective hybrid model by combining a sentence encoding-based method and a sentence interaction-based method as our base model. Extensive experiments on both paraphrase identification and natural language inference demonstrate that our base model is efficient and has promising performance compared to the competing models, and our transfer learning method can help to significantly boost the performance. Further analysis shows that the inter-domain and intra-domain relationship captured by our model are insightful. Last but not least, we deploy our transfer learning model for PI into our online chatbot system, which can bring in significant improvements over our existing system. Finally, we launch our new system on the chatbot platform Eva in our E-commerce site AliExpress. Jianfei Yu, Minghui Qiu, Jing Jiang 0001, Jun Huang 0007, Shuangyong Song, Haiqing Chen |
WSDM | 1 |
| 2017 | Recurrent Neural Networks with Auxiliary Labels for Cross-Domain Opinion Target ExtractionabstractOpinion target extraction is a fundamental task in opinion mining. In recent years, neural network based supervised learning methods have achieved competitive performance on this task. However, as with any supervised learning method, neural network based methods for this task cannot work well when the training data comes from a different domain than the test data. On the other hand, some rule-based unsupervised methods have shown to be robust when applied to different domains. In this work, we use rule-based unsupervised methods to create auxiliary labels and use neural network models to learn a hidden representation that works well for different domains. When this hidden representation is used for opinion target extraction, we find that it can outperform a number of strong baselines with a large margin. Ying Ding 0005, Jianfei Yu, Jing Jiang 0001 |
AAAI | 2 |
| 2017 | Leveraging Auxiliary Tasks for Document-Level Cross-Domain Sentiment ClassificationabstractIn this paper, we study domain adaptation with a state-of-the-art hierarchical neural network for document-level sentiment classification. We first design a new auxiliary task based on sentiment scores of domain-independent words. We then propose two neural network architectures to respectively induce document embeddings and sentence embeddings that work well for different domains. When these document and sentence embeddings are used for sentiment classification, we find that with both pseudo and external sentiment lexicons, our proposed methods can perform similarly to or better than several highly competitive domain adaptation methods on a benchmark dataset of product reviews. Jianfei Yu, Jing Jiang 0001 |
IJCNLP(1) | 1 |
| 2016 | Pairwise Relation Classification with Mirror Instances and a Combined Convolutional Neural NetworkabstractRelation classification is the task of classifying the semantic relations between entity pairs in text. Observing that existing work has not fully explored using different representations for relation instances, especially in order to better handle the asymmetry of relation types, in this paper, we propose a neural network based method for relation classification that combines the raw sequence and the shortest dependency path representations of relation instances and uses mirror instances to perform pairwise relation classification. We evaluate our proposed models on the SemEval-2010 Task 8 dataset. The empirical results show that with two additional features, our model achieves the state-of-the-art result of F1 score of 85.7. Jianfei Yu, Jing Jiang 0001 |
COLING | 1 |
| 2016 | Learning Sentence Embeddings with Auxiliary Tasks for Cross-Domain Sentiment ClassificationabstractNational Research Foundation (NRF) Singapore under International Research Centres in Singapore Funding Initiative Jianfei Yu, Jing Jiang 0001 |
EMNLP | 1 |
| 2016 | Polarity shift detection, elimination and ensemble: A three-stage model for document-level sentiment analysis
Jianfei Yu, Erik Cambria |
Inf. Process. Manag. | 3 |
| 2014 | Instance-Based Domain Adaptation in NLP via In-Target-Domain Logistic ApproximationabstractIn the field of NLP, most of the existing domain adaptation studies belong to the feature-based adaptation, while the research of instance-based adaptation is very scarce. In this work, we propose a new instance-based adaptation model, called in-target-domain logistic approximation (ILA). In ILA, we adapt the source-domain data to the target domain by a logistic approximation. The normalized in-target-domain probability is assigned as an instance weight to each of the source-domain training data. An instance-weighted classification model is trained finally for the cross-domain classification problem. Compared to the previous techniques, ILA conducts instance adaptation in a dimensionality-reduced linear feature space to ensure efficiency in high-dimensional NLP tasks. The instance weights in ILA are learnt by leveraging the criteria of both maximum likelihood and minimum statistical distance. The empirical results on two NLP tasks including text categorization and sentiment classification show that our ILA model beats the state-of-the-art instance adaptation methods significantly, in cross-domain classification accuracy, parameter stability and computational efficiency. Jianfei Yu, Shumei Wang |
AAAI | 2 |