EDBT 2026 Demo / reviewers in the wild / expert
Zhen Wu 0002
dblp:16/4485-2
· DBLP profile ↗
33ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0002-7678-103XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing RAG Rerankers with LLM Feedback via Reinforcement LearningabstractRerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation.However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process.This isolation leads to a fundamental misalignment: documents identified as topically relevant by information retrieval metrics often fail to provide the actual utility required by the LLM for precise answer generation.To bridge this gap, we introduce ReRanking Preference Optimization (RRPO) 1 , a reinforcement learning framework that directly aligns reranking with the LLM's generation quality.By formulating reranking as a sequential decision-making process, RRPO optimizes for context utility using LLM feedback, thereby eliminating the need for expensive human annotations.To ensure training stability, we further introduce a reference-anchored deterministic baseline.Extensive experiments on knowledge-intensive benchmarks demonstrate that RRPO significantly outperforms strong baselines, including the powerful list-wise reranker RankZephyr.Further analysis highlights the versatility of our framework: it generalizes seamlessly to diverse readers (e.g., GPT-4o), integrates orthogonally with query expansion modules like Query2Doc, and remains robust even when trained with noisy supervisors. Xiangqing Shen, Fanfan Wang, Cangqi Zhou, Zhen Wu 0002, Xinyu Dai |
ACL (1) | 5 |
| 2025 | Bridging Questions and Charts: A Weakly Supervised Alignment Model for Chart Question Answering
Jiangzhou Ju, Yunlin Mao, Zhen Wu 0002, Robert Ridley, Jiajun Chen 0001, Xinyu Dai |
NLPCC (3) | 3 |
| 2025 | Persona-Aware Alignment Framework for Personalized Dialogue GenerationabstractAbstract Personalized dialogue generation aims to leverage persona profiles and dialogue history to generate persona-relevant and consistent responses. Mainstream models typically rely on token-level language model training with persona dialogue data, such as Next Token Prediction, to implicitly achieve personalization, meaning that these methods tend to neglect the given personas and generate generic responses. To address this issue, we propose a novel Persona-Aware Alignment Framework (PAL), which directly treats persona alignment as the training objective of dialogue generation. Specifically, PAL employs a two-stage training method including Persona-Aware Learning and Persona Alignment, equipped with an easy-to-use inference strategy Select then Generate, to improve persona sensitivity and generate more persona-relevant responses at the semantics level. Through extensive experiments, we demonstrate that our framework outperforms many state-of-the-art personalized dialogue methods and large language models. Guanrong Li, Zhen Wu 0002, Xinyu Dai |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Dr3: Ask Large Language Models Not to Give Off-Topic Answers in Open Domain Multi-Hop Question AnsweringabstractOpen Domain Multi-Hop Question Answering (ODMHQA) plays a crucial role in Natural Language Processing (NLP) by aiming to answer complex questions through multi-step reasoning over retrieved information from external knowledge sources. Recently, Large Language Models (LLMs) have demonstrated remarkable performance in solving ODMHQA owing to their capabilities including planning, reasoning, and utilizing tools. However, LLMs may generate off-topic answers when attempting to solve ODMHQA, namely the generated answers are irrelevant to the original questions. This issue of off-topic answers accounts for approximately one-third of incorrect answers, yet remains underexplored despite its significance. To alleviate this issue, we propose the Discriminate→Re-Compose→Re- Solve→Re-Decompose (Dr3) mechanism. Specifically, the Discriminator leverages the intrinsic capabilities of LLMs to judge whether the generated answers are off-topic. In cases where an off-topic answer is detected, the Corrector performs step-wise revisions along the reversed reasoning chain (Re-Compose→Re-Solve→Re-Decompose) until the final answer becomes on-topic. Experimental results on the HotpotQA and 2WikiMultiHopQA datasets demonstrate that our Dr3 mechanism considerably reduces the occurrence of off-topic answers in ODMHQA by nearly 13%, improving the performance in Exact Match (EM) by nearly 3% compared to the baseline method without the Dr3 mechanism. Yiheng Zhu 0004, Yuanbin Cao, Yinzhi Zhou, Zhen Wu 0002, Shenglan Wu, Haoyuan Hu, Xinyu Dai |
LREC/COLING | 5 |
| 2024 | EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) have attracted increasing attention in the past few years, but they may still generate descriptions that include objects not present in the corresponding images, a phenomenon known as object hallucination.To eliminate hallucinations, existing methods manually annotate paired responses with and without hallucinations, and then employ various alignment algorithms to improve the alignment capability between images and text.However, they not only demand considerable computation resources during the finetuning stage but also require expensive human annotation to construct paired data needed by the alignment algorithms.To address these issues, we propose an efficient fine-grained unlearning framework (EFUF), which performs gradient ascent utilizing three tailored losses to eliminate hallucinations without paired data.Extensive experiments show that our method consistently reduces hallucinations while preserving the generation quality with modest computational overhead.Our code and datasets are available at https://github.com/starreeze/efuf. Shangyu Xing, Fei Zhao 0012, Zhen Wu 0002, Tuo An, Xinyu Dai |
EMNLP | 3 |
| 2024 | Focus & Gating: A Multimodal Approach for Unveiling Relations in Noisy Social MediaabstractMultimedia content's surge on the internet has made multimodal relation extraction vital for applications like intelligent search and knowledge graph construction. As a rich source of image-text data, social media plays a crucial role in populating knowledge bases. However, the noisy information present in social media poses a challenge in multimodal relation extraction. Current methods focus on extracting relevant information from images to improve model performance but often overlook the importance of global image information. In this paper, we propose a novel multimodal relation extraction method FocalMRE, which leverages image focal augmentation, focal attention, and gating mechanisms. FocalMRE enables the model to concentrate on the image's focal regions while effectively utilizing the global information in the image. Through gating mechanisms, FocalMRE optimizes the multimodal fusion strategy, allowing the model to select the most relevant augmented regions for overcoming noise interference in relation extraction. The experimental results on the public MNRE dataset reveal that FocalMRE exhibits robust and significant performance advantages in the multimodal relation extraction task, especially in scenarios with high noise, long-tail distributions, and limited resources. The code is available at https://github.com/NJUNLP/FocalMRE. Liang He 0009, Hongke Wang, Zhen Wu 0002, Xinyu Dai, Jiajun Chen 0001 |
ACM Multimedia | 3 |
| 2024 | AutoSurvey: Large Language Models Can Automatically Write SurveysabstractThis paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial intelligence. Traditional survey paper creation faces challenges due to the vast volume and complexity of information, prompting the need for efficient survey methods. While large language models (LLMs) offer promise in automating this process, challenges such as context window limitations, parametric knowledge constraints, and the lack of evaluation benchmarks remain. AutoSurvey addresses these challenges through a systematic approach that involves initial retrieval and outline generation, subsection drafting by specialized LLMs, integration and refinement, and rigorous evaluation and iteration. Our contributions include a comprehensive solution to the survey problem, a reliable evaluation method, and experimental validation demonstrating AutoSurvey's effectiveness. Yidong Wang 0003, Wenjin Yao, Xin Zhang 0097, Zhen Wu 0002, Meishan Zhang, Xinyu Dai, Min Zhang 0005, Qingsong Wen, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
NeurIPS | 6 |
| 2023 | On Prefix-tuning for Lightweight Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection, a fundamental task vexing real-world applications, has attracted growing attention in the NLP community.Recently fine-tuning based methods have made promising progress.However, it could be costly to store fine-tuned models for each scenario.In this paper, we depart from the classic fine-tuning based OOD detection toward a parameter-efficient alternative, and propose an unsupervised prefix-tuning based OOD detection framework termed PTO.Additionally, to take advantage of optional training data labels and targeted OOD data, two practical extensions of PTO are further proposed.Overall, PTO and its extensions offer several key advantages of being lightweight, easy-to-reproduce, and theoretically justified.Experimental results show that our methods perform comparably to, even better than, existing fine-tuning based OOD detection approaches under a wide range of metrics, detection settings, and OOD types. Yawen Ouyang, Yongchang Cao, Zhen Wu 0002, Xinyu Dai |
ACL (1) | 4 |
| 2023 | Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP DomainabstractIt has been well documented that a reviewer's opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication.Previous works have also shone a light on the stress and anxiety authors who are non-native English speakers experience when attempting to publish in international venues.We explore how this might be a concern in the field of Natural Language Processing (NLP) through conducting a comprehensive statistical analysis of NLP paper abstracts, identifying how authors of different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing.Through our analysis, we identify that there are a number of characteristics that are highly variable across the different corpora examined in this paper.This indicates potential for the presence of linguistic bias.Therefore, we outline a set of recommendations to publishers of academic journals and conferences regarding their guidelines and resources for prospective authors in order to help enhance inclusivity and fairness. Robert Ridley, Zhen Wu 0002, Shujian Huang, Xinyu Dai |
EMNLP | 2 |
| 2023 | M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisabstractMultimodal Aspect-based Sentiment Analysis (MABSA) is a fine-grained Sentiment Analysis task, which has attracted growing research interests recently.Existing work mainly utilizes image information to improve the performance of MABSA task.However, most of the studies overestimate the importance of images since there are many noisy images unrelated to the text in the dataset, which will have a negative impact on model learning.Although some work attempts to filter low-quality noisy images by setting thresholds, relying on thresholds will inevitably filter out a lot of useful image information.Therefore, in this work, we focus on whether the negative impact of noisy images can be reduced without filtering the data.To achieve this goal, we borrow the idea of Curriculum Learning and propose a Multi-grained Multi-curriculum Denoising Framework (M2DF), which can achieve denoising by adjusting the order of training data.Extensive experimental results show that our framework consistently outperforms state-ofthe-art work on three sub-tasks of MABSA.Our code and datasets are available at https: //github.com/grandchicken/M2DF. Fei Zhao 0012, Zhen Wu 0002, Yawen Ouyang, Xinyu Dai |
EMNLP | 3 |
| 2023 | An Empirical Study and Improvement for Speech Emotion RecognitionabstractMultimodal speech emotion recognition aims to detect speakers’ emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while neglecting the effect of different fusion strategies on emotion recognition. In this work, we consider a simple yet important problem: how to fuse audio and text modality information is more helpful for this multimodal task. Further, we propose a multimodal emotion recognition model improved by perspective loss. Empirical results show our method obtained new state-of-the-art results on the IEMOCAP dataset. The in-depth analysis explains why the improved model can achieve improvements and outperforms baselines. Zhen Wu 0002, Yizhe Lu, Xinyu Dai |
ICASSP | 1 |
| 2023 | FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Yidong Wang 0003, Hao Chen 0102, Qiang Heng, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie 0001 |
ICLR | 6 |
| 2023 | Make BERT-based Chinese Spelling Check Model Enhanced by Layerwise Attention and Gaussian Mixture ModelabstractBERT-based models have shown a remarkable ability in the Chinese Spelling Check (CSC) task recently. However, traditional BERT-based methods still suffer from two limitations. First, although previous works have identified that explicit prior knowledge like Part-Of-Speech (POS) tagging can benefit in the CSC task, they neglected the fact that spelling errors inherent in CSC data can lead to incorrect tags and therefore mislead models. Additionally, they ignored the correlation between the implicit hierarchical information encoded by BERT's intermediate layers and different linguistic phenomena. This results in sub-optimal accuracy. To alleviate the above two issues, we design a heterogeneous knowledge-infused framework to strengthen BERT-based CSC models. To incorporate explicit POS knowledge, we utilize an auxiliary task strategy driven by Gaussian mixture model. Meanwhile, to incorporate implicit hierarchical linguistic knowledge within the encoder, we propose a novel form of n-gram-based layerwise self-attention to generate a multilayer representation. Experimental results show that our proposed framework yields a stable performance boost over four strong baseline models and outperforms the previous state-of-the-art methods on two datasets. Yongchang Cao, Liang He 0009, Zhen Wu 0002, Xinyu Dai |
IJCNN | 3 |
| 2023 | MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationabstractExtracting relational facts from multimodal data is a crucial task in the field of multimedia and knowledge graphs that feeds into widespread real-world applications. The emphasis of recent studies centers on recognizing relational facts in which both entities are present in one modality and supplementary information is used from other modalities. However, such works disregard a substantial amount of multimodal relational facts that arise across different modalities, such as one entity seen in a text and another in an image. In this paper, we propose a new task, namely Multimodal Object-Entity Relation Extraction, which aims to extract "object-entity" relational facts from image and text data. To facilitate research on this task, we introduce MORE, a new dataset comprising 21 relation types and 20,136 multimodal relational facts annotated on 3,522 pairs of textual news titles and corresponding images. To show the challenges of Multimodal Object-Entity Relation Extraction, we evaluated recent state-of-the-art methods for multimodal relation extraction and conducted a comprehensive experimentation analysis on MORE. Our results demonstrate significant challenges for existing methods, underlining the need for further research on this task. Based on our experiments, we identify several promising directions for future research. The MORE dataset and code are available at https://github.com/NJUNLP/MORE. Liang He 0009, Hongke Wang, Yongchang Cao, Zhen Wu 0002, Xinyu Dai |
ACM Multimedia | 4 |
| 2023 | DRIN: Dynamic Relation Interactive Network for Multimodal Entity LinkingabstractMultimodal Entity Linking (MEL) is a task that aims to link ambiguous mentions within multimodal contexts to referential entities in a multimodal knowledge base. Recent methods for MEL adopt a common framework: they first interact and fuse the text and image to obtain representations of the mention and entity respectively, and then compute the similarity between them to predict the correct entity. However, these methods still suffer from two limitations: first, as they fuse the features of text and image before matching, they cannot fully exploit the fine-grained alignment relations between the mention and entity. Second, their alignment is static, leading to low performance when dealing with complex and diverse data. To address these issues, we propose a novel framework called Dynamic Relation Interactive Network (DRIN) for MEL tasks. DRIN explicitly models four different types of alignment between a mention and entity and builds a dynamic Graph Convolutional Network (GCN) to dynamically select the corresponding alignment relations for different input samples. Experiments on two datasets show that DRIN outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of our approach. Our code and datasets are publicly available. Shangyu Xing, Fei Zhao 0012, Zhen Wu 0002, Xinyu Dai |
ACM Multimedia | 3 |
| 2023 | Episode-Based Prompt Learning for Any-Shot Intent Detection
Dingjie Song, Yawen Ouyang, Zhen Wu 0002, Xinyu Dai |
NLPCC (1) | 4 |
| 2023 | IDOS: A Unified Debiasing Method via Word Shuffling
Yuanhang Tang, Yawen Ouyang, Zhen Wu 0002, Xinyu Dai |
NLPCC (2) | 3 |
| 2023 | Label-Correction Capsule Network for Hierarchical Text ClassificationabstractHierarchical Text Classification (HTC) aims to predict the category of a document in a given label hierarchy. Considering a parent-child relationship among labels at different levels, previous works mainly leverage the parent-level label information to guide the child-level classification and achieve promising results. However, they still suffer from two drawbacks: (1) insufficient for distinguishing similar labels at the same level; (2) fail to consider the error propagation problem caused by the incorrect parent-level predictions. For this reason, we first propose a hierarchical capsule network for the HTC task, due to the ability of capsules to distinguish similar categories. To ease the error propagation problem, we further devise two novel mechanisms in the proposed hierarchical capsule framework, i.e.,Label InjectionandLabel Re-Routing, to enhance the tolerance of the model to the incorrect parent-level predictions. Experiments on two widely used datasets prove that our model achieves competitive performance. The ablation study further demonstrates the scalability ofLabel InjectionandLabel Re-Routing. Fei Zhao 0012, Zhen Wu 0002, Liang He 0009, Xinyu Dai |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Margin Calibration for Long-Tailed Visual Recognition
Yidong Wang 0003, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki |
ACML | 4 |
| 2022 | Towards Multi-label Unknown Intent DetectionabstractMulti-class unknown intent detection has made remarkable progress recently. However, it has a strong assumption that each utterance has only one intent, which does not conform to reality because utterances often have multiple intents. In this paper, we propose a more desirable task, multi-label unknown intent detection, to detect whether the utterance contains the unknown intent, in which each utterance may contain multiple intents. In this task, the unique utterances simultaneously containing known and unknown intents make existing multi-class methods easy to fail. To address this issue, we propose an intuitive and effective method to recognize whether All Intents contained in the utterance are Known (AIK). Our high-level idea is to predict the utterance’s intent number, then check whether the utterance contains the same number of known intents. If the number of known intents is less than the number of intents, it implies that the utterance also contains unknown intents. We benchmark AIK over existing methods, and empirical results suggest that our method obtains state-of-the-art performances. For example, on the MultiWOZ 2.3 dataset, AIK significantly reduces the FPR95 by 12.25% compared to the best baseline. Yawen Ouyang, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
COLING | 2 |
| 2022 | Exploiting Unlabeled Data for Target-Oriented Opinion Words ExtractionabstractTarget-oriented Opinion Words Extraction (TOWE) is a fine-grained sentiment analysis task that aims to extract the corresponding opinion words of a given opinion target from the sentence. Recently, deep learning approaches have made remarkable progress on this task. Nevertheless, the TOWE task still suffers from the scarcity of training data due to the expensive data annotation process. Limited labeled data increase the risk of distribution shift between test data and training data. In this paper, we propose exploiting massive unlabeled data to reduce the risk by increasing the exposure of the model to varying distribution shifts. Specifically, we propose a novel Multi-Grained Consistency Regularization (MGCR) method to make use of unlabeled data and design two filters specifically for TOWE to filter noisy data at different granularity. Extensive experimental results on four TOWE benchmark datasets indicate the superiority of MGCR compared with current state-of-the-art methods. The in-depth analysis also demonstrates the effectiveness of the different-granularity filters. Yidong Wang 0003, Hao Wu 0059, Ao Liu 0008, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki, Manabu Okumura, Yue Zhang 0004 |
COLING | 5 |
| 2022 | Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment ClassificationabstractTarget-oriented multimodal sentiment classification (TMSC) is a new subtask of aspect-based sentiment analysis, which aims to determine the sentiment polarity of the opinion target mentioned in a (sentence, image) pair. Recently, dominant works employ the attention mechanism to capture the corresponding visual representations of the opinion target, and then aggregate them as evidence to make sentiment predictions. However, they still suffer from two problems: (1) The granularity of the opinion target in two modalities is inconsistent, which causes visual attention sometimes fail to capture the corresponding visual representations of the target; (2) Even though it is captured, there are still significant differences between the visual representations expressing the same mood, which brings great difficulty to sentiment prediction. To this end, we propose a novel Knowledge-enhanced Framework (KEF) in this paper, which can successfully exploit adjective-noun pairs extracted from the image to improve the visual attention capability and sentiment prediction capability of the TMSC task. Extensive experimental results show that our framework consistently outperforms state-of-the-art works on two public datasets. Fei Zhao 0012, Zhen Wu 0002, Siyu Long, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
COLING | 2 |
| 2022 | Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERabstractMultimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a (text, image) pair. However, dominant work independently models the internal matching relations in a pair of image and text, ignoring the external matching relations between different (text, image) pairs inside the dataset, though such relations are crucial for alleviating image noise in MNER task. In this paper, we primarily explore two kinds of external matching relations between different (text, image) pairs, i.e., inter-modal relations and intra-modal relations. On the basis, we propose a Relation-enhanced Graph Convolutional Network (R-GCN) for the MNER task. Specifically, we first construct an inter-modal relation graph and an intra-modal relation graph to gather the image information most relevant to the current text and image from the dataset, respectively. And then, multimodal interaction and fusion are leveraged to predict the NER label sequences. Extensive experimental results show that our model consistently outperforms state-of-the-art works on two public datasets. Our code and datasets are available at https://github.com/1429904852/R-GCN. Fei Zhao 0012, Zhen Wu 0002, Shangyu Xing, Xinyu Dai |
ACM Multimedia | 3 |
| 2022 | USB: A Unified Semi-supervised Learning Benchmark for ClassificationabstractSemi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL. Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004 |
NeurIPS | 12 |
| 2022 | SPRoBERTa: protein embedding learning with local fragment modelingabstractWell understanding protein function and structure in computational biology helps in the understanding of human beings. To face the limited proteins that are annotated structurally and functionally, the scientific community embraces the self-supervised pre-training methods from large amounts of unlabeled protein sequences for protein embedding learning. However, the protein is usually represented by individual amino acids with limited vocabulary size (e.g. 20 type proteins), without considering the strong local semantics existing in protein sequences. In this work, we propose a novel pre-training modeling approach SPRoBERTa. We first present an unsupervised protein tokenizer to learn protein representations with local fragment pattern. Then, a novel framework for deep pre-training model is introduced to learn protein embeddings. After pre-training, our method can be easily fine-tuned for different protein tasks, including amino acid-level prediction task (e.g. secondary structure prediction), amino acid pair-level prediction task (e.g. contact prediction) and also protein-level prediction task (remote homology prediction, protein function prediction). Experiments show that our approach achieves significant improvements in all tasks and outperforms the previous methods. We also provide detailed ablation studies and analysis for our protein tokenizer and training framework. Lijun Wu 0003, Chengcan Yin, Jinhua Zhu 0001, Zhen Wu 0002, Liang He 0010, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Tie-Yan Liu |
Briefings Bioinform. | 4 |
| 2022 | Pairwise tagging framework for end-to-end emotion-cause pair extraction
Zhen Wu 0002, Xinyu Dai |
Frontiers Comput. Sci. | 1 |
| 2021 | UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra CostabstractZhen Wu, Lijun Wu, Qi Meng, Yingce Xia, Shufang Xie, Tao Qin, Xinyu Dai, Tie-Yan Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhen Wu 0002, Lijun Wu 0003, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Xinyu Dai, Tie-Yan Liu |
NAACL-HLT | 1 |
| 2020 | Latent Opinions Transfer Network for Target-Oriented Opinion Words ExtractionabstractTarget-oriented opinion words extraction (TOWE) is a new subtask of ABSA, which aims to extract the corresponding opinion words for a given opinion target in a sentence. Recently, neural network methods have been applied to this task and achieve promising results. However, the difficulty of annotation causes the datasets of TOWE to be insufficient, which heavily limits the performance of neural models. By contrast, abundant review sentiment classification data are easily available at online review sites. These reviews contain substantial latent opinions information and semantic patterns. In this paper, we propose a novel model to transfer these opinions knowledge from resource-rich review sentiment classification datasets to low-resource task TOWE. To address the challenges in the transfer process, we design an effective transformation method to obtain latent opinions, then integrate them into TOWE. Extensive experimental results show that our model achieves better performance compared to other state-of-the-art methods and significantly outperforms the base model without transferring opinions knowledge. Further analysis validates the effectiveness of our model. Zhen Wu 0002, Fei Zhao 0012, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
AAAI | 1 |
| 2020 | Attention Transfer Network for Aspect-level Sentiment ClassificationabstractAspect-level sentiment classification (ASC) aims to detect the sentiment polarity of a given opinion target in a sentence.In neural network-based methods for ASC, most works employ the attention mechanism to capture the corresponding sentiment words of the opinion target, then aggregate them as evidence to infer the sentiment of the target.However, aspect-level datasets are all relatively small-scale due to the complexity of annotation.Data scarcity causes the attention mechanism sometimes to fail to focus on the corresponding sentiment words of the target, which finally weakens the performance of neural models.To address the issue, we propose a novel Attention Transfer Network (ATN) in this paper, which can successfully exploit attention knowledge from resource-rich document-level sentiment classification datasets to improve the attention capability of the aspect-level sentiment classification task.In the ATN model, we design two different methods to transfer attention knowledge and conduct experiments on two ASC benchmark datasets.Extensive experimental results show that our methods consistently outperform state-of-the-art works.Further analysis also validates the effectiveness of ATN.Our code and dataset are available at https Fei Zhao 0012, Zhen Wu 0002, Xinyu Dai |
COLING | 2 |
| 2020 | Transformer-Based Multi-aspect Modeling for Multi-aspect Multi-sentiment Analysis
Zhen Wu 0002, Chengcan Ying, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (2) | 1 |
| 2020 | Opinion Transmission Network for Jointly Improving Aspect-Oriented Opinion Words Extraction and Sentiment Classification
Chengcan Ying, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (1) | 2 |
| 2018 | Improving Review Representations With User Attention and Product Attention for Sentiment ClassificationabstractNeural network methods have achieved great success in reviews sentiment classification. Recently, some works achieved improvement by incorporating user and product information to generate a review representation. However, in reviews, we observe that some words or sentences show strong user's preference, and some others tend to indicate product's characteristic. The two kinds of information play different roles in determining the sentiment label of a review. Therefore, it is not reasonable to encode user and product information together into one representation. In this paper, we propose a novel framework to encode user and product information. Firstly, we apply two individual hierarchical neural networks to generate two representations, with user attention or with product attention. Then, we design a combined strategy to make full use of the two representations for training and final prediction. The experimental results show that our model obviously outperforms other state-of-the-art methods on IMDB and Yelp datasets. Through the visualization of attention over words related to user or product, we validate our observation mentioned above. Zhen Wu 0002, Xinyu Dai, Cunyan Yin, Shujian Huang, Jiajun Chen 0001 |
AAAI | 1 |
| 2018 | Improving Aspect Identification with Reviews Segmentation
Tianhao Ning, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (1) | 2 |