Junhui Li 0001

dblp:08/500-1 · DBLP profile ↗
← Back
55ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0001-7829-6348ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 7 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing document-level translation of large language model via translation mixed instructions
Yachao Li 0003, Junhui Li 0001, Min Zhang 0005
Neurocomputing2
2026 Multiphase and Multitask Prompt Tuning for LLM-Based Context-Aware Machine Translation
abstract
Large language models (LLMs) are typically adapted for context-aware machine translation (MT) by combining both the source sentence and its surrounding sentences into a single input. This unified input is then processed in one go, with the model producing the target translation step by step. However, this method treats the intrasentence and intersentence contexts similarly, even though they play distinct roles. In this study, we present a novel strategy called multiphase prompt tuning (MPT) to address this issue by enabling LLMs to treat these two context types differently. MPT divides the context-aware translation task into three phases: encoding the intersentence context, encoding the source sentence, and the final decoding phase. Each phase incorporates distinct continuous prompts that help the model focus on the appropriate task for each type of context. We also introduce a multitask fine-tuning approach to emphasize the distinction between intersentence and intrasentence contexts and enhance intersentence dependencies. This includes two auxiliary tasks: context-agnostic translation and cross-lingual next sentence generation, which help extract additional information and improve the model's handling of discourse-related challenges.
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE Trans. Neural Networks Learn. Syst.2
2025 Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
abstract
Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub.
Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
ACL (1)3
2025 Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
abstract
Direct speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance.
Suhang Wu, Jialong Tang, Pei Zhang 0011, Baosong Yang, Junhui Li 0001, Junfeng Yao, Min Zhang 0005, Jinsong Su
ACL (1)6
2025 Improving LLM-Based Document-Level MT with Multi-Knowledge Fusion
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006
NLPCC (3)3
2025 DialoguePFM:Prompt-based Fusion Model for Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) presents a significant challenge in natural language processing. In this study, we propose the Prompt-based Fusion Model (DialoguePFM), which innovatively introduces an emotion representation that conveys the emotion label via mask token in a pre-defined prompt. We then refine both emotion and utterance representations by capturing comprehensive dialogue information using a novel speaker-aware attention mechanism, which distinguishes between self and other speakers. Subsequently, these refined representations are merged before being inputted into the classifier. Empirical evaluations conducted on three English ERC datasets and one Chinese ERC dataset reveal that our proposed model either outperforms or matches the performance of state-of-the-art baselines, underscoring its effectiveness across different languages.
Yu Tian 0019, Junhui Li 0001, Suyang Zhu, Guodong Zhou 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2025 Enhanced Generative Framework With LLMs for Multimodal Emotion-Cause Pair Extraction in Conversations
abstract
Emotion-Cause Pair Extraction (ECPE) in conversations aims to identify the emotional utterances (even their categories) along with their corresponding causal utterances, which is crucial in understanding the cause-effect relationship in dialogues. While prior studies of ECPE have predominantly focused on purely textual dialogues and neglected the exploration on the natural scenario of the dialogues with multimodal features, i.e., Multimodal Emotion-Cause Pair Extraction (MECPE) in conversations. To attempt this scenario, we propose a Generative approach for Multimodal Emotion-Cause pair extraction (GMEC) with a single stage, thus effectively reducing errors associated with the propagation and accumulation for MECPE. This approach can not only uniformly handle the information of diverse modalities, but also address all emotion and cause analysis tasks uniformly. Additionally, instead of utilizing the fixed commonsense knowledge base as previously, we resort to the Large Language Models (LLMs), which possess a powerful ability to emerge new knowledge, thereby acting as implicit knowledge engines for MECPE. We refer to this approach as enhanced GMEC. Extensive experimental results and detailed analysis demonstrate a notable improvement in the generative approach. Moreover, the integration of external knowledge from LLMs optimizes the efficiency of data utilization, particularly in few-shot scenarios. The integration of the generative model with LLMs has resulted in a cumulative enhancement of 4.94%, 10.90% on MECPE and MECPE-C (with emotion Category).
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
IEEE Trans. Multim.3
2024 Leveraging AMR Graph Structure for Better Sequence-to-Sequence AMR Parsing
abstract
Thanks to the development of pre-trained sequence-to-sequence (seq2seq) models (e.g., BART), recent studies on AMR parsing often regard this task as a seq2seq translation problem by linearizing AMR graphs into AMR token sequences in pre-processing and recovering AMR graphs from sequences in post-processing. Seq2seq AMR parsing is a relatively simple paradigm but it unavoidably loses structural information among AMR tokens. To compensate for the loss of structural information, in this paper we explicitly leverage AMR structure in the decoding phase. Given an AMR graph, we first project the structure in the graph into an AMR token graph, i.e., structure among AMR tokens in the linearized sequence. The structures for an AMR token could be divided into two parts: structure in prediction history and structure in future. Then we propose to model structure in prediction history via a graph attention network (GAT) and learn structure in future via a multi-task scheme, respectively. Experimental results show that our approach significantly outperforms a strong baseline and achieves performance with 85.5 ±0.1 and 84.2 ±0.1 Smatch scores on AMR 2.0 and AMR 3.0, respectively
Linyu Fan, Wu Wu Yiheng, Junhui Li 0001, Fang Kong 0001, Guodong Zhou 0001
LREC/COLING4
2024 Submodular-based In-context Example Selection for LLMs-based Machine Translation
abstract
Large Language Models (LLMs) have demonstrated impressive performances across various NLP tasks with just a few prompts via in-context learning. Previous studies have emphasized the pivotal role of well-chosen examples in in-context learning, as opposed to randomly selected instances that exhibits unstable results.A successful example selection scheme depends on multiple factors, while in the context of LLMs-based machine translation, the common selection algorithms only consider the single factor, i.e., the similarity between the example source sentence and the input sentence.In this paper, we introduce a novel approach to use multiple translational factors for in-context example selection by using monotone submodular function maximization.The factors include surface/semantic similarity between examples and inputs on both source and target sides, as well as the diversity within examples.Importantly, our framework mathematically guarantees the coordination between these factors, which are different and challenging to reconcile.Additionally, our research uncovers a previously unexamined dimension: unlike other NLP tasks, the translation part of an example is also crucial, a facet disregarded in prior studies.Experiments conducted on BLOOMZ-7.1B and LLAMA2-13B, demonstrate that our approach significantly outperforms random selection and robust single-factor baselines across various machine translation tasks.
Baijun Ji, Xiangyu Duan, Zhenyu Qiu, Junhui Li 0001, Hao Yang 0006, Min Zhang 0005
LREC/COLING5
2024 Evaluation Dataset for Lexical Translation Consistency in Chinese-to-English Document-level Translation
abstract
Lexical translation consistency is one of the most common discourse phenomena in Chinese-to-English document-level translation. To better evaluate the performance of lexical translation consistency, previous researches assumes that all repeated source words should be translated consistently. However, constraining translations of repeated source words to be consistent will hurt word diversity and human translators tend to use different words in translation. Therefore, in this paper we construct a test set of 310 bilingual news articles to properly evaluate lexical translation consistency. We manually differentiate those repeated source words whose translations are consistent into two types: true consistency and false consistency. Then based on the constructed test set, we evaluate the performance of lexical translation consistency for several typical NMT systems.
Xiangyu Lei, Junhui Li 0001, Shimin Tao, Hao Yang 0006
LREC/COLING2
2024 DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators
abstract
Generally, the decoder-only large language models (LLMs) are adapted to context-aware neural machine translation (NMT) in a concatenating way, where LLMs take the concatenation of the source sentence (i.e., intrasentence context) and the inter-sentence context as the input, and then to generate the target tokens sequentially.This adaptation strategy, i.e., concatenation mode, considers intrasentence and inter-sentence contexts with the same priority, despite an apparent difference between the two kinds of contexts.In this paper, we propose an alternative adaptation approach, named Decoding-enhanced Multiphase Prompt Tuning (DeMPT), to make LLMs discriminately model and utilize the inter-and intra-sentence context and more effectively adapt LLMs to context-aware NMT.First, DeMPT divides the context-aware NMT process into three separate phases.During each phase, different continuous prompts are introduced to make LLMs discriminately model various information.Second, DeMPT employs a heuristic way to further discriminately enhance the utilization of the source-side interand intra-sentence information at the final decoding phase.Experiments show that our approach significantly outperforms the concatenation method, and further improves the performance of LLMs in discourse modeling.
Xinglin Lyu, Junhui Li 0001, Min Zhang 0042, Daimeng Wei, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP2
2024 ECFCON: Emotion Consequence Forecasting in Conversations
abstract
Conversation is a common form of human communication that includes extensive emotional interaction. Traditional approaches focused on studying emotions and their underlying causes in conversations. They try to address two issues: what emotions are present in the dialogue and what causes these emotions. However, these works often overlook the bidirectional nature of emotional interaction in dialogue: utterances can evoke emotions (cause), and emotions can also lead to certain utterances (consequence). Therefore, we propose a new issue: what consequences arise from these emotions? This leads to the introduction of a new task called Emotion Consequence Forecasting in CONversations (ECFCON). In this work, we first propose a corresponding dialogue-level dataset. Specifically, we select 2,780 video dialogues for annotation, totaling 39,950 utterances. Out of these, 12,391 utterances contain emotions, and 8,810 of these have discernible consequences. Then, we benchmark this task by conducting experiments from the perspectives of traditional methods, generalized LLMs prompting methods, and clue-driven hybrid methods. Both our dataset and benchmark codes are openly accessible to the public.
Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
ACM Multimedia4
2024 Neural Chat Translation as Online Document-to-Document Translation
Mengzhe Lyu, Huaixia Dou, Junhui Li 0001, Muhua Zhu, Guodong Zhou 0001
NLPCC (3)3
2023 Real-time Emotion Pre-Recognition in Conversations with Contrastive Multi-modal Dialogue Pre-training
abstract
This paper presents our pioneering effort in addressing a new and realistic scenario in multi-modal dialogue systems called Multi-modal Real-time Emotion Pre-recognition in Conversations (MREPC). The objective is to predict the emotion of a forthcoming target utterance that is highly likely to occur. We believe that this task can enhance the dialogue system's understanding of the interlocutor's state of mind, enabling it to prepare an appropriate response in advance. However, addressing MREPC poses the following challenges:1) Previous studies on emotion elicitation typically focus on textual modality and perform sentiment forecasting within a fixed contextual scenario. 2) Previous studies on multi-modal emotion recognition aim to predict the emotion of existing utterances, making it difficult to extend these approaches to MREPC due to the absence of the target utterance. To tackle these challenges, we construct two benchmark multi-modal datasets for MREPC and propose a task-specific multi-modal contrastive pre-training approach. This approach leverages large-scale unlabeled multi-modal dialogues to facilitate emotion pre-recognition for potential utterances of specific target speakers. Through detailed experiments and extensive analysis, we demonstrate that our proposed multi-modal contrastive pre-training architecture effectively enhances the performance of multi-modal real-time emotion pre-recognition in conversations.
Xincheng Ju, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Shoushan Li, Guodong Zhou 0001
CIKM4
2023 A Benchmark for Hierarchical Emotion Cause Extraction in Spoken Dialogues
abstract
Emotion cause extraction (ECE) seeks to find out what causes a given emotion, which has drawn much attention in natural language and signal processing. Conventional ECE normally focuses on single level (i.e., either word-level or clause-level) in a document scenario. However, as we known, single level ECE can not satisfy the wide applications, compared with both levels. Besides, the existing dialogue systems increasingly need empathy support. Therefore, in this paper, we propose to hierarchically extract both word and utterance-level emotion causes in the spoken dialogue scenario (HECE). We first construct two datasets (i.e., HECE-DD and HECE-IE) based on previous studies, then propose a hierarchical framework, which consists of a feature extractor, an utterance-level cause extractor, and a word-level cause extractor. In this framework, utterance and word levels can naturally preform interaction. Detailed experiments on two HECE datasets demonstrate that hierarchical extraction performs better than extraction on a single level.
Huanqin Ping, Dong Zhang 0013, Suyang Zhu, Junhui Li 0001, Guodong Zhou 0001
IEEE Signal Process. Lett.4
2023 P-Transformer: Towards Better Document-to-Document Neural Machine Translation
abstract
Directly training a document-to-document (Doc2Doc) neural machine translation (NMT) via Transformer from scratch, especially on small datasets, usually fails to converge. Our dedicated probing tasks show that 1) both the absolute position and relative position information gets gradually weakened or even vanished once it reaches the upper encoder layers, and 2) the vanishing of absolute position information in encoder output causes the training failure of Doc2Doc NMT. To alleviate this problem, we propose a position-aware Transformer (P-Transformer) to enhance both the absolute and relative position information in both self-attention and cross-attention. Specifically, we integrate absolute positional information, i.e., position embeddings, into the query-key pairs both in self-attention and cross-attention through a simple yet effective addition operation. Moreover, we also integrate relative position encoding in self-attention. The proposed P-Transformer utilizes sinusoidal position encoding and does not require any task-specified position embedding, segment embedding, or attention mechanism. Through the above methods, we build a Doc2Doc NMT model with P-Transformer, which ingests the source document and completely generates the target document in a sequence-to-sequence (seq2seq) way. In addition, P-Transformer can be applied to seq2seq-based document-to-sentence (Doc2Sent) and sentence-to-sentence (Sent2Sent) translations. Extensive experimental results of Doc2Doc NMT show that P-Transformer significantly outperforms strong baselines on the widely-used 9 document-level datasets in 7 language pairs, covering small-, middle-, and large-scales, and achieves a new state-of-the-art. Experimentation on discourse phenomena shows that our Doc2Doc NMT models improve the translation quality in both BLEU and discourse coherence. We make our code available on Github.
Yachao Li 0003, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Alleviating Exposure Bias for Neural Machine Translation via Contextual Augmentation and Self Distillation
abstract
In neural machine translation (NMT), most sequence-to-sequence (seq2seq) models are trained only with the teacher-forcing paradigm, where the ground truth history is used to predict the next ground truth word. At the inference stage, however, the decoder predicts the next token solely based on history generated from scratch. Both using ground truth history and predicting ground truth words potentially lead to exposure bias. On the one hand, to alleviate the issue of exposure bias caused by using ground truth history, we propose contextual augmentation by allowing substitution, insertion, and deletion of words. The contextual augmentation applies to target sequence to generate non-ground truth and natural history when predicting next words. On the other hand, to alleviate the exposure bias caused by predicting ground truth words, we further apply self distillation to guide the model to carry out optimization according to smoothed prediction distribution, i.e, enable the model to predict not only ground truth words, but also other potentially correct and reasonable words. Experimental results on WMT14 English$\leftrightarrow$German and IWSLT14 German$\rightarrow$English translation tasks demonstrate that our approach achieves significant improvements over Transformer on standard benchmarks. Detailed experimental analyses further reveal the effectiveness of our proposed approach on improving the translation quality.
Zhidong Liu, Junhui Li 0001, Muhua Zhu
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Refining History for Future-Aware Neural Machine Translation
abstract
Neural machine translation uses a decoder to generate target words auto-regressively by predicting the next target word conditioned on a given source sentence and its previously predicted target words, i.e, its translation history, which suffers from two limitations: 1) the prediction of next word depends heavily on the quality of its history information. Moreover, the discrepancy between training and inference exacerbates this limitation; 2) this left-to-right decoding way cannot make full use of the target-side future information, which leads to the issue of unbalanced outputs. On the one hand, we alleviate the first limitation with a history-refining module, which learns to examine the quality of each history word by assigning it a confidence score. The confidence score is further used as a gate to control the amount of its word embedding flowing to the decoder. On the other hand, we attack the second limitation with a future-foreseeing module, which learns the distribution of future translation at each decoding time step. More importantly, we further propose refining history for future-aware NMT since the two modules can be closely incorporated as they focus on different kinds of context. Experimental results on various translation tasks with different scaled datasets, including WMT English$\leftrightarrow${German, French, Romanian}, show that our proposed approach achieves significant improvements over strong Transformer-based NMT baselines.
Xinglin Lyu, Junhui Li 0001, Min Zhang 0005, Chenchen Ding, Hideki Tanaka, Masao Utiyama
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine Translation
abstract
In this paper we aim to relieve the issue of lexical translation inconsistency for documentlevel neural machine translation (NMT) by modeling consistency preference for lexical chains which consist of repeated words in a source-side document and provide a representation of the lexical consistency structure of the document.Specifically, we first propose lexical-consistency attention to capture consistency context among words in the same lexical chains.Then for each lexical chain we define and learn a consistency-tailored latent variable, which will guide the translation of corresponding sentences to enhance lexical translation consistency.Experimental results on Chinese→English and French→English document-level translation tasks show that our approach not only significantly improves translation performance in BLEU, but also substantially alleviates the problem of the lexical translation inconsistency.
Xinglin Lyu, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005
EMNLP2
2022 Contrastive Learning for Robust Neural Machine Translation with ASR Errors
Dongyang Hu, Junhui Li 0001
NLPCC (1)2
2022 Topic-Features for Dialogue Summarization
Junhui Li 0001
NLPCC (1)2
2022 One Type Context Is Not Enough: Global Context-aware Neural Machine Translation
abstract
How to effectively model global context has been a critical challenge for document-level neural machine translation (NMT). Both preceding and global context have been carefully explored in the sequence-to-sequence (seq2seq) framework. However, previous studies generally map global context into one vector, which is not enough to well represent the entire document since this largely ignores the hierarchy between sentences and words within. In this article, we propose to model global context for source language from both sentence level and word level. Specifically at sentence level, we extract useful global context for the current sentence, while at word level, we compute global context against words within the current sentence. On this basis, both kinds of global context can be appropriately fused before being incorporated into the state-of-the-art seq2seq model, i.e., Transformer . Detailed experimentation on various document-level translation tasks shows that global context at both sentence level and word level significantly improve translation performance. More encouraging, both kinds of global context are complementary. This leads to more improvement when both kinds of global context are used.
Linqing Chen, Junhui Li 0001, Zhengxian Gong, Min Zhang 0005, Guodong Zhou 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing
abstract
As an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus on multi-modal emotion recognition in a multi-label scenario. In this scenario, we consider not only the label-to-label dependency, but also the feature-to-label and modality-to-label dependencies. Particularly, we propose a heterogeneous hierarchical message passing network to effectively model above dependencies. Furthermore, we propose a new multi-modal multi-label emotion dataset based on partial time-series content to show predominant generalization of our model. Detailed evaluation demonstrates the effectiveness of our approach.
Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001
AAAI4
2021 Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-training
abstract
Linqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen, Weihua Luo, Min Zhang, Guodong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Linqing Chen, Junhui Li 0001, Zhengxian Gong, Boxing Chen, Weihua Luo, Min Zhang 0005, Guodong Zhou 0001
ACL/IJCNLP (1)2
2021 XLPT-AMR: Cross-Lingual Pre-Training via Multi-Task Learning for Zero-Shot AMR Parsing and Text Generation
abstract
Dongqin Xu, Junhui Li, Muhua Zhu, Min Zhang, Guodong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dongqin Xu, Junhui Li 0001, Muhua Zhu, Min Zhang 0005, Guodong Zhou 0001
ACL/IJCNLP (1)2
2021 Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation Detection
abstract
Aspect terms extraction (ATE) and aspect sentiment classification (ASC) are two fundamental and fine-grained sub-tasks in aspect-level sentiment analysis (ALSA).In the textual analysis, jointly extracting both aspect terms and sentiment polarities has been drawn much attention due to the better applications than individual sub-task.However, in the multimodal scenario, the existing studies are limited to handle each sub-task independently, which fails to model the innate connection between the above two objectives and ignores the better applications.Therefore, in this paper, we are the first to jointly perform multi-modal ATE (MATE) and multi-modal ASC (MASC), and we propose a multi-modal joint learning approach with auxiliary cross-modal relation detection for multi-modal aspect-level sentiment analysis (MALSA).Specifically, we first build an auxiliary text-image relation detection module to control the proper exploitation of visual information.Second, we adopt the hierarchical framework to bridge the multi-modal connection between MATE and MASC, as well as separately visual guiding for each sub module.Finally, we can obtain all aspect-level sentiment polarities dependent on the jointly extracted specific aspects.Extensive experiments show the effectiveness of our approach against the joint textual approaches, pipeline and collapsed multi-modal approaches.
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Shoushan Li, Min Zhang 0005, Guodong Zhou 0001
EMNLP (1)4
2021 Encouraging Lexical Translation Consistency for Document-Level Neural Machine Translation
abstract
Recently a number of approaches have been proposed to improve translation performance for document-level neural machine translation (NMT).However, few are focusing on the subject of lexical translation consistency.In this paper we apply "one translation per discourse" in NMT, and aim to encourage lexical translation consistency for document-level NMT.This is done by first obtaining a word link for each source word in a document, which tells the positions where the source word appears at.Then we encourage the translations of those words within a link to be consistent in two ways.On the one hand, when encoding sentences within a document we properly exchange context information of those words.On the other hand, we propose an auxiliary loss function to better constrain that their translations should be consistent.Experimental results on Chinese↔English and English→French translation tasks show that our approach not only achieves state-of-the-art performance in BLEU scores, but also greatly improves lexical translation consistency.
Xinglin Lyu, Junhui Li 0001, Zhengxian Gong, Min Zhang 0005
EMNLP (1)2
2021 Improving Context-Aware Neural Machine Translation with Source-side Monolingual Documents
abstract
Document context-aware machine translation remains challenging due to the lack of large-scale document parallel corpora. To make full use of source-side monolingual documents for context-aware NMT, we propose a Pre-training approach with Global Context (PGC). In particular, we first propose a novel self-supervised pre-training task, which contains two training objectives: (1) reconstructing the original sentence from a corrupted version; (2) generating a gap sentence from its left and right neighbouring sentences. Then we design a universal model for PGC which consists of a global context encoder, a sentence encoder and a decoder, with similar architecture to typical context-aware NMT models. We evaluate the effectiveness and generality of our pre-trained PGC model by adapting it to various downstream context-aware NMT models. Detailed experimentation on four different translation tasks demonstrates that our PGC approach significantly improves the translation performance of context-aware NMT. For example, based on the state-of-the-art SAN model, we achieve an averaged improvement of 1.85 BLEU scores and 1.59 Meteor scores on the four translation tasks.
Linqing Chen, Junhui Li 0001, Zhengxian Gong, Xiangyu Duan, Boxing Chen, Weihua Luo, Min Zhang 0005, Guodong Zhou 0001
IJCAI2
2021 Improving Text Generation with Dynamic Masking and Recovering
abstract
Due to different types of inputs, diverse text generation tasks may adopt different encoder-decoder frameworks. Thus most existing approaches that aim to improve the robustness of certain generation tasks are input-relevant, and may not work well for other generation tasks. Alternatively, in this paper we present a universal approach to enhance the language representation for text generation on the base of generic encoder-decoder frameworks. This is done from two levels. First, we introduce randomness by randomly masking some percentage of tokens on the decoder side when training the models. In this way, instead of using ground truth history context, we use its corrupted version to predict the next token. Then we propose an auxiliary task to properly recover those masked tokens. Experimental results on several text generation tasks including machine translation (MT), AMR-to-text generation, and image captioning show that the proposed approach can significantly improve over competitive baselines without using any task-specific techniques. This suggests the effectiveness and generality of our proposed approach.
Zhidong Liu, Junhui Li 0001, Muhua Zhu
IJCAI2
2021 Improving neural sentence alignment with word translation
Junhui Li 0001, Zhengxian Gong, Guodong Zhou 0001
Frontiers Comput. Sci.2
2021 Improving neural machine translation with latent features feedback
Yachao Li 0003, Junhui Li 0001, Min Zhang 0005
Neurocomputing2
2021 Deep Transformer modeling via grouping skip connection for neural machine translation
Yachao Li 0003, Junhui Li 0001, Min Zhang 0005
Knowl. Based Syst.2
2020 Improving AMR Parsing with Sequence-to-Sequence Pre-training
abstract
In the literature, the research on abstract meaning representation (AMR) parsing is much restricted by the size of human-curated dataset which is critical to build an AMR parser with good performance.To alleviate such data size restriction, pre-trained models have been drawing more and more attention in AMR parsing.However, previous pre-trained models, like BERT, are implemented for general purpose which may not work as expected for the specific task of AMR parsing.In this paper, we focus on sequence-to-sequence (seq2seq) AMR parsing and propose a seq2seq pre-training approach to build pre-trained models in both single and joint way on three relevant tasks, i.e., machine translation, syntactic parsing, and AMR parsing itself.Moreover, we extend the vanilla fine-tuning method to a multi-task learning fine-tuning method that optimizes for the performance of AMR parsing while endeavors to preserve the response of pre-trained models.Extensive experimental results on two English benchmark datasets show that both the single and joint pre-trained models significantly improve the performance (e.g., from 71.5 to 80.2 on AMR 2.0), which reaches the state of the art.The result is very encouraging since we achieve this with seq2seq models rather than complex models.We make our code and model available at https:// github.com/xdqkid/S2S-AMR-Parser.
Dongqin Xu, Junhui Li 0001, Muhua Zhu, Min Zhang 0005, Guodong Zhou 0001
EMNLP (1)2
2020 Multi-modal Multi-label Emotion Detection with Modality and Label Dependence
abstract
As an important research issue in the natural language processing community, multi-label emotion detection has been drawing more and more attention in the last few years. However, almost all existing studies focus on one modality (e.g., textual modality). In this paper, we focus on multi-label emotion detection in a multi-modal scenario. In this scenario, we need to consider both the dependence among different labels (label dependence) and the dependence between each predicting label and different modalities (modality dependence). Particularly, we propose a multi-modal sequence-to-set approach to effectively model both kinds of dependence in multi-modal multi-label emotion detection. The detailed evaluation demonstrates the effectiveness of our approach.
Dong Zhang 0013, Xincheng Ju, Junhui Li 0001, Shoushan Li, Qiaoming Zhu, Guodong Zhou 0001
EMNLP (1)3
2020 Transformer-based Label Set Generation for Multi-modal Multi-label Emotion Detection
abstract
Multi-modal utterance-level emotion detection has been a hot research topic in both multi-modal analysis and natural language processing communities. Different from traditional single-label multi-modal sentiment analysis, typical multi-modal emotion detection is naturally a multi-label problem where an utterance often contains multiple emotions. Existing studies normally focus on multi-modal fusion only and transform multi-label emotion classification into multiple binary classification problem independently. As a result, existing studies largely ignore two kinds of important dependency information: (1) Modality-to-label dependency, where different emotions can be inferred from different modalities, that is, different modalities contribute differently to each potential emotion. (2) Label-to-label dependency, where some emotions are more likely to coexist than those conflicting emotions. To simultaneously model above two kinds of dependency, we propose a unified approach, namely multi-modal emotion set generation network (MESGN) to generate an emotion set for an utterance. Specifically, we first employ a cross-modal transformer encoder to capture cross-modal interactions among different modalities, and a standard transformer encoder to capture temporal information for each modality-specific sequence given previous interactions. Then, we design a transformer-based discriminative decoding module equipped with modality-to-label attention to handle the modality-to-label dependency. In the meanwhile, we employ a reinforced decoding algorithm with self-critic learning to handle the label-to-label dependency. Finally, we validate the proposed MESGN architecture on a word-level aligned and unaligned multi-modal dataset. Detailed experimentation shows that our proposed MESGN architecture can effectively improve the performance of multi-modal multi-label emotion detection.
Xincheng Ju, Dong Zhang 0013, Junhui Li 0001, Guodong Zhou 0001
ACM Multimedia3
2020 Word-Pair Relevance Modeling with Multi-View Neural Attention Mechanism for Sentence Alignment
Junhui Li 0001, Zhengxian Gong, Guodong Zhou 0001
J. Comput. Sci. Technol.2
2020 Explicitly Modeling Word Translations in Neural Machine Translation
abstract
In this article, we show that word translations can be explicitly incorporated into NMT effectively to avoid wrong translations. Specifically, we propose three cross-lingual encoders to explicitly incorporate word translations into NMT: (1) Factored encoder, which encodes a word and its translation in a vertical way; (2) Gated encoder, which uses a gated mechanism to selectively control the amount of word translations moving forward; and (3) Mixed encoder, which stitchingly learns a word and its translation annotations over sequences where words and their translations are alternatively mixed. Besides, we first use a simple word dictionary approach and then a word sense disambiguation (WSD) approach to effectively model the word context for better word translation. Experimentation on Chinese-to-English translation demonstrates that all proposed encoders are able to improve the translation accuracy for both traditional RNN-based NMT and recent self-attention-based NMT (hereafter referred to as Transformer ). Specifically, Mixed encoder yields the most significant improvement of 2.0 in BLEU on the RNN-based NMT, while Gated encoder improves 1.2 in BLEU on Transformer . This indicates the usefulness of an WSD approach in modeling word context for better word translation. This also indicates the effectiveness of our proposed cross-lingual encoders in explicitly modeling word translations to avoid wrong translations in NMT. Finally, we discuss in depth how word translations benefit different NMT frameworks from several perspectives.
Junhui Li 0001, Yachao Li 0003, Min Zhang 0005, Guodong Zhou 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 Improving Neural Machine Translation with Linear Interpolation of a Short-Path Unit
abstract
In neural machine translation (NMT), the source and target words are at the two ends of a large deep neural network, normally mediated by a series of non-linear activations. The problem with such consequent non-linear activations is that they significantly decrease the magnitude of the gradient in a deep neural network, and thus gradually loosen the interaction between source words and their translations. As a result, a source word may be incorrectly translated into a target word out of its translational equivalents. In this article, we propose short-path units (SPUs) to strengthen the association of source and target words by allowing information flow over adjacent layers effectively via linear interpolation. In particular, we enrich three critical NMT components with SPUs: (1) an enriched encoding model with SPU, which interpolates source word embeddings linearly into source annotations; (2) an enriched decoding model with SPU, which enables the source context linearly flow to target-side hidden states; and (3) an enriched output model with SPU, which further allows linear interpolation of target-side hidden states into output states. Experimentation on Chinese-to-English, English-to-German, and low-resource Tibetan-to-Chinese translation tasks demonstrates that the linear interpolation of SPUs significantly improves the overall translation quality by 1.88, 1.43, and 3.75 BLEU, respectively. Moreover, detailed analysis shows that our approaches much strengthen the association of source and target words. From the preceding, we can see that our proposed model is effective both in rich- and low-resource scenarios.
Yachao Li 0003, Junhui Li 0001, Min Zhang 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Generating Multiple Diverse Responses for Short-Text Conversation
abstract
Neural generative models have become popular and achieved promising performance on short-text conversation tasks. They are generally trained to build a 1-to-1 mapping from the input post to its output response. However, a given post is often associated with multiple replies simultaneously in real applications. Previous research on this task mainly focuses on improving the relevance and informativeness of the top one generated response for each post. Very few works study generating multiple accurate and diverse responses for the same post. In this paper, we propose a novel response generation model, which considers a set of responses jointly and generates multiple diverse responses simultaneously. A reinforcement learning algorithm is designed to solve our model. Experiments on two short-text conversation tasks validate that the multiple responses generated by our model obtain higher quality and larger diversity compared with various state-ofthe-art generative models.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Shuming Shi 0001
AAAI4
2019 A Discrete CVAE for Response Generation on Short-Text Conversation
abstract
Jun Gao, Wei Bi, Xiaojiang Liu, Junhui Li, Guodong Zhou, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Guodong Zhou 0001, Shuming Shi 0001
EMNLP/IJCNLP (1)4
2019 Modeling Graph Structure in Transformer for Better AMR-to-Text Generation
abstract
Jie Zhu, Junhui Li, Muhua Zhu, Longhua Qian, Min Zhang, Guodong Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junhui Li 0001, Muhua Zhu, Longhua Qian, Min Zhang 0005, Guodong Zhou 0001
EMNLP/IJCNLP (1)2
2019 Modeling Source Syntax and Semantics for Neural AMR Parsing
abstract
Sequence-to-sequence (seq2seq) approaches formalize Abstract Meaning Representation (AMR) parsing as a translation task from a source sentence to a target AMR graph. However, previous studies generally model a source sentence as a word sequence but ignore the inherent syntactic and semantic information in the sentence. In this paper, we propose two effective approaches to explicitly modeling source syntax and semantics into neural seq2seq AMR parsing. The first approach linearizes source syntactic and semantic structure into a mixed sequence of words, syntactic labels, and semantic labels, while in the second approach we propose a syntactic and semantic structure-aware encoding scheme through a self-attentive model to explicitly capture syntactic and semantic relations between words. Experimental results on an English benchmark dataset show that our two approaches achieve significant improvement of 3.1% and 3.4% F1 scores over a strong seq2seq baseline.
DongLai Ge, Junhui Li 0001, Muhua Zhu, Shoushan Li
IJCAI2
2019 A Transformer-Based Semantic Parser for NLPCC-2019 Shared Task 2
DongLai Ge, Junhui Li 0001, Muhua Zhu
NLPCC (2)2
2018 Adaptive Weighting for Neural Machine Translation
abstract
In the popular sequence to sequence (seq2seq) neural machine translation (NMT), there exist many weighted sum models (WSMs), each of which takes a set of input and generates one output. However, the weights in a WSM are independent of each other and fixed for all inputs, suggesting that by ignoring different needs of inputs, the WSM lacks effective control on the influence of each input. In this paper, we propose adaptive weighting for WSMs to control the contribution of each input. Specifically, we apply adaptive weighting for both GRU and the output state in NMT. Experimentation on Chinese-to-English translation and English-to-German translation demonstrates that the proposed adaptive weighting is able to much improve translation accuracy by achieving significant improvement of 1.49 and 0.92 BLEU points for the two translation tasks. Moreover, we discuss in-depth on what type of information is encoded in the encoder and how information influences the generation of target words in the decoder.
Yachao Li 0003, Junhui Li 0001, Min Zhang 0005
COLING2
2018 Semi-supervised Sentiment Classification Based on Auxiliary Task Learning
Shoushan Li, Junhui Li 0001, Guodong Zhou 0001
NLPCC (2)4
2017 Modeling Source Syntax for Neural Machine Translation
abstract
Even though a linguistics-free sequence to sequence model in neural machine translation (NMT) has certain capability of implicitly learning syntactic information of source sentences, this paper shows that source syntax can be explicitly incorporated into NMT effectively to provide further improvements.Specifically, we linearize parse trees of source sentences to obtain structural label sequences.On the basis, we propose three different sorts of encoders to incorporate source syntax into NMT: 1) Parallel RNN encoder that learns word and label annotation vectors parallelly; 2) Hierarchical RNN encoder that learns word and label annotation vectors in a two-level hierarchy; and 3) Mixed RNN encoder that stitchingly learns word and label annotation vectors over sequences where words and labels are mixed.Experimentation on Chinese-to-English translation demonstrates that all the three proposed syntactic encoders are able to improve translation accuracy.It is interesting to note that the simplest RNN encoder, i.e., Mixed RNN encoder yields the best performance with an significant improvement of 1.4 BLEU points.Moreover, an in-depth analysis from several perspectives is provided to reveal how source syntax benefits NMT.
Junhui Li 0001, Deyi Xiong, Zhaopeng Tu, Muhua Zhu, Min Zhang 0005, Guodong Zhou 0001
ACL (1)1
2016 Improving Semantic Parsing with Enriched Synchronous Context-Free Grammars in Statistical Machine Translation
abstract
Semantic parsing maps a sentence in natural language into a structured meaning representation. Previous studies show that semantic parsing with synchronous context-free grammars (SCFGs) achieves favorable performance over most other alternatives. Motivated by the observation that the performance of semantic parsing with SCFGs is closely tied to the translation rules, this article explores to extend translation rules with high quality and increased coverage in three ways. First, we examine the difference between word alignments for semantic parsing and statistical machine translation (SMT) to better adapt word alignment in SMT to semantic parsing. Second, we introduce both structure and syntax informed nonterminals, better guiding the parsing in favor of well-formed structure, instead of using a uninformed nonterminal in SCFGs. Third, we address the unknown word translation issue via synthetic translation rules. Last but not least, we use a filtering approach to improve performance via predicting answer type. Evaluation on the standard GeoQuery benchmark dataset shows that our approach greatly outperforms the state of the art across various languages, including English, Chinese, Thai, German, and Greek.
Junhui Li 0001, Muhua Zhu, Wei Lu 0011, Guodong Zhou 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2015 Improving Semantic Parsing with Enriched Synchronous Context-Free Grammar
abstract
Semantic parsing maps a sentence in natural language into a structured meaning representation.Previous studies show that semantic parsing with synchronous contextfree grammars (SCFGs) achieves favorable performance over most other alternatives.Motivated by the observation that the performance of semantic parsing with SCFGs is closely tied to the translation rules, this paper explores extending translation rules with high quality and increased coverage in three ways.First, we introduce structure informed non-terminals, better guiding the parsing in favor of well formed structure, instead of using a uninformed non-terminal in SCFGs.Second, we examine the difference between word alignments for semantic parsing and statistical machine translation (SMT) to better adapt word alignment in SMT to semantic parsing.Finally, we address the unknown word translation issue via synthetic translation rules.Evaluation on the standard GeoQuery benchmark dataset shows that our approach achieves the state-of-the-art across various languages, including English, German and Greek.
Junhui Li 0001, Muhua Zhu, Wei Lu 0011, Guodong Zhou 0001
EMNLP1
2011 Tree kernel-based semantic role labeling with enriched parse tree structure
Guodong Zhou 0001, Junhui Li 0001, Jianxi Fan, Qiaoming Zhu
Inf. Process. Manag.2
2011 Unified Semantic Role Labeling for Verbal and Nominal Predicates in the Chinese Language
abstract
This article explores unified semantic role labeling (SRL) for both verbal and nominal predicates in the Chinese language. This is done by considering SRL for both verbal and nominal predicates in a unified framework. First, we systematically examine various kinds of features for verbal SRL and nominal SRL, respectively, besides those widely used ones. Then we further improve the performance of nominal SRL with various kinds of verbal evidence, that is, merging the training instances from verbal predicates and integrating various kinds of features derived from SRL for verbal predicates. Finally, we address the issue of automatic predicate recognition, which is essential for nominal SRL. Evaluation on Chinese PropBank and Chinese NomBank shows that our unified approach significantly improves the performance, in particular that of nominal SRL. To the best of our knowledge, this is the first reported work of unified verbal and nominal SRL on Chinese PropBank and NomBank.
Junhui Li 0001, Guodong Zhou 0001
ACM Trans. Asian Lang. Inf. Process.1
2010 Joint Syntactic and Semantic Parsing of Chinese
Junhui Li 0001, Guodong Zhou 0001, Hwee Tou Ng
ACL1
2010 Learning the Scope of Negation via Shallow Semantic Parsing
Junhui Li 0001, Guodong Zhou 0001, Hongling Wang, Qiaoming Zhu
COLING1
2010 A Unified Framework for Scope Learning via Simplified Shallow Semantic Parsing
Qiaoming Zhu, Junhui Li 0001, Hongling Wang, Guodong Zhou 0001
EMNLP2
2009 Improving Nominal SRL in Chinese Language with Verbal SRL Information and Automatic Predicate Recognition
Junhui Li 0001, Guodong Zhou 0001, Hai Zhao 0001, Qiaoming Zhu, Peide Qian
EMNLP1
2008 Semi-Supervised Learning for Relation Extraction
Guodong Zhou 0001, Junhui Li 0001, Longhua Qian, Qiaoming Zhu
IJCNLP2