VLDB 2026 Research / reviewers in the wild / expert
Jiajun Zhang 0001
dblp:71/6950-1
· DBLP profile ↗
113ranked-venue papers
13as first author
37since 2021 · last 2026
0000-0001-5293-7434ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 108 · 12 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOEP: Task-Specific Operator Evolution via Multi-objective Pareto Optimization for Automatic Workflow Generation
Chunlin Leng, Xiaomian Kang, Jiajun Zhang 0001 |
ICPR (14) | 5 |
| 2025 | LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMsabstractLong-context modeling has drawn more and more attention in the area of Large Language Models (LLMs).Continual training with long-context data becomes the de-facto method to equip LLMs with the ability to process long inputs.However, it still remains an open challenge to measure the quality of long-context training data.To address this issue, we propose a Long-context data selection framework with Attention-based Dependency Measurement (LADM), which can efficiently identify high-quality long-context data from a large-scale, multi-domain pre-training corpus.LADM leverages the retrieval capabilities of the attention mechanism to capture contextual dependencies, ensuring a comprehensive quality measurement of long-context data.Experimental results show that our LADM framework significantly boosts the performance of LLMs on multiple long-context tasks with only 1B tokens for continual training.1 Jianghao Chen, Junhong Wu, Yangyifan Xu, Jiajun Zhang 0001 |
ACL (1) | 4 |
| 2025 | Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsabstractIn visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits.Existing research addresses such ambiguities primarily by rephrasing questions.These approaches neglect the inherently interactive nature of user interactions with VLMs, where ambiguities can be clarified through user feedback.However, research on interactive clarification faces two major challenges: (1) Benchmarks are absent to assess VLMs' capacity for resolving ambiguities through interaction; (2) VLMs are trained to prefer answering rather than asking, preventing them from seeking clarification.To overcome these challenges, we introduce ClearVQA benchmark 1 , which targets three common categories of ambiguity in VQA context, and encompasses various VQA scenarios.Furthermore, we propose an automated pipeline to generate ambiguity-clarification question pairs.Experimental results demonstrate that training based on the automated generated data enables VLMs to ask reasonable clarification questions, thereby generating more accurate and specific answers based on user feedback. Pu Jian, Donglei Yu, Shuo Ren 0002, Jiajun Zhang 0001 |
ACL (1) | 5 |
| 2025 | Hit the Sweet Spot! Span-Level Ensemble for Large Language ModelsabstractEnsembling various LLMs to unlock their complementary potential and leverage their individual strengths is highly valuable. Previous studies typically focus on two main paradigms: sample-level and token-level ensembles. Sample-level ensemble methods either select or blend fully generated outputs, which hinders dynamic correction and enhancement of outputs during the generation process. On the other hand, token-level ensemble methods enable real-time correction through fine-grained ensemble at each generation step. However, the information carried by an individual token is quite limited, leading to suboptimal decisions at each step. To address these issues, we propose SweetSpan, a span-level ensemble method that effectively balances the need for real-time adjustments and the information required for accurate ensemble decisions. Our approach involves two key steps: First, we have each candidate model independently generate candidate spans based on the shared prefix. Second, we calculate perplexity scores to facilitate mutual evaluation among the candidate models and achieve robust span selection by filtering out unfaithful scores. To comprehensively evaluate ensemble methods, we propose a new challenging setting (ensemble models with significant performance gaps) in addition to the standard setting (ensemble the best-performing models) to assess the performance of model ensembles in more realistic scenarios. Experimental results in both standard and challenging settings across various language generation tasks demonstrate the effectiveness, robustness, and versatility of our approach compared with previous ensemble methods. Yangyifan Xu, Jianghao Chen, Junhong Wu, Jiajun Zhang 0001 |
COLING | 4 |
| 2025 | Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language ModelsabstractRecent advances in text-only "slow-thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs), for training visual reasoning models (VRMs).However, such transfer faces critical challenges: Effective "slow thinking" in VRMs requires visual reflection, the ability to check the reasoning process based on visual information.Through quantitative analysis, we observe that current VRMs exhibit limited visual reflection, as their attention to visual information diminishes rapidly with longer generated responses.To address this challenge, we propose a new VRM Reflection-V 1 , which enhances visual reflection based on reasoning data construction for cold-start and reward design for reinforcement learning (RL).Firstly, we construct vision-centered reasoning data by leveraging an agent that interacts between VLMs and reasoning LLMs, enabling cold-start learning of visual reflection patterns.Secondly, a visual attention based reward model is employed during RL to encourage reasoning based on visual information.Therefore, Reflection-V demonstrates significant improvements across multiple visual reasoning benchmarks.Furthermore, Reflection-V maintains a stronger and more consistent reliance on visual information during visual reasoning, indicating effective enhancement in visual reflection capabilities. Pu Jian, Junhong Wu, Shuo Ren 0002, Jiajun Zhang 0001 |
EMNLP | 6 |
| 2025 | Collaborative Beam Search: Enhancing LLM Reasoning via Collective ConsensusabstractComplex multi-step reasoning remains challenging for large language models (LLMs).While parallel inference-time scaling methods, such as step-level beam search, offer a promising solution, existing approaches typically depend on either domain-specific external verifiers, or self-evaluation which is brittle and prompt-sensitive.To address these issues, we propose Collaborative Beam Search (CBS), an iterative framework that harnesses the collective intelligence of multiple LLMs across both generation and verification stages.For generation, CBS leverages multiple LLMs to explore a broader search space, resulting in more diverse candidate steps.For verifications, CBS employs a perplexity-based collective consensus among these models, eliminating reliance on an external verifier or complex prompts.Between iterations, CBS leverages a dynamic quota allocation strategy that reassigns generation budget based on each model's past performance, striking a balance between candidate diversity and quality.Experimental results on six tasks across arithmetic, logical, and commonsense reasoning show that CBS outperforms single-model scaling and multi-model ensemble baselines by over 4 percentage points in average accuracy, demonstrating its effectiveness and general applicability. Yangyifan Xu, Shuo Ren 0002, Jiajun Zhang 0001 |
EMNLP | 3 |
| 2025 | Language Imbalance Driven Rewarding for Multilingual Self-improvingabstractLarge Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such as English and Chinese, leaving many other languages underrepresented. This imbalance, while limiting broader applications, generates a natural preference ranking between languages, offering an opportunity to bootstrap the multilingual capabilities of LLM in a self-improving manner. Thus, we propose $\textit{Language Imbalance Driven Rewarding}$, where the inherent imbalance between dominant and non-dominant languages within LLMs is leveraged as a reward signal. Iterative DPO training demonstrates that this approach not only enhances LLM performance in non-dominant languages but also improves the dominant language's capacity, thereby yielding an iterative reward signal. Fine-tuning Meta-Llama-3-8B-Instruct over two iterations of this approach results in continuous improvements in multilingual performance across instruction-following and arithmetic reasoning tasks, evidenced by an average improvement of 7.46\% win rate on the X-AlpacaEval leaderboard and 13.9\% accuracy on the MGSM benchmark. This work serves as an initial exploration, paving the way for multilingual self-improvement of LLMs. Junhong Wu, Chengqing Zong, Jiajun Zhang 0001 |
ICLR | 5 |
| 2025 | KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical ReasoningabstractRecent advances have demonstrated that integrating reinforcement learning with rule-based rewards can significantly enhance the reasoning capabilities of large language models (LLMs), even without supervised fine-tuning (SFT). However, prevalent reinforcement learning algorithms such as GRPO and its variants like DAPO, suffer from a coarse granularity issue when computing the advantage. Specifically, they compute rollout-level advantages that assign identical values to every token within a sequence, failing to capture token-specific contributions. To address this limitation, we propose Key-token Advantage Estimation (KTAE)—a novel algorithm that estimates fine-grained, token-level advantages without introducing additional models. KTAE leverages the correctness of sampled rollouts and applies statistical analysis to quantify the importance of individual tokens within a sequence to the final outcome. This quantified token-level importance is then combined with the rollout-level advantage to obtain a more fine-grained token-level advantage estimation. Empirical results show that models trained with GRPO+KTAE and DAPO+KTAE outperform baseline methods across five mathematical reasoning benchmarks. Notably, they achieve higher accuracy with shorter responses and even surpass R1-Distill-Qwen-1.5B using the same base model. Pu Jian, Qianlong Du, Fuwei Cui, Shuo Ren 0002, Jiajun Zhang 0001 |
NeurIPS | 7 |
| 2025 | Beyond One-Size-Fits-All: Adaptive Fine-Tuning for LLMs Based on Data Inherent Heterogeneity
Wanyue Zhang, Yangyifan Xu, Shuo Ren 0002, Jiajun Zhang 0001 |
NLPCC (1) | 4 |
| 2024 | Improving Unsupervised Neural Machine Translation via Training Data Self-CorrectionabstractUnsupervised neural machine translation (UNMT) models are trained with pseudo-parallel sentences constructed by on-the-fly back-translation using monolingual corpora. However, the quality of pseudo-parallel sentences cannot be guaranteed, which hinders the final performance of UNMT. This paper demonstrates that although UNMT usually generates mistakes during pseudo-parallel data construction, some of them can be corrected by the token-level translations that exist in the embedding table. Therefore, we propose a self-correction method to automatically improve the quality of pseudo-parallel sentences during training, thereby enhancing translation performance. Specifically, for a pseudo sentence pair, our self-correction method first estimates the alignment relations between tokens by treating and solving it as an optimal transport problem. Then, we measure the translation reliability for each token and detect the mis-translated ones. Finally, the mis-translated tokens are corrected with real-time computed token-by-token translations based on the embedding table, yielding a better training example. Considering that the modified examples are semantically equivalent to the original ones when UNMT converges, we introduce second-phase training to strengthen the output consistency between them, further improving the generalization capability and translation performance. Empirical results on widely used UNMT datasets demonstrate the effectiveness of our method and it significantly outperforms several strong baselines. Jinliang Lu, Jiajun Zhang 0001 |
LREC/COLING | 2 |
| 2024 | Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAabstractVisual question answering (VQA) tasks, often performed by visual language model (VLM), face challenges with long-tail knowledge.Recent retrieval-augmented VQA (RA-VQA) systems address this by retrieving and integrating external knowledge sources.However, these systems still suffer from redundant visual information irrelevant to the question during retrieval.To address these issues, in this paper, we propose LLM-RA , a novel method leveraging the reasoning capability of a large language model (LLM) to identify key visual entities, thus minimizing the impact of irrelevant information in the query of retriever.Furthermore, key visual entities are independently encoded for multimodal joint retrieval, preventing cross-entity interference.Experimental results demonstrate that our method outperforms other strong RA-VQA systems.In two knowledge-intensive VQA benchmarks, our method achieves the new state-of-the-art performance among those with similar scale of parameters and even performs comparably to models with 1-2 orders larger parameters. Pu Jian, Donglei Yu, Jiajun Zhang 0001 |
EMNLP | 3 |
| 2024 | BLSP-Emo: Towards Empathetic Large Speech-Language ModelsabstractThe recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions.While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible.In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speechlanguage model capable of understanding both semantics and emotions in speech and generate empathetic responses.BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process.The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data.The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data.Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations. 1 Minpeng Liao, Zhongqiang Huang, Junhong Wu, Chengqing Zong, Jiajun Zhang 0001 |
EMNLP | 6 |
| 2024 | Knowledge Graph Guided Neural Machine Translation with Dynamic Reinforce-selected TriplesabstractPrevious methods incorporating knowledge graphs (KGs) into neural machine translation (NMT) adopt a static knowledge utilization strategy, that introduces many useless knowledge triples and makes the useful triples difficult to be utilized by NMT. To address this problem, we propose a KG guided NMT model with dynamic reinforce-selected triples. The proposed methods could dynamically select the different useful knowledge triples for different source sentences. Specifically, the proposed model contains two components: (1) knowledge selector, that dynamically selects useful knowledge triples for a source sentence, and (2) knowledge guided NMT (KgNMT), that utilizes the selected triples to guide the translation of NMT. Meanwhile, to overcome the non-differentiable problem and guide the training procedure, we propose a policy gradient strategy to encourage the model to select useful triples and improve the generation probability of gold target sentence. Various experimental results show that the proposed method can significantly outperform the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Xiaomian Kang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsabstractFine-tuning all parameters of large language models (LLMs) requires significant computational resources and is time-consuming.Recent parameter-efficient tuning methods such as Adapter tuning, Prefix tuning, and LoRA allow updating a small subset of parameters in large language models.However, they can only save approximately 30% of the training memory requirements because gradient computation and backpropagation are still necessary for these methods.This paper proposes a novel parameter-efficient tuning method for LLMs without calculating their gradients.Leveraging the discernible similarities between the parameter-efficient modules of the same task learned by both large and small language models, we put forward a strategy for transferring the parameter-efficient modules derived initially from small language models to much larger ones.To ensure a smooth and effective adaptation process, we introduce a Bridge model to guarantee dimensional consistency while stimulating a dynamic interaction between the models.We demonstrate the effectiveness of our method using the T5 and GPT-2 series of language models on the SuperGLUE benchmark.Our method achieves comparable performance to fine-tuning and parameterefficient tuning on large language models without needing gradient-based optimization.Additionally, our method achieves up to 5.7× memory reduction compared to parameter-efficient tuning. Feihu Jin, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 2 |
| 2023 | Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningabstractTransformer-based models, even though achieving super-human performance on several downstream tasks, are often regarded as a black box and used as a whole.It is still unclear what mechanisms they have learned, especially their core module: multi-head attention.Inspired by functional specialization in the human brain, which helps to efficiently handle multiple tasks, this work attempts to figure out whether the multi-head attention module will evolve similar function separation under multitasking training.If it is, can this mechanism further improve the model performance?To investigate these questions, we introduce an interpreting method to quantify the degree of functional specialization in multi-head attention.We further propose a simple multi-task training method to increase functional specialization and mitigate negative information transfer in multi-task learning.Experimental results on seven pre-trained transformer models have demonstrated that multi-head attention does evolve functional specialization phenomenon after multi-task training which is affected by the similarity of tasks.Moreover, the multi-task training strategy based on functional specialization boosts performance in both multi-task learning and transfer learning without adding any parameters. 1 Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 4 |
| 2023 | Unified Prompt Learning Makes Pre-Trained Language Models Better Few-Shot LearnersabstractLanguage prompting induces the model to produce a textual output during the training phase, which achieves remarkable performance in few-shot learning scenarios. However, current prompt-based methods either use the same task-specific prompts for each instance, losing the particularity of instance-dependent information, or generate an instance-dependent prompt for each instance, lacking shared information about the task. In this paper, we propose an efficient few-shot learning method to dynamically decide the degree to which task-specific and instance-dependent information are incorporated according to different task and instance characteristics, enriching the prompt with task-specific and instance-dependent information. Extensive experiments on a wide range of natural language understanding tasks demonstrate that our approach obtains significant improvements compared to prompt-based fine-tuning baselines in a few-shot setting with about 0.1% parameters tuned. Moreover, our approach outperforms existing state-of-the-art efficient few-shot learning methods on several natural language understanding tasks. Feihu Jin, Jinliang Lu, Jiajun Zhang 0001 |
ICASSP | 3 |
| 2023 | Adapter Tuning With Task-Aware Attention MechanismabstractAdapter-tuning inserts simple feed-forward layers (adapters) in pre-trained language models (PLMs) and just tunes the adapters when transferring to downstream tasks, having become the state-of-the-art parameter-efficient tuning (PET) strategy. Although the adapters aim to learn task-related representations, their inputs are still obtained from the task-independent and frozen multi-head attention (MHA) modules, leading to insufficient utilization of contextual information for various downstream tasks. Intuitively, MHA should be task-dependent and could attend to different contexts in different downstream tasks. Thus, this paper proposes the task-aware attention mechanism (TAM) to enhance adapter tuning. Specifically, we first utilize the task-dependent adapter to generate token-wise task embedding. Then, we apply the task embedding to influence MHA which task-dependently aggregates the contextual information. Experimental results on a wide range of natural language understanding and generation tasks demonstrate the effectiveness of our method. Furthermore, extensive analyses demonstrate that the generated task embedding corresponds with the difficulty of tasks. Jinliang Lu, Feihu Jin, Jiajun Zhang 0001 |
ICASSP | 3 |
| 2023 | Contrastive Adversarial Training for Multi-Modal Machine TranslationabstractThe multi-modal machine translation task is to improve translation quality with the help of additional visual input. It is expected to disambiguate or complement semantics while there are ambiguous words or incomplete expressions in the sentences. Existing methods have tried many ways to fuse visual information into text representations. However, only a minority of sentences need extra visual information as complementary. Without guidance, models tend to learn text-only translation from the major well-aligned translation pairs. In this article, we propose a contrastive adversarial training approach to enhance visual participation in semantic representation learning. By contrasting multi-modal input with the adversarial samples, the model learns to identify the most informed sample that is coupled with a congruent image and several visual objects extracted from it. This approach can prevent the visual information from being ignored and further fuse cross-modal information. We examine our method in three multi-modal language pairs. Experimental results show that our model is capable of improving translation accuracy. Further analysis shows that our model is more sensitive to visual information. Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Instance-Aware Prompt Learning for Language Understanding and GenerationabstractPrompt learning has emerged as a new paradigm for leveraging pre-trained language models (PLMs) and has shown promising results in downstream tasks with only a slight increase in parameters. However, the current usage of fixed prompts, whether discrete or continuous, assumes that all samples within a task share the same prompt. This assumption may not hold for tasks with diverse samples that require different prompt information. To address this issue, we propose an instance-aware prompt learning method that learns a different prompt for each instance. Specifically, we suppose that each learnable prompt token has a different contribution to different instances, and we learn the contribution by calculating the relevance score between an instance and each prompt token. The contribution-weighted prompt would be instance aware. We apply our method to both unidirectional and bidirectional PLMs on both language understanding and generation tasks. Extensive experiments demonstrate that our method achieves comparable results using as few as 1.5% of the parameters of PLMs tuned and obtains considerable improvements compared with strong baselines. In particular, our method achieves state-of-the-art results using ALBERT-xxlarge-v2 on the SuperGLUE few-shot learning benchmark. 1 Feihu Jin, Jinliang Lu, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Topic-Oriented Dialogue SummarizationabstractA multi-turn dialogue often contains multiple discussion topics. In several scenarios (e.g., customer service dispute, public opinion monitoring), people are only interested in the gist of a specific topic in the dialogue. Therefore, we propose a novel summarization task, i.e., Topic-Oriented Dialogue Summarization (TODS). Given a dialogue with a topic label, TODS aims to produce a summary covering the main content of the given topic in the dialogue. To model the relationship between dialogues and topics, three key abilities are needed for TODS: (1) Learning the semantic information of different topics. (2) Locating the topic-related content in the dialogue. (3) Distinguishing summaries for different topics in the same dialogue. Thus, we propose three topic-related auxiliary tasks to make the summarization model learn the three abilities above. First, the topic identification task aims at generating all the topics in the dialogue. Second, the topic attention restriction task tries to constrain the attention distribution on topic-related utterances. Third, the topic summary distinguishing task focuses on increasing the difference of summaries for different topics in the same dialogue. Experimental results on two public TODS datasets show that all auxiliary tasks are critical for TODS and help generate high-quality summaries. We also point out the expansions and challenges in TODS for future research. Haitao Lin 0001, Junnan Zhu, Lu Xiang, Feifei Zhai, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Towards Unified Multi-Domain Machine Translation With Mixture of Domain ExpertsabstractMulti-domain machine translation (MDMT) aims to construct models with mixed-domain training corpora to switch translation between different domains. Previous studies either assume that the domain information is given and leverage the domain knowledge to guide the translation process, or suppose that the domain information is unknown and utilize the model to automatically recognize it. However, the cases are mixed in practical scenarios, which means that some sentences are labeled with domain information while others are unlabeled, which is beyond the capacity of the previous methods. In this paper, we propose a unified MDMT model with a mixture of sub-networks (experts) to address the cases with or without domain labels. The mixture of sub-networks in our MDMT model includes a shared expert and multiple domain-specific experts. For the inputs with domain labels, our MDMT model goes through the shared and the corresponding domain-specific experts. For the unlabeled inputs, our MDMT model activates all the experts, each of which makes a dynamic contribution. Experimental results on multiple diverse domains in De$\rightarrow$En, Fr$\rightarrow$En, and En$\rightarrow$Ro demonstrate that our method can outperform the strong baselines in both scenarios with or without domain labels. Further analyses show that our model has good generalization ability when transferring into new domains. Jinliang Lu, Jiajun Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Probing Word Syntactic Representations in the Brain by a Feature Elimination MethodabstractNeuroimaging studies have identified multiple brain regions that are associated with semantic and syntactic processing when comprehending language. However, existing methods cannot explore the neural correlates of fine-grained word syntactic features, such as part-of-speech and dependency relations. This paper proposes an alternative framework to study how different word syntactic features are represented in the brain. To separate each syntactic feature, we propose a feature elimination method, called Mean Vector Null space Projection (MVNP). This method can remove a specific feature from word representations, resulting in one-feature-removed representations. Then we respectively associate one-feature-removed and the original word vectors with brain imaging data to explore how the brain represents the removed feature. This paper for the first time studies the cortical representations of multiple fine-grained syntactic features simultaneously and suggests some possible contributions of several brain regions to the complex division of syntactic processing. These findings indicate that the brain foundations of syntactic information processing might be broader than those suggested by classical studies. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 4 |
| 2022 | Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role InteractionsabstractRole-oriented dialogue summarization is to generate summaries for different roles in the dialogue, e.g., merchants and consumers.Existing methods handle this task by summarizing each role's content separately and thus are prone to ignore the information from other roles.However, we believe that other roles' content could benefit the quality of summaries, such as the omitted information mentioned by other roles.Therefore, we propose a novel role interaction enhanced method for role-oriented dialogue summarization.It adopts cross attention and decoder self-attention interactions to interactively acquire other roles' critical information.The cross attention interaction aims to select other roles' critical dialogue utterances, while the decoder self-attention interaction aims to obtain key information from other roles' summaries.Experimental results have shown that our proposed method significantly outperforms strong baselines on two public role-oriented dialogue summarization datasets.Extensive analyses have demonstrated that other roles' content could help generate summaries with more complete semantics and correct topic structures. 1 Haitao Lin 0001, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 5 |
| 2022 | Learning Confidence for Transformer-based Neural Machine TranslationabstractConfidence estimation aims to quantify the confidence of the model prediction, providing an expectation of success.A well-calibrated confidence estimate enables accurate failure prediction and proper risk measurement when given noisy samples and out-of-distribution data in real-world settings.However, this task remains a severe challenge for neural machine translation (NMT), where probabilities from softmax distribution fail to describe when the model is probably mistaken.To address this problem, we propose an unsupervised confidence estimate learning jointly with the training of the NMT model.We explain confidence as how many hints the NMT model needs to make a correct prediction, and more hints indicate low confidence.Specifically, the NMT model is given the option to ask for hints to improve translation accuracy at the cost of some slight penalty.Then, we approximate their level of confidence by counting the number of hints the model uses.We demonstrate that our learned confidence estimate achieves high accuracy on extensive sentence/word-level quality estimation tasks.Analytical results verify that our confidence estimate can correctly assess underlying risk in two real-world scenarios: (1) discovering noisy samples and (2) detecting out-of-domain data.We further propose a novel confidence-based instance-specific label smoothing approach based on our learned confidence estimate, which outperforms standard label smoothing 1 . Jiali Zeng, Jiajun Zhang 0001, Shuangzhi Wu, Mu Li 0001 |
ACL (1) | 3 |
| 2022 | Norm-based Noisy Corpora Filtering and Refurbishing in Neural Machine TranslationabstractRecent advances in neural machine translation depend on massive parallel corpora, which are collected from any open source without much guarantee of quality.It stresses the need for noisy corpora filtering, but existing methods are insufficient to solve this issue.They spend much time ensembling multiple scorers trained on clean bitexts, unavailable for low-resource languages in practice.In this paper, we propose a norm-based noisy corpora filtering and refurbishing method with no external data and costly scorers.The noisy and clean samples are separated based on how much information from the source and target sides the model requires to fit the given translation.For the unparallel sentence, the target-side history translation is much more important than the source context, contrary to the parallel ones.The amount of these two information flows can be measured by norms of source-/target-side context vectors.Moreover, we propose to reuse the discovered noisy data by generating pseudo labels via online knowledge distillation.Extensive experiments show that our proposed filtering method performs comparably with state-ofthe-art noisy corpora filtering techniques but is more efficient and easier to operate.Noisy sample refurbishing further enhances the performance by making the most of the given data 1 . Jiajun Zhang 0001 |
EMNLP | 2 |
| 2022 | Discrete Cross-Modal Alignment Enables Zero-Shot Speech TranslationabstractEnd-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions.However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain.Fortunately, the supervised data for automatic speech recognition (ASR) and machine translation (MT) are usually more accessible, making zero-shot speech translation a potential direction.Existing zero-shot methods fail to align the two modalities of speech and text into a shared semantic space, resulting in much worse performance compared to the supervised ST methods.In order to enable zero-shot ST, we propose a novel Discrete Cross-Modal Alignment (DCMA) method that employs a shared discrete vocabulary space to accommodate and match both modalities of speech and text.Specifically, we introduce a vector quantization module to discretize the continuous representations of speech and text into a finite set of virtual tokens, and use ASR data to map corresponding speech and text to the same virtual token in a shared codebook.This way, source language speech can be embedded in the same semantic space as the source language text, which can be then transformed into target language text with an MT module.Experiments on multiple language pairs demonstrate that our zero-shot ST method significantly improves the SOTA, and even performs on par with the strong supervised ST baselines 1 . Yuchen Liu 0007, Boxing Chen, Jiajun Zhang 0001, Zhongqiang Huang, Chengqing Zong |
EMNLP | 4 |
| 2022 | Enhancing Lexical Translation Consistency for Document-Level Neural Machine TranslationabstractDocument-level neural machine translation (DocNMT) has yielded attractive improvements. In this article, we systematically analyze the discourse phenomena in Chinese-to-English translation, and focus on the most obvious ones, namely lexical translation consistency. To alleviate the lexical inconsistency, we propose an effective approach that is aware of the words which need to be translated consistently and constrains the model to produce more consistent translations. Specifically, we first introduce a global context extractor to extract the document context and consistency context, respectively. Then, the two types of global context are integrated into a encoder enhancer and a decoder enhancer to improve the lexical translation consistency. We create a test set to evaluate the lexical consistency automatically. Experiments demonstrate that our approach can significantly alleviate the lexical translation inconsistency. In addition, our approach can also substantially improve the translation quality compared to sentence-level Transformer. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | Dual-View Conditional Variational Auto-Encoder for Emotional Dialogue GenerationabstractEmotional dialogue generation aims to generate appropriate responses that are content relevant with the query and emotion consistent with the given emotion tag. Previous work mainly focuses on incorporating emotion information into the sequence to sequence or conditional variational auto-encoder (CVAE) models, and they usually utilize the given emotion tag as a conditional feature to influence the response generation process. However, emotion tag as a feature cannot well guarantee the emotion consistency between the response and the given emotion tag. In this article, we propose a novel Dual-View CVAE model to explicitly model the content relevance and emotion consistency jointly. These two views gather the emotional information and the content-relevant information from the latent distribution of responses, respectively. We jointly model the dual-view via VAE to get richer and complementary information. Extensive experiments on both English and Chinese emotion dialogue datasets demonstrate the effectiveness of our proposed Dual-View CVAE model, which significantly outperforms the strong baseline models in both aspects of content relevance and emotion consistency. Jiajun Zhang 0001, Lu Xiang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Attention Analysis and Calibration for Transformer in Natural Language GenerationabstractAttention mechanism has been ubiquitous in neural machine translation by dynamically selecting relevant contexts for different translations. Apart from performance gains, attention weights assigned to input tokens are often utilized to explain that high-attention tokens contribute more to the prediction. However, many works question whether this assumption holds in text classification by manually manipulating attention weights and observing decision flips. This article extends this question to Transformer-based neural machine translation, which heavily relies on cross-lingual attention to produce accurate translations but is relatively understudied in this context. We first design a mask perturbation model which automatically assesses each input’s contribution to model outputs. We then test whether the token contributing most to the current translation receives the highest attention weight. We find that it sometimes does not, which closely depends on the entropy of attention weights, the syntactic role of the current generation, and language pairs. We also rethink the discrepancy between attention weights and word alignments from the view of unreliable attention weights. Our observations further motivate us to calibrate the cross-lingual multi-head attention by attaching more attention to indispensable tokens, whose removal leads to a dramatic performance drop. Empirical experiments on different-scale translation tasks and text summarization tasks demonstrate that our calibration methods significantly outperform strong baselines. Jiajun Zhang 0001, Jiali Zeng, Shuangzhi Wu, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Synchronous Inference for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation allows a single model to translate between multiple language pairs, which greatly reduces the cost of model training and receives much attention recently. Previous studies mainly focus on training stage optimization and improve positive knowledge transfer among languages with different levels of parameter sharing, but ignore the multilingual knowledge transfer during inference although the translation in one language may help the generation of other languages. This work enhances knowledge sharing among multiple target languages in the inference phase. To achieve this, we propose a synchronous inference method that can simultaneously generate translations in multiple languages. During generation, the model predicts the next word of each language not only based on source sentence and previously predicted segments, but also based on predicted words of other target languages. To maximize the inference stage knowledge sharing, we design a cross-lingual attention module which allows the model to dynamically select the most relevant information from multiple target languages. The synchronous inference model requires multi-way parallel training data which is scarce. We therefore propose to adopt multi-task learning to incorporate large-scale bilingual data. We evaluate our method on three multilingual translation datasets and prove that the proposed method significantly improve the translation quality and the decoding efficiency compared to strong bilingual and multilingual baselines. Qian Wang 0061, Jiajun Zhang 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Synchronous Interactive Decoding for Multilingual Neural Machine TranslationabstractTo simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future information, and therefore may suffer from some issues like the unbalanced outputs. In this paper, we present a new approach for synchronous interactive multilingual neural machine translation (SimNMT), which predicts each target language output simultaneously and interactively using historical and future information of all target languages. Specifically, we first propose a synchronous cross-interactive decoder in which generation of each target output does not only depend on its generated sequences, but also relies on its future information, as well as history and future contexts of other target languages. Then, we present a new interactive multilingual beam search algorithm that enables synchronous interactive decoding of all target languages in a single model. We take two target languages as an example to illustrate and evaluate the proposed SimNMT model on IWSLT datasets. The experimental results demonstrate that our method achieves significant improvements over several advanced NMT and MNMT models. Qian Wang 0061, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 5 |
| 2021 | Attention Calibration for Transformer in Neural Machine TranslationabstractYu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu, Mu Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiali Zeng, Jiajun Zhang 0001, Shuangzhi Wu, Mu Li 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue SummarizationabstractDialogue summarization has drawn much attention recently.Especially in the customer service domain, agents could use dialogue summaries to help boost their works by quickly knowing customer's issues and service progress.These applications require summaries to contain the perspective of a single speaker and have a clear topic flow structure, while neither are available in existing datasets.Therefore, in this paper, we introduce a novel Chinese dataset for Customer Service Dialogue Summarization (CSDS).CSDS improves the abstractive summaries in two aspects: (1) In addition to the overall summary for the whole dialogue, role-oriented summaries are also provided to acquire different speakers' viewpoints.(2) All the summaries sum up each topic separately, thus containing the topic-level structure of the dialogue.We define tasks in CSDS as generating the overall summary and different role-oriented summaries for a given dialogue.Next, we compare various summarization methods on CSDS, and experiment results show that existing methods are prone to generate redundant and incoherent summaries.Besides, the performance becomes much worse when analyzing the performance on role-oriented summaries and topic structures.We hope that this study could benchmark Chinese dialogue summarization and benefit further studies. Haitao Lin 0001, Liqun Ma, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
EMNLP (1) | 6 |
| 2021 | Augmenting Slot Values and Contexts for Spoken Language Understanding with Pretrained ModelsabstractSpoken Language Understanding (SLU) is one essential step in building a dialogue system. Due to the expensive cost of obtaining the labeled data, SLU suffers from the data scarcity problem. Therefore, in this paper, we focus on data augmentation for slot filling task in SLU. To achieve that, we aim at generating more diverse data based on existing data. Specifically, we try to exploit the latent language knowledge from pretrained language models by finetuning them. We propose two strategies for finetuning process: value-based and context-based augmentation. Experimental results on two public SLU datasets have shown that compared with existing data augmentation methods, our proposed method can generate more diverse sentences and significantly improve the performance on SLU. Haitao Lin 0001, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
Interspeech | 4 |
| 2021 | Chinese Spelling Error Detection Using a Fusion Lattice LSTMabstractSpelling error detection serves as a crucial preprocessing in many natural language processing applications. Unlike English, where every single word is directly typed by keyboard, we have to use an input method to input Chinese characters. The pinyin input method is the most widely used. By intuition, pinyin should be helpful in detecting spelling errors. However, when detect spelling errors, most of the current methods ignore the pinyin information and adopt a pipeline framework that leads to error propagation. In this article, we propose a fusion lattice-LSTM model under the end-to-end framework to integrate character, word, and pinyin features for error detection. Experiments on the SIGHAN Bake-off-2015 dataset show that pinyin is a discriminating feature, and our end-to-end model outperforms the baseline models obviously. Hao Wang 0018, Jianyong Duan, Jiajun Zhang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | Graph-based Multimodal Ranking Models for Multimodal SummarizationabstractMultimodal summarization aims to extract the most important information from the multimedia input. It is becoming increasingly popular due to the rapid growth of multimedia data in recent years. There are various researches focusing on different multimodal summarization tasks. However, the existing methods can only generate single-modal output or multimodal output. In addition, most of them need a lot of annotated samples for training, which makes it difficult to be generalized to other tasks or domains. Motivated by this, we propose a unified framework for multimodal summarization that can cover both single-modal output summarization and multimodal output summarization. In our framework, we consider three different scenarios and propose the respective unsupervised graph-based multimodal summarization models without the requirement of any manually annotated document-summary pairs for training: (1) generic multimodal ranking, (2) modal-dominated multimodal ranking, and (3) non-redundant text-image multimodal ranking. Furthermore, an image-text similarity estimation model is introduced to measure the semantic similarity between image and text. Experiments show that our proposed models outperform the single-modal summarization methods on both automatic and human evaluation metrics. Besides, our models can also improve the single-modal summarization with the guidance of the multimedia information. This study can be applied as the benchmark for further study on multimodal summarization task. Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | Neural Encoding and Decoding With Distributed Sentence RepresentationsabstractBuilding computational models to account for the cortical representation of language plays an important role in understanding the human linguistic system. Recent progress in distributed semantic models (DSMs), especially transformer-based methods, has driven advances in many language understanding tasks, making DSM a promising methodology to probe brain language processing. DSMs have been shown to reliably explain cortical responses to word stimuli. However, characterizing the brain activities for sentence processing is much less exhaustively explored with DSMs, especially the deep neural network-based methods. What is the relationship between cortical sentence representations against DSMs? What linguistic features that a DSM catches better explain its correlation with the brain activities aroused by sentence stimuli? Could distributed sentence representations help to reveal the semantic selectivity of different brain areas? We address these questions through the lens of neural encoding and decoding, fueled by the latest developments in natural language representation learning. We begin by evaluating the ability of a wide range of 12 DSMs to predict and decipher the functional magnetic resonance imaging (fMRI) images from humans reading sentences. Most models deliver high accuracy in the left middle temporal gyrus (LMTG) and left occipital complex (LOC). Notably, encoders trained with transformer-based DSMs consistently outperform other unsupervised structured models and all the unstructured baselines. With probing and ablation tasks, we further find that differences in the performance of the DSMs in modeling brain activities can be at least partially explained by the granularity of their semantic representations. We also illustrate the DSM's selectivity for concept categories and show that the topics are represented by spatially overlapping and distributed cortical patterns. Our results corroborate and extend previous findings in understanding the relation between DSMs and neural activation patterns and contribute to building solid brain-machine interfaces with deep neural network representations. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Keywords-Guided Abstractive Sentence SummarizationabstractWe study the problem of generating a summary for a given sentence. Existing researches on abstractive sentence summarization ignore that keywords in the input sentence provide significant clues for valuable content, and humans tend to write summaries covering these keywords. In this paper, we propose an abstractive sentence summarization method by applying guidance signals of keywords to both the encoder and the decoder in the sequence-to-sequence model. A multi-task learning framework is adopted to jointly learn to extract keywords and generate a summary for the input sentence. We apply keywords-guided selective encoding strategies to filter source information by investigating the interactions between the input sentence and the keywords. We extend pointer-generator network by a dual-attention and a dual-copy mechanism, which can integrate the semantics of the input sentence and the keywords, and copy words from both the input sentence and the keywords. We demonstrate that multi-task learning and keywords-oriented guidance facilitate sentence summarization task, achieving better performance than the competitive models on the English Gigaword sentence summarization dataset. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong, Xiaodong He 0001 |
AAAI | 3 |
| 2020 | Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingabstractSpeech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential benefits of lower latency, smaller model size, and less error propagation. However, it is notoriously difficult to implement such a model without transcriptions as intermediate. Existing works generally apply multi-task learning to improve translation quality by jointly training end-to-end ST along with automatic speech recognition (ASR). However, different tasks in this method cannot utilize information from each other, which limits the improvement. Other works propose a two-stage model where the second model can use the hidden state from the first one, but its cascade manner greatly affects the efficiency of training and inference process. In this paper, we propose a novel interactive attention mechanism which enables ASR and ST to perform synchronously and interactively in a single model. Specifically, the generation of transcriptions and translations not only relies on its previous outputs but also the outputs predicted in the other task. Experiments on TED speech translation corpora have shown that our proposed model can outperform strong baselines on the quality of speech translation and achieve better speech recognition performances as well. Yuchen Liu 0007, Jiajun Zhang 0001, Hao Xiong 0005, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
AAAI | 2 |
| 2020 | Probing Brain Activation Patterns by Dissociating Semantics and Syntax in SentencesabstractThe relation between semantics and syntax and where they are represented in the neural level has been extensively debated in neurosciences. Existing methods use manually designed stimuli to distinguish semantic and syntactic information in a sentence that may not generalize beyond the experimental setting. This paper proposes an alternative framework to study the brain representation of semantics and syntax. Specifically, we embed the highly-controlled stimuli as objective functions in learning sentence representations and propose a disentangled feature representation model (DFRM) to extract semantic and syntactic information in sentences. This model can generate one semantic and one syntactic vector for each sentence. Then we associate these disentangled feature vectors with brain imaging data to explore brain representation of semantics and syntax. Results have shown that semantic feature is represented more robustly than syntactic feature across the brain including the default-mode, frontoparietal, visual networks, etc.. The brain representations of semantics and syntax are largely overlapped, but there are brain regions only sensitive to one of them. For instance, several frontal and temporal regions are specific to the semantic feature; parts of the right superior frontal and right inferior parietal gyrus are specific to the syntactic feature. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 2 |
| 2020 | Multimodal Summarization with Guidance of Multimodal ReferenceabstractMultimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Order-ranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments. Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Haoran Li 0001, Chengqing Zong, Changliang Li |
AAAI | 3 |
| 2020 | Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual SummarizationabstractCross-lingual summarization aims at summarizing a document in one language (e.g., Chinese) into another language (e.g., English).In this paper, we propose a novel method inspired by the translation pattern in the process of obtaining a cross-lingual summary.We first attend to some words in the source text, then translate them into the target language, and summarize to get the final summary.Specifically, we first employ the encoder-decoder attention distribution to attend to the source words.Second, we present three strategies to acquire the translation probability, which helps obtain the translation candidates for each source word.Finally, each summary word is generated either from the neural distribution or from the translation candidates of source words.Experimental results on Chinese-to-English and English-to-Chinese summarization tasks have shown that our proposed method can significantly outperform the baselines, achieving comparable performance with the state-of-the-art. Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL | 3 |
| 2020 | Multimodal Sentence Summarization via Multimodal Selective EncodingabstractThis paper studies the problem of generating a summary for a given sentence-image pair.Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news event or a document.Thus, we propose a multimodal selective gate network that considers reciprocal relationships between textual and multi-level visual features, including global image descriptor, activation grids, and object proposals, to select highlights of the event when encoding the source sentence.In addition, we introduce a modality regularization to encourage the summary to capture the highlights embedded in the image more accurately.To verify the generalization of our model, we adopt the multimodal selective gate to the text-based decoder and multimodal-based decoder.Experimental results on a public multimodal sentence summarization dataset demonstrate the advantage of our models over baselines.Further analysis suggests that our proposed multimodal selective gate network can effectively select important information in the input sentence. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Xiaodong He 0001, Chengqing Zong |
COLING | 3 |
| 2020 | Distill and Replay for Continual Language LearningabstractAccumulating knowledge to tackle new tasks without necessarily forgetting the old ones is a hallmark of human-like intelligence.But the current dominant paradigm of machine learning is still to train a model that works well on static datasets.When learning tasks in a stream where data distribution may fluctuate, fitting on new tasks often leads to forgetting on the previous ones.We propose a simple yet effective framework that continually learns natural language understanding tasks with one model.Our framework distills knowledge and replays experience from previous tasks when fitting on a new task, thus named DnR (distill and replay).The framework is based on language models and can be smoothly built with different language model architectures.Experimental results demonstrate that DnR outperfoms previous state-of-the-art models in continually learning tasks of the same type but from different domains, as well as tasks of radically different types.With the distillation method, we further show that it's possible for DnR to incrementally compress the model size while still outperforming most of the baselines.We hope that DnR could promote the empirical application of continual language learning, and contribute to building human-level language intelligence minimally bothered by catastrophic forgetting. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
COLING | 3 |
| 2020 | Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity GranularityabstractPrevious studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized.ii) Granularity mismatch: the current KG methods utilize the entity as the basic granularity, while NMT utilizes the sub-word as the granularity, making the KG different to be utilized in NMT.To alleviate above problems, we propose a multi-task learning method on sub-entity granularity.Specifically, we first split the entities in KG and sentence pairs into sub-entity granularity by using joint BPE.Then we utilize the multi-task learning to combine the machine translation task and knowledge reasoning task.The extensive experiments on various translation tasks have demonstrated that our method significantly outperforms the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Lu Xiang, Junnan Zhu, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 4 |
| 2020 | Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningabstractDocument-level neural machine translation has yielded attractive improvements.However, majority of existing methods roughly use all context sentences in a fixed scope.They neglect the fact that different source sentences need different sizes of context.To address this problem, we propose an effective approach to select dynamic context so that the document-level translation model can utilize the more useful selected context sentences to produce better translations.Specifically, we introduce a selection module that is independent of the translation module to score each candidate context sentence.Then, we propose two strategies to explicitly select a variable number of context sentences and feed them into the translation module.We train the two modules end-to-end via reinforcement learning.A novel reward is proposed to encourage the selection and utilization of dynamic context sentences.Experiments demonstrate that our approach can select adaptive context sentences for different source sentences, and significantly improves the performance of document-level translation methods. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
EMNLP (1) | 3 |
| 2020 | Knowledge Graphs Enhanced Neural Machine TranslationabstractKnowledge graphs (KGs) store much structured information on various entities, many of which are not covered by the parallel sentence pairs of neural machine translation (NMT). To improve the translation quality of these entities, in this paper we propose a novel KGs enhanced NMT method. Specifically, we first induce the new translation results of these entities by transforming the source and target KGs into a unified semantic space. We then generate adequate pseudo parallel sentence pairs that contain these induced entity pairs. Finally, NMT model is jointly trained by the original and pseudo sentence pairs. The extensive experiments on Chinese-to-English and Englishto-Japanese translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling the induced entities. Yang Zhao 0007, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IJCAI | 2 |
| 2020 | Non-autoregressive Neural Machine Translation with Distortion Model
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
NLPCC (1) | 2 |
| 2020 | Synchronous bidirectional inference for neural sequence generation
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
Artif. Intell. | 1 |
| 2020 | Fine-grained neural decoding with distributed word representations
Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
Inf. Sci. | 2 |
| 2020 | Structurally Comparative Hinge Loss for Dependency-Based Neural Text RepresentationabstractDependency-based graph convolutional networks (DepGCNs) are proven helpful for text representation to handle many natural language tasks. Almost all previous models are trained with cross-entropy (CE) loss, which maximizes the posterior likelihood directly. However, the contribution of dependency structures is not well considered by CE loss. As a result, the performance improvement gained by using the structure information can be narrow due to the failure in learning to rely on this structure information. To face the challenge, we propose the novel structurally comparative hinge (SCH) loss function for DepGCNs. SCH loss aims at enlarging the margin gained by structural representations over non-structural ones. From the perspective of information theory, this is equivalent to improving the conditional mutual information of model decision and structure information given text. Our experimental results on both English and Chinese datasets show that by substituting SCH loss for CE loss on various tasks, for both induced structures and structures from an external parser, performance is improved without additional learnable parameters. Furthermore, the extent to which certain types of examples rely on the dependency structure can be measured directly by the learned margin, which results in better interpretability. In addition, through detailed analysis, we show that this structure margin has a positive correlation with task performance and structure induction of DepGCNs, and SCH loss can help model focus more on the shortest dependency path between entities. We achieve the new state-of-the-art results on TACRED, IMDB, and Zh. Literature datasets, even compared with ensemble and BERT baselines. Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Deep Neural Network-based Machine Translation System CombinationabstractDeep neural networks (DNNs) have provably enhanced the state-of-the-art natural language process (NLP) with their capability of feature learning and representation. As one of the more challenging NLP tasks, neural machine translation (NMT) becomes a new approach to machine translation and generates much more fluent results compared to statistical machine translation (SMT). However, SMT is usually better than NMT in translation adequacy and word coverage. It is therefore a promising direction to combine the advantages of both NMT and SMT. In this article, we propose a deep neural network--based system combination framework leveraging both minimum Bayes-risk decoding and multi-source NMT, which take as input the N-best outputs of NMT and SMT systems and produce the final translation. In particular, we apply the proposed model to both RNN and self-attention networks with different segmentation granularity. We verify our approach empirically through a series of experiments on resource-rich Chinese⇒English and low-resource English⇒Vietnamese translation tasks. Experimental results demonstrate the effectiveness and universality of our proposed approach, which significantly outperforms the conventional system combination methods and the best individual system output. Jiajun Zhang 0001, Xiaomian Kang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Towards Sentence-Level Brain Decoding with Distributed RepresentationsabstractDecoding human brain activities based on linguistic representations has been actively studied in recent years. However, most previous studies exclusively focus on word-level representations, and little is learned about decoding whole sentences from brain activation patterns. This work is our effort to mend the gap. In this paper, we build decoders to associate brain activities with sentence stimulus via distributed representations, the currently dominant sentence representation approach in natural language processing (NLP). We carry out a systematic evaluation, covering both widely-used baselines and state-of-the-art sentence representation models. We demonstrate how well different types of sentence representations decode the brain activation patterns and give empirical explanations of the performance difference. Moreover, to explore how sentences are neurally represented in the brain, we further compare the sentence representation’s correspondence to different brain areas associated with high-level cognitive functions. We find the supervised structured representation models most accurately probe the language atlas of human brain. To the best of our knowledge, this work is the first comprehensive evaluation of distributed sentence representations for brain decoding. We hope this work can contribute to decoding brain activities with NLP representation models, and understanding how linguistic items are neurally represented. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 3 |
| 2019 | Addressing the Under-Translation Problem from the Entropy PerspectiveabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance in recent years. However, the under-translation problem still remains a big challenge. In this paper, we focus on the under-translation problem and attempt to find out what kinds of source words are more likely to be ignored. Through analysis, we observe that a source word with a large translation entropy is more inclined to be dropped. To address this problem, we propose a coarse-to-fine framework. In coarse-grained phase, we introduce a simple strategy to reduce the entropy of highentropy words through constructing the pseudo target sentences. In fine-grained phase, we propose three methods, including pre-training method, multitask method and two-pass method, to encourage the neural model to correctly translate these high-entropy words. Experimental results on various translation tasks show that our method can significantly improve the translation quality and substantially reduce the under-translation cases of high-entropy words. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhongjun He, Hua Wu 0003 |
AAAI | 2 |
| 2019 | Memory Consolidation for Contextual Spoken Language Understanding with Dialogue Logistic InferenceabstractDialogue contexts are proven helpful in the spoken language understanding (SLU) system and they are typically encoded with explicit memory representations.However, most of the previous models learn the context memory with only one objective to maximizing the SLU performance, leaving the context memory under-exploited.In this paper, we propose a new dialogue logistic inference (DLI) task to consolidate the context memory jointly with SLU in the multi-task framework.DLI is defined as sorting a shuffled dialogue session into its original logical order and shares the same memory encoder and retrieval mechanism as the SLU model.Our experimental results show that various popular contextual SLU models can benefit from our approach, and improvements are quite impressive, especially in slot filling. He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 3 |
| 2019 | Incremental Learning from Scratch for Task-Oriented Dialogue SystemsabstractClarifying user needs is essential for existing task-oriented dialogue systems.However, in real-world applications, developers can never guarantee that all possible user demands are taken into account in the design phase.Consequently, existing systems will break down when encountering unconsidered user needs.To address this problem, we propose a novel incremental learning framework to design task-oriented dialogue systems, or for short Incremental Dialogue System (IDS), without pre-defining the exhaustive list of user needs.Specifically, we introduce an uncertainty estimation module to evaluate the confidence of giving correct responses.If there is high confidence, IDS will provide responses to users.Otherwise, humans will be involved in the dialogue process, and IDS can learn from human intervention through an online learning module.To evaluate our method, we propose a new dataset which simulates unanticipated user needs in the deployment stage.Experiments show that IDS is robust to unconsidered user actions, and can update itself online by smartly selecting only the most effective training data, and hence attains better performance with less annotation cost. 1 Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong, Zhifei Li 0001 |
ACL (1) | 2 |
| 2019 | A Compact and Language-Sensitive Multilingual Translation MethodabstractMultilingual neural machine translation (Multi-NMT) with one encoder-decoder model has made remarkable progress due to its simple deployment.However, this multilingual translation paradigm does not make full use of language commonality and parameter sharing between encoder and decoder.Furthermore, this kind of paradigm cannot outperform the individual models trained on bilingual corpus in most cases.In this paper, we propose a compact and language-sensitive method for multilingual translation.To maximize parameter sharing, we first present a universal representor to replace both encoder and decoder models.To make the representor sensitive for specific languages, we further introduce language-sensitive embedding, attention, and discriminator with the ability to enhance model performance.We verify our methods on various translation scenarios, including one-to-many, many-to-many and zero-shot.Extensive experiments demonstrate that our proposed methods remarkably outperform strong standard multilingual translation systems on WMT and IWSLT datasets.Moreover, we find that our model is especially helpful in low-resource and zero-shot translation scenarios. Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong |
ACL (1) | 3 |
| 2019 | Are You for Real? Detecting Identity Fraud via Dialogue InteractionsabstractWeikang Wang, Jiajun Zhang, Qian Li, Chengqing Zong, Zhifei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiajun Zhang 0001, Chengqing Zong, Zhifei Li 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Synchronously Generating Two Languages with Interactive DecodingabstractYining Wang, Jiajun Zhang, Long Zhou, Yuchen Liu, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiajun Zhang 0001, Yuchen Liu 0007, Chengqing Zong |
EMNLP/IJCNLP (1) | 2 |
| 2019 | NCLS: Neural Cross-Lingual SummarizationabstractJunnan Zhu, Qian Wang, Yining Wang, Yu Zhou, Jiajun Zhang, Shaonan Wang, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junnan Zhu, Qian Wang 0061, Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Sequence Generation: From Both Sides to the MiddleabstractThe encoder-decoder framework has achieved promising process for many sequence generation tasks, such as neural machine translation and text summarization. Such a framework usually generates a sequence token by token from left to right, hence (1) this autoregressive decoding procedure is time-consuming when the output sentence becomes longer, and (2) it lacks the guidance of future context which is crucial to avoid under-translation. To alleviate these issues, we propose a synchronous bidirectional sequence generation (SBSG) model which predicts its outputs from both sides to the middle simultaneously. In the SBSG model, we enable the left-to-right (L2R) and right-to-left (R2L) generation to help and interact with each other by leveraging interactive bidirectional attention network. Experiments on neural machine translation (En-De, Ch-En, and En-Ro) and text summarization tasks show that the proposed model significantly speeds up decoding while improving the generation quality compared to the autoregressive Transformer. Jiajun Zhang 0001, Chengqing Zong, Heng Yu 0006 |
IJCAI | 2 |
| 2019 | End-to-End Speech Translation with Knowledge DistillationabstractEnd-to-end speech translation (ST), which directly translates from source language speech into target language text, has attracted intensive attentions in recent years.Compared to conventional pipepine systems, end-to-end ST models have advantages of lower latency, smaller model size and less error propagation.However, the combination of speech recognition and text translation in one model is more difficult than each of these two tasks.In this paper, we propose a knowledge distillation approach to improve ST model by transferring the knowledge from text translation model.Specifically, we first train a text translation model, regarded as a teacher model, and then ST model is trained to learn output probabilities from teacher model through knowledge distillation.Experiments on English-French Augmented LibriSpeech and English-Chinese TED corpus show that end-to-end ST is possible to implement on both similar and dissimilar language pairs.In addition, with the instruction of teacher model, end-to-end ST model can gain significant improvements by over 3.5 BLEU points. Yuchen Liu 0007, Hao Xiong 0005, Jiajun Zhang 0001, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
INTERSPEECH | 3 |
| 2019 | Select the Best Translation from Different Systems Without Reference
Jinliang Lu, Jiajun Zhang 0001 |
NLPCC (1) | 2 |
| 2019 | Synchronous Bidirectional Neural Machine TranslationabstractAbstract Existing approaches to neural machine translation (NMT) generate the target language sequence token-by-token from left to right. However, this kind of unidirectional decoding framework cannot make full use of the target-side future contexts which can be produced in a right-to-left decoding direction, and thus suffers from the issue of unbalanced outputs. In this paper, we introduce a synchronous bidirectional–neural machine translation (SB-NMT) that predicts its outputs using left-to-right and right-to-left decoding simultaneously and interactively, in order to leverage both of the history and future information at the same time. Specifically, we first propose a new algorithm that enables synchronous bidirectional decoding in a single model. Then, we present an interactive decoding model in which left-to-right (right-to-left) generation does not only depend on its previously generated outputs, but also relies on future contexts predicted by right-to-left (left-to-right) decoding. We extensively evaluate the proposed SB-NMT model on large-scale NIST Chinese-English, WMT14 English-German, and WMT18 Russian-English translation tasks. Experimental results demonstrate that our model achieves significant improvements over the strong Transformer model by 3.92, 1.49, and 1.04 BLEU points, respectively, and obtains the state-of-the-art per- formance on Chinese-English and English- German translation tasks.1 Jiajun Zhang 0001, Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Input Method for Human Translators: A Novel Approach to Integrate Machine Translation Effectively and ImperceptiblyabstractComputer-aided translation (CAT) systems are the most popular tool for helping human translators efficiently perform language translation. To further improve the translation efficiency, there is an increasing interest in applying machine translation (MT) technology to upgrade CAT. To thoroughly integrate MT into CAT systems, in this article, we propose a novel approach: a new input method that makes full use of the knowledge adopted by MT systems, such as translation rules, decoding hypotheses, and n-best translation lists. The proposed input method contains two parts: a phrase generation model, allowing human translators to type target sentences quickly, and an n-gram prediction model, helping users choose perfect MT fragments smoothly. In addition, to tune the underlying MT system to generate the input method preferable results, we design a new evaluation metric for the MT system. The proposed input method integrates MT effectively and imperceptibly, and it is particularly suitable for many target languages with complex characters, such as Chinese and Japanese. The extensive experiments demonstrate that our method saves more than 23% in time and over 42% in keystrokes, and it also improves the translation quality by more than 5 absolute BLEU scores compared with the strong baseline, i.e., post-editing using Google Pinyin. Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Experience-based Causality Learning for Intelligent AgentsabstractUnderstanding causality in text is crucial for intelligent agents. In this article, inspired by human causality learning, we propose an experience-based causality learning framework. Comparing to traditional approaches, which attempt to handle the causality problem relying on textual clues and linguistic resources, we are the first to use experience information for causality learning. Specifically, we first construct various scenarios for intelligent agents, thus, the agents can gain experience from interaction in these scenarios. Then, human participants build a number of training instances for agents of causality learning based on these scenarios. Each instance contains two sentences and a label. Each sentence describes an event that an agent experienced in a scenario, and the label indicates whether the sentence (event) pair has a causal relation. Accordingly, we propose a model that can infer the causality in text using experience by accessing the corresponding event information based on the input sentence pair. Experiment results show that our method can achieve impressive performance on the grounded causality corpus and significantly outperform the conventional approaches. Our work suggests that experience is very important for intelligent agents to understand causality. Yang Liu 0085, Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Attention With Sparsity Regularization for Neural Machine Translation and SummarizationabstractThe attention mechanism has become thede factostandard component in neural sequence to sequence tasks, such as machine translation and abstractive summarization. It dynamically determines which parts in the input sentence should be focused on when generating each word in the output sequence. Ideally, only few relevant input words should be attended to at each decoding time step and the attention weight distribution should be sparse and sharp. However, previous methods have no good mechanism to control this attention weight distribution. In this paper, we propose a sparse attention model in which a sparsity regularization term is designed to augment the objective function. We explore two kinds of regularizations:$L_{\infty }$-norm regularization and minimum entropy regularization, both of which aim to sharpen the attention weight distribution. Extensive experiments on both neural machine translation and abstractive summarization demonstrate that our proposed sparse attention model can substantially outperform the strong baselines. And the detailed analyses reveal that the final attention distribution indeed becomes sparse and sharp. Jiajun Zhang 0001, Yang Zhao 0007, Haoran Li 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and VideoabstractAutomatic text summarization is a fundamental natural language processing (NLP) application that aims to condense a source text into a shorter version. The rapid increase in multimedia data transmission over the Internet necessitates multi-modal summarization (MMS) from asynchronous collections of text, image, audio, and video. In this work, we propose an extractive MMS method that unites the techniques of NLP, speech processing, and computer vision to explore the rich information contained in multi-modal data and to improve the quality of multimedia news summarization. The key idea is to bridge the semantic gaps between multi-modal content. Audio and visual are main modalities in the video. For audio information, we design an approach to selectively use its transcription and to infer the salience of the transcription with audio signals. For visual information, we learn the joint representations of text and images using a neural network. Then, we capture the coverage of the generated summary for important visual information through text-image matching or multi-modal topic modeling. Finally, all the multi-modal aspects are considered to generate a textual summary by maximizing the salience, non-redundancy, readability, and coverage through the budgeted optimization of submodular functions. We further introduce a publicly available MMS corpus in English and Chinese.1 The experimental results obtained on our dataset demonstrate that our methods based on image matching and image topic framework outperform other competitive baseline methods. Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Investigating Inner Properties of Multimodal Representation and Semantic Compositionality With Brain-Based Componential SemanticsabstractMultimodal models have been proven to outperform text-based approaches on learning semantic representations. However, it still remains unclear what properties are encoded in multimodal representations, in what aspects do they outperform the single-modality representations, and what happened in the process of semantic compositionality in different input modalities. Considering that multimodal models are originally motivated by human concept representations, we assume that correlating multimodal representations with brain-based semantics would interpret their inner properties to answer the above questions. To that end, we propose simple interpretation methods based on brain-based componential semantics. First we investigate the inner properties of multimodal representations by correlating them with corresponding brain-based property vectors. Then we map the distributed vector space to the interpretable brain-based componential space to explore the inner properties of semantic compositionality. Ultimately, the present paper sheds light on the fundamental questions of natural language understanding, such as how to represent the meaning of words and how to combine word meanings into larger units. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 2 |
| 2018 | Learning Multimodal Word Representation via Dynamic Fusion MethodsabstractMultimodal models have been proven to outperform text-based models on learning semantic word representations. Almost all previous multimodal models typically treat the representations from different modalities equally. However, it is obvious that information from different modalities contributes differently to the meaning of words. This motivates us to build a multimodal model that can dynamically fuse the semantic representations from different modalities according to different types of words. To that end, we propose three novel dynamic fusion methods to assign importance weights to each modality, in which weights are learned under the weak supervision of word association pairs. The extensive experiments have demonstrated that the proposed methods outperform strong unimodal baselines and state-of-the-art multimodal models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 2 |
| 2018 | Source Critical Reinforcement Learning for Transferring Spoken Language Understanding to a New LanguageabstractTo deploy a spoken language understanding (SLU) model to a new language, language transferring is desired to avoid the trouble of acquiring and labeling a new big SLU corpus. An SLU corpus is a monolingual corpus with domain/intent/slot labels. Translating the original SLU corpus into the target language is an attractive strategy. However, SLU corpora consist of plenty of semantic labels (slots), which general-purpose translators cannot handle well, not to mention additional culture differences. This paper focuses on the language transferring task given a small in-domain parallel SLU corpus. The in-domain parallel corpus can be used as the first adaptation on the general translator. But more importantly, we show how to use reinforcement learning (RL) to further adapt the adapted translator, where translated sentences with more proper slot tags receive higher rewards. Our reward is derived from the source input sentence exclusively, unlike reward via actor-critical methods or computing reward with a ground truth target sentence. Hence we can adapt the translator the second time, using the big monolingual SLU corpus from the source language. We evaluate our approach on Chinese to English language transferring for SLU systems. The experimental results show that the generated English SLU corpus via adaptation and reinforcement learning gives us over 97% in the slot F1 score and over 84% accuracy in domain classification. It demonstrates the effectiveness of the proposed language transferring method. Compared with naive translation, our proposed method improves domain classification accuracy by relatively 22%, and the slot filling F1 score by relatively more than 71%. He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong |
COLING | 3 |
| 2018 | Ensure the Correctness of the Summary: Incorporate Entailment Knowledge into Abstractive Sentence SummarizationabstractIn this paper, we investigate the sentence summarization task that produces a summary from a source sentence. Neural sequence-to-sequence models have gained considerable success for this task, while most existing approaches only focus on improving the informativeness of the summary, which ignore the correctness, i.e., the summary should not contain unrelated information with respect to the source sentence. We argue that correctness is an essential requirement for summarization systems. Considering a correct summary is semantically entailed by the source sentence, we incorporate entailment knowledge into abstractive summarization models. We propose an entailment-aware encoder under multi-task framework (i.e., summarization generation and entailment recognition) and an entailment-aware decoder by entailment Reward Augmented Maximum Likelihood (RAML) training. Experiment results demonstrate that our models significantly outperform baselines from the aspects of informativeness and correctness. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong |
COLING | 3 |
| 2018 | Associative Multichannel Autoencoder for Multimodal Word RepresentationabstractIn this paper we address the problem of learning multimodal word representations by integrating textual, visual and auditory inputs.Inspired by the re-constructive and associative nature of human memory, we propose a novel associative multichannel autoencoder (AMA).Our model first learns the associations between textual and perceptual modalities, so as to predict the missing perceptual information of concepts.Then the textual and predicted perceptual representations are fused through reconstructing their original and associated embeddings.Using a gating mechanism our model assigns different weights to each modality according to the different concepts.Results on six benchmark concept similarity tests show that the proposed method significantly outperforms strong unimodal baselines and state-of-the-art multimodal models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 2 |
| 2018 | A Teacher-Student Framework for Maintainable Dialog ManagerabstractReinforcement learning (RL) is an attractive solution for task-oriented dialog systems.However, extending RL-based systems to handle new intents and slots requires a system redesign.The high maintenance cost makes it difficult to apply RL methods to practical systems on a large scale.To address this issue, we propose a practical teacherstudent framework to extend RL-based dialog systems without retraining from scratch.Specifically, the "student" is an extended dialog manager based on a new ontology, and the "teacher" is existing resources used for guiding the learning process of the "student".By specifying constraints held in the new dialog manager, we transfer knowledge of the "teacher" to the "student" without additional resources.Experiments show that the performance of the extended system is comparable to the system trained from scratch.More importantly, the proposed framework makes no assumption about the unsupported intents and slots, which makes it possible to improve RL-based systems incrementally.U: I'm looking for a Sichuan restaurant.S: "Spicy Little Girl" Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong, Zhifei Li 0001 |
EMNLP | 2 |
| 2018 | Three Strategies to Improve One-to-Many Multilingual TranslationabstractDue to the benefits of model compactness, multilingual translation (including many-toone, many-to-many and one-to-many) based on a universal encoder-decoder architecture attracts more and more attention.However, previous studies show that one-to-many translation based on this framework cannot perform on par with the individually trained models.In this work, we introduce three strategies to improve one-to-many multilingual translation by balancing the shared and unique features.Within the architecture of one decoder for all target languages, we first exploit the use of unique initial states for different target languages.Then, we employ language-dependent positional embeddings.Finally and especially, we propose to divide the hidden cells of the decoder into shared and language-dependent ones.The extensive experiments demonstrate that our proposed methods can obtain remarkable improvements over the strong baselines.Moreover, our strategies can achieve comparable or even better performance than the individually trained translation models. Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong |
EMNLP | 2 |
| 2018 | Addressing Troublesome Words in Neural Machine TranslationabstractOne of the weaknesses of Neural Machine Translation (NMT) is in handling lowfrequency and ambiguous words, which we refer as troublesome words.To address this problem, we propose a novel memoryenhanced NMT method.First, we investigate different strategies to define and detect the troublesome words.Then, a contextual memory is constructed to memorize which target words should be produced in what situations.Finally, we design a hybrid model to dynamically access the contextual memory so as to correctly translate the troublesome words.The extensive experiments on Chineseto-English and English-to-German translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling troublesome words. Yang Zhao 0007, Jiajun Zhang 0001, Zhongjun He, Chengqing Zong, Hua Wu 0003 |
EMNLP | 2 |
| 2018 | MSMO: Multimodal Summarization with Multimodal OutputabstractMultimodal summarization has drawn much attention due to the rapid growth of multimedia data.The output of the current multimodal summarization systems is usually represented in texts.However, we have found through experiments that multimodal output can significantly improve user satisfaction for informativeness of summaries.In this paper, we propose a novel task, multimodal summarization with multimodal output (MSMO).To handle this task, we first collect a large-scale dataset for MSMO research.We then propose a multimodal attention model to jointly generate text and select the most relevant image from the multimodal input.Finally, to evaluate multimodal outputs, we construct a novel multimodal automatic evaluation (MMAE) method which considers both intramodality salience and intermodality relevance.The experimental results show the effectiveness of MMAE. Junnan Zhu, Haoran Li 0001, Tianshang Liu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 5 |
| 2018 | Multi-modal Sentence Summarization with Modality Attention and Image FilteringabstractIn this paper, we introduce a multi-modal sentence summarization task that produces a short summary from a pair of sentence and image. This task is more challenging than sentence summarization. It not only needs to effectively incorporate visual features into standard text summarization framework, but also requires to avoid noise of image. To this end, we propose a modality-based attention mechanism to pay different attention to image patches and text units, and we design image filters to selectively use visual information to enhance the semantics of the input sentence. We construct a multimodal sentence summarization dataset and extensive experiments on this dataset demonstrate that our models significantly outperform conventional models which only employ text as input. Further analyses suggest that sentence summarization task can benefit from visually grounded representations from a variety of aspects. Haoran Li 0001, Junnan Zhu, Tianshang Liu, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 4 |
| 2018 | Phrase Table as Recommendation Memory for Neural Machine TranslationabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance recently. However, several studies indicate that NMT often generates fluent but unfaithful translations. In this paper, we propose a method to alleviate this problem by using a phrase table as recommendation memory. The main idea is to add bonus to words worthy of recommendation, so that NMT can make correct predictions. Specifically, we first derive a prefix tree to accommodate all the candidate target phrases by searching the phrase translation table according to the source sentence.Then, we construct a recommendation word set by matching between candidate target phrases and previously translated target words by NMT. After that, we determine the specific bonus value for each recommendable word by using the attention vector and phrase translation probability. Finally,we integrate this bonus value into NMT to improve the translation results. The extensive experiments demonstrate that the proposed methods obtain remarkable improvements over the strong attention based NMT. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 3 |
| 2018 | One Sentence One Model for Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
LREC | 2 |
| 2018 | Exploiting Pre-Ordering for Neural Machine Translation
Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
LREC | 2 |
| 2018 | A Comparable Study on Model Averaging, Ensembling and Reranking in NMT
Yuchen Liu 0007, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
NLPCC (2) | 5 |
| 2018 | Empirical Exploring Word-Character Relationship for Chinese Sentence RepresentationabstractThis article addresses the problem of learning compositional Chinese sentence representations, which represent the meaning of a sentence by composing the meanings of its constituent words. In contrast to English, a Chinese word is composed of characters, which contain rich semantic information. However, this information has not been fully exploited by existing methods. In this work, we introduce a novel, mixed character-word architecture to improve the Chinese sentence representations by utilizing rich semantic information of inner-word characters. We propose two novel strategies to reach this purpose. The first one is to use a mask gate on characters, learning the relation among characters in a word. The second one is to use a max-pooling operation on words to adaptively find the optimal mixture of the atomic and compositional word representations. Finally, the proposed architecture is applied to various sentence composition models, which achieves substantial performance gains over baseline models on sentence similarity task. To further verify the generalization ability of our model, we employ the learned sentence representations as features in sentence classification task, question classification task, and sentence entailment task. Results have shown that the proposed mixed character-word sentence representation models outperform both the character-based and word-based models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2017 | A Dynamic Window Neural Network for CCG SupertaggingabstractCombinatory Category Grammar (CCG) supertagging is a task to assign lexical categories to each word in a sentence. Almost all previous methods use fixed context window sizes to encode input tokens. However, it is obvious that different tags usually rely on different context window sizes. This motivates us to build a supertagger with a dynamic window approach, which can be treated as an attention mechanism on the local contexts. We find that applying dropout on the dynamic filters is superior to the regular dropout on word embeddings. We use this approach to demonstrate the state-of-the-art CCG supertagging performance on the standard test set. Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 2 |
| 2017 | Multi-modal Summarization for Asynchronous Collection of Text, Image, Audio and VideoabstractThe rapid increase in multimedia data transmission over the Internet necessitates the multi-modal summarization (MMS) from collections of text, image, audio and video.In this work, we propose an extractive multi-modal summarization method that can automatically generate a textual summary given a set of documents, images, audios and videos related to a specific topic.The key idea is to bridge the semantic gaps between multi-modal content.For audio information, we design an approach to selectively use its transcription.For visual information, we learn the joint representations of text and images using a neural network.Finally, all of the multimodal aspects are considered to generate the textual summary by maximizing the salience, non-redundancy, readability and coverage through the budgeted optimization of submodular functions.We further introduce an MMS corpus in English and Chinese, which is released to the public 1 .The experimental results obtained on this dataset demonstrate that our method outperforms other competitive baseline methods. Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 4 |
| 2017 | Exploiting Word Internal Structures for Generic Chinese Sentence RepresentationabstractWe introduce a novel mixed characterword architecture to improve Chinese sentence representations, by utilizing rich semantic information of word internal structures.Our architecture uses two key strategies.The first is a mask gate on characters, learning the relation among characters in a word.The second is a maxpooling operation on words, adaptively finding the optimal mixture of the atomic and compositional word representations.Finally, the proposed architecture is applied to various sentence composition models, which achieves substantial performance gains over baseline models on sentence similarity task. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 2 |
| 2017 | Learning Sentence Representation with Guidance of Human AttentionabstractRecently, much progress has been made in learning general-purpose sentence representations that can be used across domains. However, most of the existing models typically treat each word in a sentence equally. In contrast, extensive studies have proven that human read sentences efficiently by making a sequence of fixation and saccades. This motivates us to improve sentence representations by assigning different weights to the vectors of the component words, which can be treated as an attention mechanism on single sentences. To that end, we propose two novel attention models, in which the attention weights are derived using significant predictors of human reading time, i.e., Surprisal, POS tags and CCG supertags. The extensive experiments demonstrate that the proposed methods significantly improve upon the state-of-the-art sentence representation models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 2 |
| 2017 | Towards Neural Machine Translation with Partially Aligned CorporaabstractWhile neural machine translation (NMT) has become the new paradigm, the parameter optimization requires large-scale parallel data which is scarce in many domains and language pairs. In this paper, we address a new translation scenario in which there only exists monolingual corpora and phrase pairs. We propose a new method towards translation with partially aligned sentence pairs which are derived from the phrase pairs and monolingual corpora. To make full use of the partially aligned corpora, we adapt the conventional NMT training method in two aspects. On one hand, different generation strategies are designed for aligned and unaligned target words. On the other hand, a different objective function is designed to model the partially aligned parts. The experiments demonstrate that our method can achieve a relatively good result in such a translation scenario, and tiny bitexts can boost translation quality to a large extent. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhengshan Xue |
IJCNLP(1) | 3 |
| 2017 | Shortcut Sequence Tagging
Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
NLPCC | 2 |
| 2017 | Look-Ahead Attention for Generation in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
NLPCC | 2 |
| 2017 | Augmenting Neural Sentence Summarization Through Extractive Summarization
Junnan Zhu, Haoran Li 0001, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
NLPCC | 4 |
| 2017 | Implicit Discourse Relation Recognition for English and Chinese with Multiview Modeling and Effective Representation LearningabstractDiscourse relations between two text segments play an important role in many Natural Language Processing (NLP) tasks. The connectives strongly indicate the sense of discourse relations, while in fact, there are no connectives in a large proportion of discourse relations, that is, implicit discourse relations. Compared with explicit relations, implicit relations are much harder to detect and have drawn significant attention. Until now, there have been many studies focusing on English implicit discourse relations, and few studies address implicit relation recognition in Chinese even though the implicit discourse relations in Chinese are more common than those in English. In our work, both the English and Chinese languages are our focus. The key to implicit relation prediction is to properly model the semantics of the two discourse arguments, as well as the contextual interaction between them. To achieve this goal, we propose a neural network based framework that consists of two hierarchies. The first one is the model hierarchy, in which we propose a max-margin learning method to explore the implicit discourse relation from multiple views. The second one is the feature hierarchy, in which we learn multilevel distributed representations from words, arguments, and syntactic structures to sentences. We have conducted experiments on the standard benchmarks of English and Chinese, and the results show that compared with several methods our proposed method can achieve the best performance in most cases. Haoran Li 0001, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | An Empirical Exploration of Skip Connections for Sequential TaggingabstractIn this paper, we empirically explore the effects of various kinds of skip connections in stacked bidirectional LSTMs for sequential tagging. We investigate three kinds of skip connections connecting to LSTM cells: (a) skip connections to the gates, (b) skip connections to the internal states and (c) skip connections to the cell outputs. We present comprehensive experiments showing that skip connections to cell outputs outperform the remaining two. Furthermore, we observe that using gated identity functions as skip mappings works pretty well. Based on this novel skip connections, we successfully train deep stacked bidirectional LSTM models and obtain state-of-the-art results on CCG supertagging and comparable results on POS tagging. Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
COLING | 2 |
| 2016 | Exploiting Source-side Monolingual Data in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 1 |
| 2016 | Towards Zero Unknown Word in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 2 |
| 2016 | A Bilingual Discourse Corpus and Its Applications
Yang Liu 0085, Jiajun Zhang 0001, Chengqing Zong, Yating Yang, Xi Zhou 0007 |
LREC | 2 |
| 2016 | Abstractive Cross-Language Summarization via Translation Model Enhanced Predicate Argument Structure FusingabstractCross-language multidocument summarization is the task to generate a summary in a target language (e.g., Chinese) from a collection of documents in a different source language (e.g., English). Previous methods such as the extractive and compressive algorithms focus only on single sentence selection and compression, which cannot make full use of the similar sentences containing complementary information. Furthermore, the translation model knowledge is not fully explored in previous approaches. To address these two problems, we propose in this paper an abstractive cross-language summarization framework. First, the source language documents are translated into target language with a machine translation system. Then, the method constructs a pool of bilingual concepts and facts represented by the bilingual elements of the source-side predicate-argument structures (PAS) and their target-side counterparts. Finally, new summary sentences are produced by fusing bilingual PAS elements with the integer linear programming algorithm to maximize both of the salience and translation quality of the PAS elements. The experimental results on English-to-Chinese cross-language summarization demonstrate that our proposed method outperforms the state-of-the-art extractive systems in both automatic and manual evaluations. Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | A New Input Method for Human Translators: Integrating Machine Translation Effectively and Imperceptibly
Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IJCAI | 2 |
| 2015 | Towards Machine Translation in Semantic Vector SpaceabstractMeasuring the quality of the translation rules and their composition is an essential issue in the conventional statistical machine translation (SMT) framework. To express the translation quality, the previous lexical and phrasal probabilities are calculated only according to the co-occurrence statistics in the bilingual corpus and may be not reliable due to the data sparseness problem. To address this issue, we propose measuring the quality of the translation rules and their composition in the semantic vector embedding space (VES). We present a recursive neural network (RNN)-based translation framework, which includes two submodels. One is the bilingually-constrained recursive auto-encoder, which is proposed to convert the lexical translation rules into compact real-valued vectors in the semantic VES. The other is a type-dependent recursive neural network, which is proposed to perform the decoding process by minimizing the semantic gap (meaning distance) between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function that maximizes the margin between the reference translation and the n-best translations in forced decoding. In the experiments, we first show that the proposed vector representations for the translation rules are very reliable for application in translation modeling. We further show that the proposed type-dependent, RNN-based model can significantly improve the translation quality in the large-scale, end-to-end Chinese-to-English translation evaluation. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2014 | Mind the Gap: Machine Translation by Minimizing the Semantic Gap in Embedding SpaceabstractThe conventional statistical machine translation (SMT) methods perform the decoding process by compositing a set of the translation rules which are associated with high probabilities. However, the probabilities of the translation rules are calculated only according to the cooccurrence statistics in the bilingual corpus rather than the semantic meaning similarity. In this paper, we propose a Recursive Neural Network (RNN) based model that converts each translation rule into a compact real-valued vector in the semantic embedding space and performs the decoding process by minimizing the semantic gap between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function. Extensive experiments on Chinese-to-English translation show that our RNN-based model can significantly improve the translation quality by up to 1.68 BLEU score. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
AAAI | 1 |
| 2014 | Bilingually-constrained Phrase Embeddings for Machine TranslationabstractWe propose Bilingually-constrained Recursive Auto-encoders (BRAE) to learn semantic phrase embeddings (compact vector representations for phrases), which can distinguish the phrases with different semantic meanings.The BRAE is trained in a way that minimizes the semantic distance of translation equivalents and maximizes the semantic distance of nontranslation pairs simultaneously.After training, the model learns how to embed each phrase semantically in two languages and also learns how to transform semantic embedding space in one language to the other.We evaluate our proposed method on two end-to-end SMT tasks (phrase table pruning and decoding with phrasal semantic similarities) which need to measure semantic similarity between a source phrase and its translation candidates.Extensive experiments show that the BRAE is remarkably effective in these two tasks. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACL (1) | 1 |
| 2013 | Handling Ambiguities of Bilingual Predicate-Argument Structures for Statistical Machine Translation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 2 |
| 2013 | Learning a Phrase-based Translation Model from Monolingual Data with Application to Domain Adaptation
Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 1 |
| 2013 | A Substitution-Translation-Restoration Framework for Handling Unknown Words in Statistical Machine Translation
Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
J. Comput. Sci. Technol. | 1 |
| 2013 | Unsupervised Tree Induction for Tree-based TranslationabstractIn current research, most tree-based translation models are built directly from parse trees. In this study, we go in another direction and build a translation model with an unsupervised tree structure derived from a novel non-parametric Bayesian model. In the model, we utilize synchronous tree substitution grammars (STSG) to capture the bilingual mapping between language pairs. To train the model efficiently, we develop a Gibbs sampler with three novel Gibbs operators. The sampler is capable of exploring the infinite space of tree structures by performing local changes on the tree nodes. Experimental results show that the string-to-tree translation system using our Bayesian tree structures significantly outperforms the strong baseline string-to-tree system using parse trees. Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Syntax-Based Translation With Bilingually Lexicalized Synchronous Tree Substitution GrammarsabstractSyntax-based models can significantly improve the translation performance due to their grammatical modeling on one or both language side(s). However, the translation rules such as the non-lexical rule “ VP→(x0x1,VP:x1PP:x0)” in string-to-tree models do not consider any lexicalized information on the source or target side. The rule is so generalized that any subtree rooted at VP can substitute for the nonterminal VP:x1. Because rules containing nonterminals are frequently used when generating the target-side tree structures, there is a risk that rules of this type will potentially be severely misused in decoding due to a lack of lexicalization guidance. In this article, inspired by lexicalized PCFG, which is widely used in monolingual parsing, we propose to upgrade the STSG (synchronous tree substitution grammars)-based syntax translation model with bilingually lexicalized STSG. Using the string-to-tree translation model as a case study, we present generative and discriminative models to integrate lexicalized STSG into the translation model. Both small- and large-scale experiments on Chinese-to-English translation demonstrate that the proposed lexicalized STSG can provide superior rule selection in decoding and substantially improve the translation quality. Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Machine Translation by Modeling Predicate-Argument Structure Transformation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 2 |
| 2012 | Tree-based Translation without using Parse Trees
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 2 |
| 2012 | Towards a Chinese Common and Common Sense Knowledge Base for Sentiment Analysis
Erik Cambria, Amir Hussain 0001, Tariq S. Durrani, Jiajun Zhang 0001 |
IEA/AIE | 4 |
| 2011 | Augmenting String-to-Tree Translation Models with Fuzzy Use of Source-side Syntax
Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
EMNLP | 1 |
| 2011 | Simple but Effective Approaches to Improving Tree-to-tree Model
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
MTSummit | 2 |
| 2009 | A Framework for Effectively Integrating Hard and Soft Syntactic Rules into Phrase Based Translation
Jiajun Zhang 0001, Chengqing Zong |
PACLIC | 1 |
| 2008 | Sentence Type Based Reordering Model for Statistical Machine Translation
Jiajun Zhang 0001, Chengqing Zong, Shoushan Li |
COLING | 1 |