VLDB 2026 Research / reviewers in the wild / expert
Chengqing Zong
dblp:38/6093
· DBLP profile ↗
214ranked-venue papers
7as first author
66since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 196 · 4 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DART: Disambiguation-Aware Reasoning for Video-guided Machine TranslationabstractVideo-guided Machine Translation (VMT) seeks to enhance translation quality by incorporating contextual information derived from paired short video clips.However, many VMT samples are text-sufficient; even when visual information is needed, only minimal cues are required.Aiming to tackle these issues, we propose a novel framework DART (Disambiguation-Aware Reasoning for Videoguided Machine Translation).Reinforcement learning is used to incorporate multimodal large language models' multimodal reasoning into VMT.The model dynamically switches between text-only processing and multimodal integration, contingent on the necessity of visual disambiguation.Furthermore, we present TVRF (Translation-oriented Video Relevance Filtering), a systematic pipeline for constructing training data based on multimodal relevance to translation.This pipeline filters samples where video information is translationrelevant, mitigating training collapse caused by video-irrelevant data in conventional VMT.Experimental results show that our approach improves multimodal information utilization in VMT, yielding gains in both translation quality and computational efficiency.* Equal corresponding authors.Reasoning: Okay, I need to translate the input sentence ... So the answer is "哦!".(809 tokens in total) DA T Existing LMRMs … "It's hitting his yacht."Translation: 哦!(Oh!) Reasoning: Okay, so I need to translate the sentence .. Boyu Guan, Chuang Han, Yang Zhao 0007, Chengqing Zong |
ACL (1) | 4 |
| 2026 | EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldabstractCurrent evaluation paradigms for emotional support conversations tend to reward generic empathetic responses, yet they fail to assess whether the support is genuinely personalized to users' unique psychological profiles and contextual needs.We introduce EmoHarbor, an automated evaluation framework that adopts a User-as-a-Judge paradigm by simulating the user's inner world.EmoHarbor employs a Chain-of-Agent architecture that decomposes users' internal processes into three specialized roles, enabling agents to interact with supporters and complete assessments in a manner similar to human users.We instantiate this benchmark using 100 real-world user profiles that cover a diverse range of personality traits and situations, and define 10 evaluation dimensions of personalized support quality.Comprehensive evaluation of 20 advanced LLMs on EmoHarbor reveals a critical insight: while these models excel at generating empathetic responses, they consistently fail to tailor support to individual user contexts.This finding reframes the central challenge, shifting research focus from merely enhancing generic empathy to developing truly user-aware emotional support.EmoHarbor provides a reproducible and scalable framework to guide the development and evaluation of more nuanced and user-aware emotional support systems 1 . Lu Xiang, Chengqing Zong |
ACL (1) | 4 |
| 2026 | ASMem: Anchor sparse memory for multi-domain knowledge editing of large language models
Guanyu Zheng, Xv Wang, Haochang Wang, Tiejun Zhao, Chengqing Zong |
Neural Networks | 8 |
| 2026 | Reason-Align-Respond: Aligning LLM Reasoning With Knowledge Graphs for KGQAabstractLarge language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, yet they often suffer from hallucinations and lack reliable factual grounding. Meanwhile, knowledge graphs (KGs) provide structured factual knowledge, but lack the flexible reasoning abilities of LLMs. In this paper, we present Reason-Align-Respond (RAR), a novel framework that systematically integrates LLM reasoning with knowledge graphs for knowledge graph question answering (KGQA). Our approach consists of three key components: a Reasoner that generates human-like natural language reasoning chains, an Aligner that maps these chains to valid KG paths, and a Responser that synthesizes the final answer. We formulate this process as a latent variable mixture model and optimize it using the Expectation-Maximization algorithm, which iteratively refines the reasoning chains and knowledge paths. Extensive experiments on multiple benchmarks demonstrate the effectiveness of RAR, achieving state-of-the-art performance with Hit scores of 93.3% and 91.0% on WebQSP and CWQ respectively. Human evaluation confirms that RAR generates high-quality, interpretable reasoning chains well-aligned with KG paths while maintaining computational efficiency during inference. Xiangqing Shen, Fanfan Wang, Zinong Yang, Wenli Du, Chengqing Zong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | TokAlign: Efficient Vocabulary Adaptation via Token AlignmentabstractTokenization serves as a foundational step for Large Language Models (LLMs) to process text.In new domains or languages, the inefficiency of the tokenizer will slow down the training and generation of LLM.The mismatch in vocabulary also hinders deep knowledge transfer between LLMs like token-level distillation.To mitigate this gap, we propose an efficient method named TokAlign to replace the vocabulary of LLM from the token co-occurrences view, and further transfer the token-level knowledge between models.It first aligns the source vocabulary to the target one by learning a oneto-one mapping matrix for token IDs.Model parameters, including embeddings, are rearranged and progressively fine-tuned for the new vocabulary.Our method significantly improves multilingual text compression rates and vocabulary initialization for LLMs, decreasing the perplexity from 3.4e 2 of strong baseline methods to 1.2e 2 after initialization.Experimental results on models across multiple parameter scales demonstrate the effectiveness and generalization of TokAlign, which costs as few as 5k steps to restore the performance of the vanilla model.After unifying vocabularies between LLMs, token-level distillation can remarkably boost (+4.4% than sentence-level distillation) the base model, costing only 235M tokens. 1 Chengqing Zong |
ACL (1) | 3 |
| 2025 | Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine TranslationabstractYupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yupu Liang, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
ACL (1) | 6 |
| 2025 | TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship ClassificationabstractLLMs have achieved remarkable fluency and coherence in text generation, yet their widespread adoption has raised concerns about content reliability and accountability. In high-stakes domains, it is crucial to understand where and how the content is created. To address this, we introduce the Text pROVEnance (TROVE) challenge, designed to trace each sentence of a target text back to specific source sentences within potentially lengthy or multi-document inputs. Beyond identifying sources, TROVE annotates the fine-grained relationships (quotation, compression, inference, and others), providing a deep understanding of how each target sentence is formed.To benchmark TROVE, we construct our dataset by leveraging three public datasets covering 11 diverse scenarios (e.g., QA and summarization) in English and Chinese, spanning source texts of varying lengths (0–5k, 5–10k, 10k+), emphasizing the multi-document and long-document settings essential for provenance. To ensure high-quality data, we employ a three-stage annotation process: sentence retrieval, GPT-4o provenance, and human provenance. We evaluate 11 LLMs under direct prompting and retrieval-augmented paradigms, revealing that retrieval is essential for robust performance, larger models perform better in complex relationship classification, and closed-source models often lead, yet open-source models show significant promise, particularly with retrieval augmentation. We make our dataset available here: https://github.com/ZNLP/ZNLP-Dataset. Junnan Zhu, Feifei Zhai, Chengqing Zong |
ACL (1) | 6 |
| 2025 | TriFine: A Large-Scale Dataset of Vision-Audio-Subtitle for Tri-Modal Machine Translation and Benchmark with Fine-Grained Annotated TagsabstractCurrent video-guided machine translation (VMT) approaches primarily use coarse-grained visual information, resulting in information redundancy, high computational overhead, and neglect of audio content. Our research demonstrates the significance of fine-grained visual and audio information in VMT from both data and methodological perspectives. From the data perspective, we have developed a large-scale dataset TriFine, the first vision-audio-subtitle tri-modal VMT dataset with annotated multimodal fine-grained tags. Each entry in this dataset not only includes the triples found in traditional VMT datasets but also encompasses seven fine-grained annotation tags derived from visual and audio modalities. From the methodological perspective, we propose a Fine-grained Information-enhanced Approach for Translation (FIAT). Experimental results have shown that, in comparison to traditional coarse-grained methods and text-only models, our fine-grained approach achieves superior performance with lower computational overhead. These findings underscore the pivotal role of fine-grained annotated information in advancing the field of VMT. Boyu Guan, Yang Zhao 0007, Chengqing Zong |
COLING | 4 |
| 2025 | SweetieChat: A Strategy-Enhanced Role-playing Framework for Diverse Scenarios Handling Emotional Support AgentabstractLarge Language Models (LLMs) have demonstrated promising potential in providing empathetic support during interactions. However, their responses often become verbose or overly formulaic, failing to adequately address the diverse emotional support needs of real-world scenarios. To tackle this challenge, we propose an innovative strategy-enhanced role-playing framework, designed to simulate authentic emotional support conversations. Specifically, our approach unfolds in two steps: (1) Strategy-Enhanced Role-Playing Interactions, which involve three pivotal roles—Seeker, Strategy Counselor, and Supporter—engaging in diverse scenarios to emulate real-world interactions and promote a broader range of dialogues; and (2) Emotional Support Agent Training, achieved through fine-tuning LLMs using our specially constructed dataset. Within this framework, we develop the ServeForEmo dataset, comprising an extensive collection of 3.7K+ multi-turn dialogues and 62.8K+ utterances. We further present SweetieChat, an emotional support agent capable of handling diverse open-domain scenarios. Extensive experiments and human evaluations confirm the framework’s effectiveness in enhancing emotional support, highlighting its unique ability to provide more nuanced and tailored assistance. Lu Xiang, Chengqing Zong |
COLING | 4 |
| 2025 | From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image TranslationabstractDocument Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in complex-layout scenarios, frequently loses the crucial document coherence, leading to chaotic text. To overcome this problem, we introduce a novel end-to-end network, named Zoom-out DIT (ZoomDIT), inspired by human translation procedures. It jointly accomplishes the multi-level tasks including word positioning, sentence recognition & translation, and document organization, based on a fine-to-coarse zoom-out framework, to progressively realize “chaotic words to coherent document” and improve translation. We further contribute a new large-scale DIT dataset with multi-level fine-grained labels. Extensive experiments on public and our new dataset demonstrate significant improvements in translation quality towards complex-layout document images, offering a robust solution for reorganizing the chaotic OCR outputs to a coherent document translation. Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
COLING | 7 |
| 2025 | SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationabstractVideo-guided Machine Translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips.Mainstream VMT approaches typically incorporate multimodal information by uniformly sampling frames from the input videos.However, this paradigm frequently incurs significant computational overhead and introduces redundant multimodal content, which degrades both efficiency and translation quality.To tackle these challenges, we propose SHIFT (Selected Helpful Informative Frame for Translation).It is a lightweight, plug-andplay framework designed for VMT with Multimodal Large Language Models (MLLMs).SHIFT adaptively selects a single informative key frame when visual context is necessary; otherwise, it relies solely on textual input.This process is guided by a dedicated clustering module and a selector module.Experimental results demonstrate that SHIFT enhances the performance of MLLMs on the VMT task while simultaneously reducing computational cost, without sacrificing generalization ability. Boyu Guan, Chuang Han, Yupu Liang, Yang Zhao 0007, Chengqing Zong |
EMNLP | 7 |
| 2025 | ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (5) | 8 |
| 2025 | Language Imbalance Driven Rewarding for Multilingual Self-improvingabstractLarge Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such as English and Chinese, leaving many other languages underrepresented. This imbalance, while limiting broader applications, generates a natural preference ranking between languages, offering an opportunity to bootstrap the multilingual capabilities of LLM in a self-improving manner. Thus, we propose $\textit{Language Imbalance Driven Rewarding}$, where the inherent imbalance between dominant and non-dominant languages within LLMs is leveraged as a reward signal. Iterative DPO training demonstrates that this approach not only enhances LLM performance in non-dominant languages but also improves the dominant language's capacity, thereby yielding an iterative reward signal. Fine-tuning Meta-Llama-3-8B-Instruct over two iterations of this approach results in continuous improvements in multilingual performance across instruction-following and arithmetic reasoning tasks, evidenced by an average improvement of 7.46\% win rate on the X-AlpacaEval leaderboard and 13.9\% accuracy on the MGSM benchmark. This work serves as an initial exploration, paving the way for multilingual self-improvement of LLMs. Junhong Wu, Chengqing Zong, Jiajun Zhang 0001 |
ICLR | 4 |
| 2025 | SimulPL: Aligning Human Preferences in Simultaneous Machine TranslationabstractSimultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios have distinct preferences, such as accurate translations, simpler syntax, and no unnecessary latency. Aligning SiMT models with these human preferences is crucial to improve their performances. However, this issue still remains unexplored. Additionally, preference optimization for SiMT task is also challenging. Existing methods focus solely on optimizing the generated responses, ignoring human preferences related to latency and the optimization of read/write policy during the preference optimization phase. To address these challenges, we propose Simultaneous Preference Learning (SimulPL), a preference learning framework tailored for the SiMT task. In the SimulPL framework, we categorize SiMT human preferences into five aspects: **translation quality preference**, **monotonicity preference**, **key point preference**, **simplicity preference**, and **latency preference**. By leveraging the first four preferences, we construct human preference prompts to efficiently guide GPT-4/4o in generating preference data for the SiMT task. In the preference optimization phase, SimulPL integrates **latency preference** into the optimization objective and enables SiMT models to improve the read/write policy, thereby aligning with human preferences more effectively. Experimental results indicate that SimulPL exhibits better alignment with human preferences across all latency levels in Zh$\rightarrow$En, De$\rightarrow$En and En$\rightarrow$Zh SiMT tasks. Our data and code will be available at https://github.com/EurekaForNLP/SimulPL. Donglei Yu, Yang Zhao 0007, Yangyifan Xu, Yu Zhou 0001, Chengqing Zong |
ICLR | 6 |
| 2025 | Pay More Attention to Images: Numerous Images-Oriented Multimodal SummarizationabstractMin Xiao, Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Investigating Hallucinations in Simultaneous Machine Translation: Knowledge Distillation Solution and Components AnalysisabstractDonglei Yu, Xiaomian Kang, Yuchen Liu, Feifei Zhai, Nanchang Cheng, Yu Zhou, Chengqing Zong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Donglei Yu, Xiaomian Kang, Yuchen Liu 0007, Feifei Zhai, Nanchang Cheng, Yu Zhou 0001, Chengqing Zong |
NAACL (Long Papers) | 7 |
| 2025 | Boosting Document Image Translation via Layout-Aware Semantic Paragraph Clustering
Yupu Liang, Yunfei Lu, Dandan Tu, Chengqing Zong, Yu Zhou 0001 |
PRCV (7) | 8 |
| 2025 | EmoDial-Reason: Unveiling Affective Reasoning in Speech-Emotion Dialogue
Shubei Tang, Lu Xiang, Chengqing Zong |
PRCV (12) | 4 |
| 2025 | Learning Emotion Category Representation to Detect Emotion Relations Across LanguagesabstractUnderstanding human emotions is crucial for a myriad of applications, from psychological research to advancements in Natural Language Processing (NLP). Traditionally, emotions are categorized into distinct basic groups, which has led to the development of various emotion detection tasks within NLP. However, these tasks typically rely on one-hot vectors to represent emotions, a method that fails to capture the relations between different emotion categories. In this study, we challenge the assumption that emotion categories are mutually exclusive and argue that the connections and boundaries between them are complex and often blurred. To better represent these nuanced interconnections, we introduce an innovative framework as well as two algorithms to learn distributed representations of emotion categories by leveraging soft labels from trained neural network models. For the first time, our approach enables the detection of emotion relations across different languages through an NLP lens, a feat unattainable with traditional one-hot representations. Validation experiments confirm the superior ability of our distributed representation algorithms to articulate these emotional connections. Moreover, application experiments corroborate several interdisciplinary insights into cross-linguistic emotion relations, findings that align with research in psychology and linguistics. This work not only presents a breakthrough in emotion detection but also bridges the gap between computational models and humanistic understanding of emotions. Chengqing Zong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image TranslationabstractDocument Image Translation (DIT) aims to translate texts on document images from one language to another. It is a multi-modal task involving cooperation of text and layout. Current approaches either handle layout and translation as separate processes, risking accumulative errors, or use vanilla end-to-end encoder-decoder models to capture layout implicitly, often suffering inadequate layout incorporation. We argue that a favorable framework should explicitly engage layout-specific modules and properly organize them toward translation. For this, we first revisit two key layouts: the geometric layout reflecting word's spatial positions, and the logical layout depicting word's logical order. Then, a novel pipeline (understand layout $\rightarrow$→ translate text) is determined to prioritize layouts such that preceding layouts contribute to translation. Following this pipeline, we introduce Unified Document Image Translation (UniDIT), a comprehensive framework that unifies layout with translation in one network. It is devised to leverage each module's advantage, and provide an elaborate feature-conductive flow for module communication globally. A novel bridging mechanism is also introduced to adapt layout features conducive to translation. We further contribute DITransv2, a large-scale fine-grained benchmark that includes heterogeneous and complex document layouts. Extensive experiments on DITransv2 and additional established benchmarks demonstrate UniDIT outperforms previous state-of-the-arts in all aspects. Yupu Liang, Cong Ma 0002, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | DIUSum: Dynamic Image Utilization for Multimodal SummarizationabstractExisting multimodal summarization approaches focus on fusing image features in the encoding process, ignoring the individualized needs for images when generating different summaries. However, whether intuitively or empirically, not all images can improve summary quality. Therefore, we propose a novel Dynamic Image Utilization framework for multimodal Summarization (DIUSum) to select and utilize valuable images for summarization. First, to predict whether an image helps produce a high-quality summary, we propose an image selector to score the usefulness of each image. Second, to dynamically utilize the multimodal information, we incorporate the hard and soft guidance from the image selector. Under the guidance, the image information is plugged into the decoder to generate a summary. Experimental results have shown that DIUSum outperforms multiple strong baselines and achieves SOTA on two public multimodal summarization datasets. Further analysis demonstrates that the image selector can reflect the improved level of summary quality brought by the images. Junnan Zhu, Feifei Zhai, Yu Zhou 0001, Chengqing Zong |
AAAI | 5 |
| 2024 | Self-Modifying State Modeling for Simultaneous Machine TranslationabstractSimultaneous Machine Translation (SiMT) generates target outputs while receiving stream source inputs and requires a read/write policy to decide whether to wait for the next source token or generate a new target token, whose decisions form a decision path.Existing SiMT methods, which learn the policy by exploring various decision paths in training, face inherent limitations.These methods not only fail to precisely optimize the policy due to the inability to accurately assess the individual impact of each decision on SiMT performance, but also cannot sufficiently explore all potential paths because of their vast number.Besides, building decision paths requires unidirectional encoders to simulate streaming source inputs, which impairs the translation quality of SiMT models.To solve these issues, we propose Self-Modifying State Modeling (SM 2 ), a novel training paradigm for SiMT task.Without building decision paths, SM 2 individually optimizes decisions at each state during training.To precisely optimize the policy, SM 2 introduces Self-Modifying process to independently assess and adjust decisions at each state.For sufficient exploration, SM 2 proposes Prefix Sampling to efficiently traverse all potential states.Moreover, SM 2 ensures compatibility with bidirectional encoders, thus achieving higher translation quality.Experiments show that SM 2 outperforms strong baselines.Furthermore, SM 2 allows offline machine translation models to acquire SiMT ability with fine-tuning 1 . Donglei Yu, Xiaomian Kang, Yuchen Liu 0007, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 5 |
| 2024 | Navigating Brain Language Representations: A Comparative Analysis of Neural Language Models and Psychologically Plausible Models
Shaonan Wang, Xinyi Dong, Jiajun Yu, Chengqing Zong |
CogSci | 5 |
| 2024 | Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine TranslationabstractText image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained knowledge supervision to make it consistent between recognition and translation modules. In this paper, we propose a novel TIMT method named as BabyNet, which is optimized with hierarchical parental supervision to improve translation performance. Inspired by genetic recombination and variation in the field of genetics, the proposed BabyNet is inherited from the recognition and translation parent models with a variation module of which parameters can be updated when training on the TIMT task. Meanwhile, hierarchical and multi-granularity supervision from parent models is introduced to bridge the gap between inherited modules in BabyNet. Extensive experiments on both synthetic and real-world TIMT tests show that our proposed method significantly outperforms existing methods. Further analyses of various parent model combinations show the good generalization of our method. Cong Ma 0002, Yupu Liang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
LREC/COLING | 7 |
| 2024 | A Hybrid Approach towards Chinese Spelling and Splitting Error CorrectionabstractExisting Chinese spelling check (CSC) methods have limitations in correcting variable-length error characters, requiring the input and output to be the same length. They mainly focus on modelling Chinese characters’ phonetic information and generating candidates for each position. In contrast, few approaches delve into the intricacies of splitting Chinese characters to address glyph errors and splitting variable-length corrections. We define the Chinese Splitting Error Correction (CSEC) task and develop CSEC datasets in news and social media domains to address this issue. We then propose Soft-Masked Multi-feature Error Correction (SoMu) model, which first generates semantic, phonetic, graphic, and unique Chinese Wubi embeddings, then integrates those features through selective gating fusion, followed by a soft-mask strategy to filter incorrect tokens and finally use transformer layers to predict the correct ones. This model effectively addresses both spelling and splitting errors. Extensive analysis shows that our model significantly improves character-splitting information modelling for CSEC. Our dataset is available at https://github.com/Skywalker-Harrison/SoMu. Junhong Liang, Junnan Zhu, Feifei Zhai, Nanchang Cheng, Chengqing Zong, Yu Zhou 0001 |
ECAI | 5 |
| 2024 | BLSP-Emo: Towards Empathetic Large Speech-Language ModelsabstractThe recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions.While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible.In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speechlanguage model capable of understanding both semantics and emotions in speech and generate empathetic responses.BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process.The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data.The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data.Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations. 1 Minpeng Liao, Zhongqiang Huang, Junhong Wu, Chengqing Zong, Jiajun Zhang 0001 |
EMNLP | 5 |
| 2024 | Vector Quantization Knowledge Transfer for End-to-End Text Image Machine TranslationabstractEnd-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning continuous features from related tasks of text image recognition (TIR) or machine translation (MT) to alleviate the problem of data limitation, but the alignment in continuous vector space is extremely difficult and it inevitably introduces fitting errors resulting in significant performance degradation. To better align TIMT features with MT semantic features, we propose a novel Vector Quantization Knowledge Transfer (VQKT) method that employs a trainable codebook to quantize continuous features into discrete space. The quantization distribution of the MT feature is utilized as the teacher distribution to guide the TIMT model to generate similar discrete codes. Through alignment and knowledge transfer based on probability distribution, the TIMT model can better imitate the feature representation of the MT teacher model and generate high-quality target language translation. Extensive experiments demonstrate VQKT significantly outperforms the existing end-to-end TIMT performance. Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICASSP | 5 |
| 2024 | Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual AlignmentabstractChong Li, Shaonan Wang, Jiajun Zhang, Chengqing Zong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shaonan Wang, Chengqing Zong |
NAACL-HLT | 4 |
| 2024 | Document Image Machine Translation with Dynamic Multi-pre-trained Models AssemblingabstractYupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yupu Liang, Cong Ma 0002, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
NAACL-HLT | 7 |
| 2024 | F-MALLOC: Feed-forward Memory Allocation for Continual Learning in Neural Machine TranslationabstractJunhong Wu, Yuchen Liu, Chengqing Zong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junhong Wu, Yuchen Liu 0007, Chengqing Zong |
NAACL-HLT | 3 |
| 2024 | MapGuide: A Simple yet Effective Method to Reconstruct Continuous Language from Brain ActivitiesabstractXinpei Zhao, Jingyuan Sun, Shaonan Wang, Jing Ye, Xiaohan Zhang, Chengqing Zong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xinpei Zhao, Shaonan Wang, Chengqing Zong |
NAACL-HLT | 6 |
| 2024 | 🚀 TableRocket: An Efficient and Effective Framework for Table Reconstruction
Liucheng Pang, Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
PRCV (7) | 6 |
| 2024 | Knowledge Graph Guided Neural Machine Translation with Dynamic Reinforce-selected TriplesabstractPrevious methods incorporating knowledge graphs (KGs) into neural machine translation (NMT) adopt a static knowledge utilization strategy, that introduces many useless knowledge triples and makes the useful triples difficult to be utilized by NMT. To address this problem, we propose a KG guided NMT model with dynamic reinforce-selected triples. The proposed methods could dynamically select the different useful knowledge triples for different source sentences. Specifically, the proposed model contains two components: (1) knowledge selector, that dynamically selects useful knowledge triples for a source sentence, and (2) knowledge guided NMT (KgNMT), that utilizes the selected triples to guide the translation of NMT. Meanwhile, to overcome the non-differentiable problem and guide the training procedure, we propose a policy gradient strategy to encourage the model to select useful triples and improve the generation probability of gold target sentence. Various experimental results show that the proposed method can significantly outperform the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Xiaomian Kang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2024 | Modal Contrastive Learning Based End-to-End Text Image Machine TranslationabstractText image machine translation (TIMT) aims at directly translating text in the source language embedded in images into the target language. Most existing systems follow the cascaded pipeline diagram from recognition to translation, which suffers from the problem of error propagation, parameter redundancy, and information reduction. The end-to-end model has the potential to alleviate these issues via bridging the recognition and translation models. However, the challenge is the data limitation and modality gap between text and image. In this paper, we propose a novel end-to-end model, namely Modal contrastive learning based End-to-end Text Image Machine Translation (METIMT), which alleviates these issues through end-to-end text image machine translation architecture and modal contrastive learning. Specifically, an image encoder is designed to encode images into the same feature space of corresponding text sentences, with the guidance of an intramodal and inter-modal contrastive learning module. To further promote the research of text image machine translation, we have constructed one synthetic and two real-world datasets. Extensive experiments show that our lighter, faster model outperforms not only existing pipeline methods but also state-of-the-art end-to-end models on both synthetic and real-world evaluation sets. Our code and dataset will be released to the public. Cong Ma 0002, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | Multilingual Knowledge Graph Completion with Language-Sensitive Multi-Graph AttentionabstractMultilingual Knowledge Graph Completion (KGC) aims to predict missing links with multilingual knowledge graphs.However, existing approaches suffer from two main drawbacks: (a) alignment dependency: the multilingual KGC is always realized with joint entity or relation alignment, which introduces additional alignment models and increases the complexity of the whole framework; (b) training inefficiency: the trained model will only be used for the completion of one target KG, although the data from all KGs are used simultaneously.To address these drawbacks, we propose a novel multilingual KGC framework with languagesensitive multi-graph attention such that the missing links on all given KGs can be inferred by a universal knowledge completion model.Specifically, we first build a relational graph neural network by sharing the embeddings of aligned nodes to transfer language-independent knowledge.Meanwhile, a language-sensitive multi-graph attention (LSMGA) is proposed to deal with the information inconsistency among different KGs.Experimental results show that our model achieves significant improvements on the DBP-5L and E-PKG datasets.1 Rongchuan Tang, Yang Zhao 0007, Chengqing Zong, Yu Zhou 0001 |
ACL (1) | 3 |
| 2023 | CFSum Coarse-to-Fine Contribution Network for Multimodal SummarizationabstractMultimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear.Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are useful.Therefore, we propose a novel Coarse-to-Fine contribution network for multimodal Summarization (CFSum) to consider different contributions of images for summarization.First, to eliminate the interference of useless images, we propose a pre-filter module to abandon useless images.Second, to make accurate use of useful images, we propose two levels of visual complement modules, word level and phrase level.Specifically, image contributions are calculated and are adopted to guide the attention of both textual and visual modalities.Experimental results have shown that CFSum significantly outperforms multiple strong baselines on the standard benchmark.Furthermore, the analysis verifies that useful images can even help generate nonvisual words which are implicitly represented in the image 1 . Junnan Zhu, Haitao Lin 0001, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 5 |
| 2023 | Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsabstractFine-tuning all parameters of large language models (LLMs) requires significant computational resources and is time-consuming.Recent parameter-efficient tuning methods such as Adapter tuning, Prefix tuning, and LoRA allow updating a small subset of parameters in large language models.However, they can only save approximately 30% of the training memory requirements because gradient computation and backpropagation are still necessary for these methods.This paper proposes a novel parameter-efficient tuning method for LLMs without calculating their gradients.Leveraging the discernible similarities between the parameter-efficient modules of the same task learned by both large and small language models, we put forward a strategy for transferring the parameter-efficient modules derived initially from small language models to much larger ones.To ensure a smooth and effective adaptation process, we introduce a Bridge model to guarantee dimensional consistency while stimulating a dynamic interaction between the models.We demonstrate the effectiveness of our method using the T5 and GPT-2 series of language models on the SuperGLUE benchmark.Our method achieves comparable performance to fine-tuning and parameterefficient tuning on large language models without needing gradient-based optimization.Additionally, our method achieves up to 5.7× memory reduction compared to parameter-efficient tuning. Feihu Jin, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 3 |
| 2023 | Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningabstractTransformer-based models, even though achieving super-human performance on several downstream tasks, are often regarded as a black box and used as a whole.It is still unclear what mechanisms they have learned, especially their core module: multi-head attention.Inspired by functional specialization in the human brain, which helps to efficiently handle multiple tasks, this work attempts to figure out whether the multi-head attention module will evolve similar function separation under multitasking training.If it is, can this mechanism further improve the model performance?To investigate these questions, we introduce an interpreting method to quantify the degree of functional specialization in multi-head attention.We further propose a simple multi-task training method to increase functional specialization and mitigate negative information transfer in multi-task learning.Experimental results on seven pre-trained transformer models have demonstrated that multi-head attention does evolve functional specialization phenomenon after multi-task training which is affected by the similarity of tasks.Moreover, the multi-task training strategy based on functional specialization boosts performance in both multi-task learning and transfer learning without adding any parameters. 1 Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 5 |
| 2023 | Multi-teacher Knowledge Distillation for End-to-End Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (1) | 6 |
| 2023 | E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (6) | 6 |
| 2023 | Multi-source domain adaptation method for textual emotion classification using deep and broad learning
Sancheng Peng, Lihong Cao, Jianwei Niu 0002, Chengqing Zong, Guodong Zhou 0001 |
Knowl. Based Syst. | 6 |
| 2023 | Zero-shot language extension for dialogue state tracking via pre-trained models and multi-auxiliary-tasks fine-tuning
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong |
Knowl. Based Syst. | 5 |
| 2023 | Contrastive Adversarial Training for Multi-Modal Machine TranslationabstractThe multi-modal machine translation task is to improve translation quality with the help of additional visual input. It is expected to disambiguate or complement semantics while there are ambiguous words or incomplete expressions in the sentences. Existing methods have tried many ways to fuse visual information into text representations. However, only a minority of sentences need extra visual information as complementary. Without guidance, models tend to learn text-only translation from the major well-aligned translation pairs. In this article, we propose a contrastive adversarial training approach to enhance visual participation in semantic representation learning. By contrasting multi-modal input with the adversarial samples, the model learns to identify the most informed sample that is coupled with a congruent image and several visual objects extracted from it. This approach can prevent the visual information from being ignored and further fuse cross-modal information. We examine our method in three multi-modal language pairs. Experimental results show that our model is capable of improving translation accuracy. Further analysis shows that our model is more sensitive to visual information. Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Instance-Aware Prompt Learning for Language Understanding and GenerationabstractPrompt learning has emerged as a new paradigm for leveraging pre-trained language models (PLMs) and has shown promising results in downstream tasks with only a slight increase in parameters. However, the current usage of fixed prompts, whether discrete or continuous, assumes that all samples within a task share the same prompt. This assumption may not hold for tasks with diverse samples that require different prompt information. To address this issue, we propose an instance-aware prompt learning method that learns a different prompt for each instance. Specifically, we suppose that each learnable prompt token has a different contribution to different instances, and we learn the contribution by calculating the relevance score between an instance and each prompt token. The contribution-weighted prompt would be instance aware. We apply our method to both unidirectional and bidirectional PLMs on both language understanding and generation tasks. Extensive experiments demonstrate that our method achieves comparable results using as few as 1.5% of the parameters of PLMs tuned and obtains considerable improvements compared with strong baselines. In particular, our method achieves state-of-the-art results using ALBERT-xxlarge-v2 on the SuperGLUE few-shot learning benchmark. 1 Feihu Jin, Jinliang Lu, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Learning Category Distribution for Text ClassificationabstractLabel smoothing has a wide range of applications in the machine learning field. Nonetheless, label smoothing only softens the targets by adding a uniform distribution into a one-hot vector, which cannot truthfully reflect the underlying relations among categories. However, learning category relations is of vital importance in many fields such as emotion taxonomy and open set recognition. In this work, we propose a method to obtain the label distribution for each category (category distribution) to reveal category relations. Furthermore, based on the learned category distribution, we calculate new soft targets to improve the performance of model classification. Compared with existing methods, our algorithm can improve neural network models without any side information or additional neural network module by considering category relations. Extensive experiments have been conducted on four original datasets and 10 constructed noisy datasets with three basic neural network models to validate our algorithm. The results demonstrate the effectiveness of our algorithm on the classification task. In addition, three experiments (arrangement, clustering, and similarity) are also conducted to validate the intrinsic quality of the learned category distribution. The results indicate that the learned category distribution can well express underlying relations among categories. Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Combination of Loss-based Active Learning and Semi-supervised Learning for Recognizing Entities in Chinese Electronic Medical RecordsabstractThe recognition of entities in an electronic medical record (EMR) is especially important to downstream tasks, such as clinical entity normalization and medical dialogue understanding. However, in the medical professional field, training a high-quality named entity recognition system always requires large-scale annotated datasets, which are highly expensive to obtain. In this article, to lower the cost of data annotation and maximizing the use of unlabeled data, we propose a hybrid approach to recognizing the entities in Chinese electronic medical record, which is in combination of loss-based active learning and semi-supervised learning. Specifically, we adopted a dynamic balance strategy to dynamically balance the minimum loss predicted by a named entity recognition decoder and a loss prediction module at different stages in the process. Experimental results demonstrated our proposed framework’s effectiveness and efficiency, achieving higher performances than existing approaches on Chinese EMR entity recognition datasets under limited labeling resources. Jinghui Yan, Chengqing Zong, Jin An Xu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Topic-Oriented Dialogue SummarizationabstractA multi-turn dialogue often contains multiple discussion topics. In several scenarios (e.g., customer service dispute, public opinion monitoring), people are only interested in the gist of a specific topic in the dialogue. Therefore, we propose a novel summarization task, i.e., Topic-Oriented Dialogue Summarization (TODS). Given a dialogue with a topic label, TODS aims to produce a summary covering the main content of the given topic in the dialogue. To model the relationship between dialogues and topics, three key abilities are needed for TODS: (1) Learning the semantic information of different topics. (2) Locating the topic-related content in the dialogue. (3) Distinguishing summaries for different topics in the same dialogue. Thus, we propose three topic-related auxiliary tasks to make the summarization model learn the three abilities above. First, the topic identification task aims at generating all the topics in the dialogue. Second, the topic attention restriction task tries to constrain the attention distribution on topic-related utterances. Third, the topic summary distinguishing task focuses on increasing the difference of summaries for different topics in the same dialogue. Experimental results on two public TODS datasets show that all auxiliary tasks are critical for TODS and help generate high-quality summaries. We also point out the expansions and challenges in TODS for future research. Haitao Lin 0001, Junnan Zhu, Lu Xiang, Feifei Zhai, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2022 | Probing Word Syntactic Representations in the Brain by a Feature Elimination MethodabstractNeuroimaging studies have identified multiple brain regions that are associated with semantic and syntactic processing when comprehending language. However, existing methods cannot explore the neural correlates of fine-grained word syntactic features, such as part-of-speech and dependency relations. This paper proposes an alternative framework to study how different word syntactic features are represented in the brain. To separate each syntactic feature, we propose a feature elimination method, called Mean Vector Null space Projection (MVNP). This method can remove a specific feature from word representations, resulting in one-feature-removed representations. Then we respectively associate one-feature-removed and the original word vectors with brain imaging data to explore how the brain represents the removed feature. This paper for the first time studies the cortical representations of multiple fine-grained syntactic features simultaneously and suggests some possible contributions of several brain regions to the complex division of syntactic processing. These findings indicate that the brain foundations of syntactic information processing might be broader than those suggested by classical studies. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 5 |
| 2022 | Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role InteractionsabstractRole-oriented dialogue summarization is to generate summaries for different roles in the dialogue, e.g., merchants and consumers.Existing methods handle this task by summarizing each role's content separately and thus are prone to ignore the information from other roles.However, we believe that other roles' content could benefit the quality of summaries, such as the omitted information mentioned by other roles.Therefore, we propose a novel role interaction enhanced method for role-oriented dialogue summarization.It adopts cross attention and decoder self-attention interactions to interactively acquire other roles' critical information.The cross attention interaction aims to select other roles' critical dialogue utterances, while the decoder self-attention interaction aims to obtain key information from other roles' summaries.Experimental results have shown that our proposed method significantly outperforms strong baselines on two public role-oriented dialogue summarization datasets.Extensive analyses have demonstrated that other roles' content could help generate summaries with more complete semantics and correct topic structures. 1 Haitao Lin 0001, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 6 |
| 2022 | Discrete Cross-Modal Alignment Enables Zero-Shot Speech TranslationabstractEnd-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions.However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain.Fortunately, the supervised data for automatic speech recognition (ASR) and machine translation (MT) are usually more accessible, making zero-shot speech translation a potential direction.Existing zero-shot methods fail to align the two modalities of speech and text into a shared semantic space, resulting in much worse performance compared to the supervised ST methods.In order to enable zero-shot ST, we propose a novel Discrete Cross-Modal Alignment (DCMA) method that employs a shared discrete vocabulary space to accommodate and match both modalities of speech and text.Specifically, we introduce a vector quantization module to discretize the continuous representations of speech and text into a finite set of virtual tokens, and use ASR data to map corresponding speech and text to the same virtual token in a shared codebook.This way, source language speech can be embedded in the same semantic space as the source language text, which can be then transformed into target language text with an MT module.Experiments on multiple language pairs demonstrate that our zero-shot ST method significantly improves the SOTA, and even performs on par with the strong supervised ST baselines 1 . Yuchen Liu 0007, Boxing Chen, Jiajun Zhang 0001, Zhongqiang Huang, Chengqing Zong |
EMNLP | 7 |
| 2022 | Is the Brain Mechanism for Hierarchical Structure Building Universal Across Languages? An fMRI Study of Chinese and EnglishabstractEvidence from psycholinguistic studies suggests that the human brain builds a hierarchical syntactic structure during language comprehension.However, it is still unknown whether the neural basis of such structures is universal across languages.In this paper, we first analyze the differences in language structure between two diverse languages: Chinese and English.By computing the working memory requirements when applying parsing strategies to different language structures, we find that top-down parsing generates less memory load for the right-branching English and bottomup parsing is less memory-demanding for Chinese.Then we use functional magnetic resonance imaging (fMRI) to investigate whether the brain has different syntactic adaptation strategies in processing Chinese and English.Specifically, for both Chinese and English, we extract predictors from the implementations of different parsing strategies, i.e., bottom-up and top-down.Then, these predictors are separately associated with fMRI signals.Results show that for Chinese and English, the brain utilizes bottom-up and top-down parsing strategies separately.These results suggest that the brain adopts parsing strategies with less memory load according to different language structures. Shaonan Wang, Chengqing Zong |
EMNLP | 4 |
| 2022 | How Does the Experimental Setting Affect the Conclusions of Neural Encoding Models?abstractRecent years have witnessed the tendency of neural encoding models on exploring brain language processing using naturalistic stimuli. Neural encoding models are data-driven methods that require an encoding model to investigate the mystery of brain mechanisms hidden in the data. As a data-driven method, the performance of encoding models is very sensitive to the experimental setting. However, it is unknown how the experimental setting further affects the conclusions of neural encoding models. This paper systematically investigated this problem and evaluated the influence of three experimental settings, i.e., the data size, the cross-validation training method, and the statistical testing method. Results demonstrate that inappropriate cross-validation training and small data size can substantially decrease the performance of encoding models, especially in the temporal lobe and the frontal lobe. And different null hypotheses in significance testing lead to highly different significant brain regions. Based on these results, we suggest a block-wise cross-validation training method and an adequate data size for increasing the performance of linear encoding models. We also propose two strict null hypotheses to control false positive discovery rates. Shaonan Wang, Chengqing Zong |
LREC | 3 |
| 2022 | Enhancing Lexical Translation Consistency for Document-Level Neural Machine TranslationabstractDocument-level neural machine translation (DocNMT) has yielded attractive improvements. In this article, we systematically analyze the discourse phenomena in Chinese-to-English translation, and focus on the most obvious ones, namely lexical translation consistency. To alleviate the lexical inconsistency, we propose an effective approach that is aware of the words which need to be translated consistently and constrains the model to produce more consistent translations. Specifically, we first introduce a global context extractor to extract the document context and consistency context, respectively. Then, the two types of global context are integrated into a encoder enhancer and a decoder enhancer to improve the lexical translation consistency. We create a test set to evaluate the lexical consistency automatically. Experiments demonstrate that our approach can significantly alleviate the lexical translation inconsistency. In addition, our approach can also substantially improve the translation quality compared to sentence-level Transformer. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Dual-View Conditional Variational Auto-Encoder for Emotional Dialogue GenerationabstractEmotional dialogue generation aims to generate appropriate responses that are content relevant with the query and emotion consistent with the given emotion tag. Previous work mainly focuses on incorporating emotion information into the sequence to sequence or conditional variational auto-encoder (CVAE) models, and they usually utilize the given emotion tag as a conditional feature to influence the response generation process. However, emotion tag as a feature cannot well guarantee the emotion consistency between the response and the given emotion tag. In this article, we propose a novel Dual-View CVAE model to explicitly model the content relevance and emotion consistency jointly. These two views gather the emotional information and the content-relevant information from the latent distribution of responses, respectively. We jointly model the dual-view via VAE to get richer and complementary information. Extensive experiments on both English and Chinese emotion dialogue datasets demonstrate the effectiveness of our proposed Dual-View CVAE model, which significantly outperforms the strong baseline models in both aspects of content relevance and emotion consistency. Jiajun Zhang 0001, Lu Xiang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | One-Shot Relation Learning for Knowledge Graphs via Neighborhood Aggregation and Paths EncodingabstractThe relation learning between two entities is an essential task in knowledge graph (KG) completion that has received much attention recently. Previous work almost exclusively focused on relations widely seen in the original KGs, which means that enough training data are available for modeling. However, long-tail relations that only show in a few triples are actually much more common in practical KGs. Without sufficiently large training data, the performance of existing models on predicting long-tail relations drops impressively. This work aims to predict the relation under a challenging setting where only one instance is available for training. We propose a path-based one-shot relation prediction framework, which can extract neighborhood information of an entity based on the relation query attention mechanism to learn transferable knowledge among the same relation. Simultaneously, to reduce the impact of long-tail entities on relation prediction, we selectively fuse path information between entity pairs as auxiliary information of relation features. Experiments in three one-shot relation learning datasets show that our proposed framework substantially outperforms existing models on one-shot link prediction and relation prediction. Jian Sun 0035, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | Attention Analysis and Calibration for Transformer in Natural Language GenerationabstractAttention mechanism has been ubiquitous in neural machine translation by dynamically selecting relevant contexts for different translations. Apart from performance gains, attention weights assigned to input tokens are often utilized to explain that high-attention tokens contribute more to the prediction. However, many works question whether this assumption holds in text classification by manually manipulating attention weights and observing decision flips. This article extends this question to Transformer-based neural machine translation, which heavily relies on cross-lingual attention to produce accurate translations but is relatively understudied in this context. We first design a mask perturbation model which automatically assesses each input’s contribution to model outputs. We then test whether the token contributing most to the current translation receives the highest attention weight. We find that it sometimes does not, which closely depends on the entropy of attention weights, the syntactic role of the current generation, and language pairs. We also rethink the discrepancy between attention weights and word alignments from the view of unreliable attention weights. Our observations further motivate us to calibrate the cross-lingual multi-head attention by attaching more attention to indispensable tokens, whose removal leads to a dramatic performance drop. Empirical experiments on different-scale translation tasks and text summarization tasks demonstrate that our calibration methods significantly outperform strong baselines. Jiajun Zhang 0001, Jiali Zeng, Shuangzhi Wu, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Synchronous Inference for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation allows a single model to translate between multiple language pairs, which greatly reduces the cost of model training and receives much attention recently. Previous studies mainly focus on training stage optimization and improve positive knowledge transfer among languages with different levels of parameter sharing, but ignore the multilingual knowledge transfer during inference although the translation in one language may help the generation of other languages. This work enhances knowledge sharing among multiple target languages in the inference phase. To achieve this, we propose a synchronous inference method that can simultaneously generate translations in multiple languages. During generation, the model predicts the next word of each language not only based on source sentence and previously predicted segments, but also based on predicted words of other target languages. To maximize the inference stage knowledge sharing, we design a cross-lingual attention module which allows the model to dynamically select the most relevant information from multiple target languages. The synchronous inference model requires multi-way parallel training data which is scarce. We therefore propose to adopt multi-task learning to incorporate large-scale bilingual data. We evaluate our method on three multilingual translation datasets and prove that the proposed method significantly improve the translation quality and the decoding efficiency compared to strong bilingual and multilingual baselines. Qian Wang 0061, Jiajun Zhang 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Synchronous Interactive Decoding for Multilingual Neural Machine TranslationabstractTo simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future information, and therefore may suffer from some issues like the unbalanced outputs. In this paper, we present a new approach for synchronous interactive multilingual neural machine translation (SimNMT), which predicts each target language output simultaneously and interactively using historical and future information of all target languages. Specifically, we first propose a synchronous cross-interactive decoder in which generation of each target output does not only depend on its generated sequences, but also relies on its future information, as well as history and future contexts of other target languages. Then, we present a new interactive multilingual beam search algorithm that enables synchronous interactive decoding of all target languages in a single model. We take two target languages as an example to illustrate and evaluate the proposed SimNMT model on IWSLT datasets. The experimental results demonstrate that our method achieves significant improvements over several advanced NMT and MNMT models. Qian Wang 0061, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 6 |
| 2021 | Distributed Representations of Emotion Categories in Emotion SpaceabstractXiangyu Wang, Chengqing Zong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chengqing Zong |
ACL/IJCNLP (1) | 2 |
| 2021 | CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue SummarizationabstractDialogue summarization has drawn much attention recently.Especially in the customer service domain, agents could use dialogue summaries to help boost their works by quickly knowing customer's issues and service progress.These applications require summaries to contain the perspective of a single speaker and have a clear topic flow structure, while neither are available in existing datasets.Therefore, in this paper, we introduce a novel Chinese dataset for Customer Service Dialogue Summarization (CSDS).CSDS improves the abstractive summaries in two aspects: (1) In addition to the overall summary for the whole dialogue, role-oriented summaries are also provided to acquire different speakers' viewpoints.(2) All the summaries sum up each topic separately, thus containing the topic-level structure of the dialogue.We define tasks in CSDS as generating the overall summary and different role-oriented summaries for a given dialogue.Next, we compare various summarization methods on CSDS, and experiment results show that existing methods are prone to generate redundant and incoherent summaries.Besides, the performance becomes much worse when analyzing the performance on role-oriented summaries and topic structures.We hope that this study could benchmark Chinese dialogue summarization and benefit further studies. Haitao Lin 0001, Liqun Ma, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
EMNLP (1) | 7 |
| 2021 | Augmenting Slot Values and Contexts for Spoken Language Understanding with Pretrained ModelsabstractSpoken Language Understanding (SLU) is one essential step in building a dialogue system. Due to the expensive cost of obtaining the labeled data, SLU suffers from the data scarcity problem. Therefore, in this paper, we focus on data augmentation for slot filling task in SLU. To achieve that, we aim at generating more diverse data based on existing data. Specifically, we try to exploit the latent language knowledge from pretrained language models by finetuning them. We propose two strategies for finetuning process: value-based and context-based augmentation. Experimental results on two public SLU datasets have shown that compared with existing data augmentation methods, our proposed method can generate more diverse sentences and significantly improve the performance on SLU. Haitao Lin 0001, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
Interspeech | 5 |
| 2021 | Zero-Shot Deployment for Cross-Lingual Dialogue System
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong |
NLPCC (2) | 5 |
| 2021 | Robust Cross-lingual Task-oriented DialogueabstractCross-lingual dialogue systems are increasingly important in e-commerce and customer service due to the rapid progress of globalization. In real-world system deployment, machine translation (MT) services are often used before and after the dialogue system to bridge different languages. However, noises and errors introduced in the MT process will result in the dialogue system's low robustness, making the system's performance far from satisfactory. In this article, we propose a novel MT-oriented noise enhanced framework that exploits multi-granularity MT noises and injects such noises into the dialogue system to improve the dialogue system's robustness. Specifically, we first design a method to automatically construct multi-granularity MT-oriented noises and multi-granularity adversarial examples, which contain abundant noise knowledge oriented to MT. Then, we propose two strategies to incorporate the noise knowledge: (i) Utterance-level adversarial learning and (ii) Knowledge-level guided method. The former adopts adversarial learning to learn a perturbation-invariant encoder, guiding the dialogue system to learn noise-independent hidden representations. The latter explicitly incorporates the multi-granularity noises, which contain the noise tokens and their possible correct forms, into the training and inference process, thus improving the dialogue system's robustness. Experimental results on three dialogue models, two dialogue datasets, and two language pairs have shown that the proposed framework significantly improves the performance of the cross-lingual dialogue system. Lu Xiang, Junnan Zhu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | Graph-based Multimodal Ranking Models for Multimodal SummarizationabstractMultimodal summarization aims to extract the most important information from the multimedia input. It is becoming increasingly popular due to the rapid growth of multimedia data in recent years. There are various researches focusing on different multimodal summarization tasks. However, the existing methods can only generate single-modal output or multimodal output. In addition, most of them need a lot of annotated samples for training, which makes it difficult to be generalized to other tasks or domains. Motivated by this, we propose a unified framework for multimodal summarization that can cover both single-modal output summarization and multimodal output summarization. In our framework, we consider three different scenarios and propose the respective unsupervised graph-based multimodal summarization models without the requirement of any manually annotated document-summary pairs for training: (1) generic multimodal ranking, (2) modal-dominated multimodal ranking, and (3) non-redundant text-image multimodal ranking. Furthermore, an image-text similarity estimation model is introduced to measure the semantic similarity between image and text. Experiments show that our proposed models outperform the single-modal summarization methods on both automatic and human evaluation metrics. Besides, our models can also improve the single-modal summarization with the guidance of the multimedia information. This study can be applied as the benchmark for further study on multimodal summarization task. Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | Medical Term and Status Generation From Chinese Clinical Dialogue With Multi-Granularity TransformerabstractThis paper describes a generative model for extracting medical terms and their status from Chinese medical dialogues. Notably, the extracted semantic information plays an essential role in downstream tasks such as automatic medical scribe and automatic diagnosis system. However, how to effectively leverage dialogue context to generate medical terms and their corresponding status accurately remains less explored. Existing generative methods treat dialogue text as concentrated long text without considering the characteristics of conversation, such as colloquialism, redundancy, interactions, etc. Various colloquial medical information is frequently discussed between doctor and patient. Each of the speakers (doctor and patient) plays a specific role in the goals of interaction. Thus the role information and interactions between utterances are vital. Besides, current generative methods only utilize character-level tokens ignoring the word-level tokens, which is the smallest meaningful utterance in Chinese. In this paper, we propose a Multi-granularity Transformer (MGT) model to enhance the dialogue context understanding from multi-granularity features. We introduce word-level information by adapting a Lattice-based encoder with our proposed relative position encoding method. We further introduce utterance-level interaction information by proposing a Role Access Controlled Attention (RaCa) mechanism. Experimental results on two benchmark datasets illustrate our model's validity and effectiveness, achieving state-of-the-art performance on both datasets. Lu Xiang, Xiaomian Kang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Neural Encoding and Decoding With Distributed Sentence RepresentationsabstractBuilding computational models to account for the cortical representation of language plays an important role in understanding the human linguistic system. Recent progress in distributed semantic models (DSMs), especially transformer-based methods, has driven advances in many language understanding tasks, making DSM a promising methodology to probe brain language processing. DSMs have been shown to reliably explain cortical responses to word stimuli. However, characterizing the brain activities for sentence processing is much less exhaustively explored with DSMs, especially the deep neural network-based methods. What is the relationship between cortical sentence representations against DSMs? What linguistic features that a DSM catches better explain its correlation with the brain activities aroused by sentence stimuli? Could distributed sentence representations help to reveal the semantic selectivity of different brain areas? We address these questions through the lens of neural encoding and decoding, fueled by the latest developments in natural language representation learning. We begin by evaluating the ability of a wide range of 12 DSMs to predict and decipher the functional magnetic resonance imaging (fMRI) images from humans reading sentences. Most models deliver high accuracy in the left middle temporal gyrus (LMTG) and left occipital complex (LOC). Notably, encoders trained with transformer-based DSMs consistently outperform other unsupervised structured models and all the unstructured baselines. With probing and ablation tasks, we further find that differences in the performance of the DSMs in modeling brain activities can be at least partially explained by the granularity of their semantic representations. We also illustrate the DSM's selectivity for concept categories and show that the topics are represented by spatially overlapping and distributed cortical patterns. Our results corroborate and extend previous findings in understanding the relation between DSMs and neural activation patterns and contribute to building solid brain-machine interfaces with deep neural network representations. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Keywords-Guided Abstractive Sentence SummarizationabstractWe study the problem of generating a summary for a given sentence. Existing researches on abstractive sentence summarization ignore that keywords in the input sentence provide significant clues for valuable content, and humans tend to write summaries covering these keywords. In this paper, we propose an abstractive sentence summarization method by applying guidance signals of keywords to both the encoder and the decoder in the sequence-to-sequence model. A multi-task learning framework is adopted to jointly learn to extract keywords and generate a summary for the input sentence. We apply keywords-guided selective encoding strategies to filter source information by investigating the interactions between the input sentence and the keywords. We extend pointer-generator network by a dual-attention and a dual-copy mechanism, which can integrate the semantics of the input sentence and the keywords, and copy words from both the input sentence and the keywords. We demonstrate that multi-task learning and keywords-oriented guidance facilitate sentence summarization task, achieving better performance than the competitive models on the English Gigaword sentence summarization dataset. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong, Xiaodong He 0001 |
AAAI | 4 |
| 2020 | Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingabstractSpeech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential benefits of lower latency, smaller model size, and less error propagation. However, it is notoriously difficult to implement such a model without transcriptions as intermediate. Existing works generally apply multi-task learning to improve translation quality by jointly training end-to-end ST along with automatic speech recognition (ASR). However, different tasks in this method cannot utilize information from each other, which limits the improvement. Other works propose a two-stage model where the second model can use the hidden state from the first one, but its cascade manner greatly affects the efficiency of training and inference process. In this paper, we propose a novel interactive attention mechanism which enables ASR and ST to perform synchronously and interactively in a single model. Specifically, the generation of transcriptions and translations not only relies on its previous outputs but also the outputs predicted in the other task. Experiments on TED speech translation corpora have shown that our proposed model can outperform strong baselines on the quality of speech translation and achieve better speech recognition performances as well. Yuchen Liu 0007, Jiajun Zhang 0001, Hao Xiong 0005, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
AAAI | 8 |
| 2020 | Probing Brain Activation Patterns by Dissociating Semantics and Syntax in SentencesabstractThe relation between semantics and syntax and where they are represented in the neural level has been extensively debated in neurosciences. Existing methods use manually designed stimuli to distinguish semantic and syntactic information in a sentence that may not generalize beyond the experimental setting. This paper proposes an alternative framework to study the brain representation of semantics and syntax. Specifically, we embed the highly-controlled stimuli as objective functions in learning sentence representations and propose a disentangled feature representation model (DFRM) to extract semantic and syntactic information in sentences. This model can generate one semantic and one syntactic vector for each sentence. Then we associate these disentangled feature vectors with brain imaging data to explore brain representation of semantics and syntax. Results have shown that semantic feature is represented more robustly than syntactic feature across the brain including the default-mode, frontoparietal, visual networks, etc.. The brain representations of semantics and syntax are largely overlapped, but there are brain regions only sensitive to one of them. For instance, several frontal and temporal regions are specific to the semantic feature; parts of the right superior frontal and right inferior parietal gyrus are specific to the syntactic feature. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 4 |
| 2020 | Multimodal Summarization with Guidance of Multimodal ReferenceabstractMultimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Order-ranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments. Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Haoran Li 0001, Chengqing Zong, Changliang Li |
AAAI | 5 |
| 2020 | Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual SummarizationabstractCross-lingual summarization aims at summarizing a document in one language (e.g., Chinese) into another language (e.g., English).In this paper, we propose a novel method inspired by the translation pattern in the process of obtaining a cross-lingual summary.We first attend to some words in the source text, then translate them into the target language, and summarize to get the final summary.Specifically, we first employ the encoder-decoder attention distribution to attend to the source words.Second, we present three strategies to acquire the translation probability, which helps obtain the translation candidates for each source word.Finally, each summary word is generated either from the neural distribution or from the translation candidates of source words.Experimental results on Chinese-to-English and English-to-Chinese summarization tasks have shown that our proposed method can significantly outperform the baselines, achieving comparable performance with the state-of-the-art. Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL | 4 |
| 2020 | Multimodal Sentence Summarization via Multimodal Selective EncodingabstractThis paper studies the problem of generating a summary for a given sentence-image pair.Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news event or a document.Thus, we propose a multimodal selective gate network that considers reciprocal relationships between textual and multi-level visual features, including global image descriptor, activation grids, and object proposals, to select highlights of the event when encoding the source sentence.In addition, we introduce a modality regularization to encourage the summary to capture the highlights embedded in the image more accurately.To verify the generalization of our model, we adopt the multimodal selective gate to the text-based decoder and multimodal-based decoder.Experimental results on a public multimodal sentence summarization dataset demonstrate the advantage of our models over baselines.Further analysis suggests that our proposed multimodal selective gate network can effectively select important information in the input sentence. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Xiaodong He 0001, Chengqing Zong |
COLING | 5 |
| 2020 | Distill and Replay for Continual Language LearningabstractAccumulating knowledge to tackle new tasks without necessarily forgetting the old ones is a hallmark of human-like intelligence.But the current dominant paradigm of machine learning is still to train a model that works well on static datasets.When learning tasks in a stream where data distribution may fluctuate, fitting on new tasks often leads to forgetting on the previous ones.We propose a simple yet effective framework that continually learns natural language understanding tasks with one model.Our framework distills knowledge and replays experience from previous tasks when fitting on a new task, thus named DnR (distill and replay).The framework is based on language models and can be smoothly built with different language model architectures.Experimental results demonstrate that DnR outperfoms previous state-of-the-art models in continually learning tasks of the same type but from different domains, as well as tasks of radically different types.With the distillation method, we further show that it's possible for DnR to incrementally compress the model size while still outperforming most of the baselines.We hope that DnR could promote the empirical application of continual language learning, and contribute to building human-level language intelligence minimally bothered by catastrophic forgetting. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
COLING | 4 |
| 2020 | Dual Attention Network for Cross-lingual Entity AlignmentabstractCross-lingual Entity alignment is an essential part of building a knowledge graph, which can help integrate knowledge among different language knowledge graphs.In the real KGs, there exists an imbalance among the information in the same hierarchy of corresponding entities, which results in the heterogeneity of neighborhood structure, making this task challenging.To tackle this problem, we propose a dual attention network for cross-lingual entity alignment (DAEA).Specifically, our dual attention consists of relation-aware graph attention and hierarchical attention.The relation-aware graph attention aims at selectively aggregating multi-hierarchy neighborhood information to alleviate the difference of heterogeneity among counterpart entities.The hierarchical attention adaptively aggregates the low-hierarchy and the high-hierarchy information, which is beneficial to balance the neighborhood information of counterpart entities and distinguish noncounterpart entities with similar structures.Finally, we treat cross-lingual entity alignment as a process of linking prediction.Experimental results on three real-world cross-lingual entity alignment datasets have shown the effectiveness of DAEA. Jian Sun 0035, Yu Zhou 0001, Chengqing Zong |
COLING | 3 |
| 2020 | Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity GranularityabstractPrevious studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized.ii) Granularity mismatch: the current KG methods utilize the entity as the basic granularity, while NMT utilizes the sub-word as the granularity, making the KG different to be utilized in NMT.To alleviate above problems, we propose a multi-task learning method on sub-entity granularity.Specifically, we first split the entities in KG and sentence pairs into sub-entity granularity by using joint BPE.Then we utilize the multi-task learning to combine the machine translation task and knowledge reasoning task.The extensive experiments on various translation tasks have demonstrated that our method significantly outperforms the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Lu Xiang, Junnan Zhu, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 6 |
| 2020 | Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningabstractDocument-level neural machine translation has yielded attractive improvements.However, majority of existing methods roughly use all context sentences in a fixed scope.They neglect the fact that different source sentences need different sizes of context.To address this problem, we propose an effective approach to select dynamic context so that the document-level translation model can utilize the more useful selected context sentences to produce better translations.Specifically, we introduce a selection module that is independent of the translation module to score each candidate context sentence.Then, we propose two strategies to explicitly select a variable number of context sentences and feed them into the translation module.We train the two modules end-to-end via reinforcement learning.A novel reward is proposed to encourage the selection and utilization of dynamic context sentences.Experiments demonstrate that our approach can select adaptive context sentences for different source sentences, and significantly improves the performance of document-level translation methods. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
EMNLP (1) | 4 |
| 2020 | A Knowledge-driven Generative Model for Multi-implication Chinese Medical Procedure Entity NormalizationabstractMedical entity normalization, which links medical mentions in the text to entities in knowledge bases, is an important research topic in medical natural language processing.In this paper, we focus on Chinese medical procedure entity normalization.However, nonstandard Chinese expressions and combined procedures present challenges in our problem.The existing strategies relying on the discriminative model are poorly to cope with normalizing combined procedure mentions.We propose a sequence generative framework to directly generate all the corresponding medical procedure entities.we adopt two strategies: category-based constraint decoding and category-based model refining to avoid unrealistic results.The method is capable of linking entities when a mention contains multiple procedure concepts and our comprehensive experiments demonstrate that the proposed model can achieve remarkable improvements over existing baselines, particularly significant in the case of multi-implication Chinese medical procedures. Jinghui Yan, Lu Xiang, Yu Zhou 0001, Chengqing Zong |
EMNLP (1) | 5 |
| 2020 | Knowledge Graphs Enhanced Neural Machine TranslationabstractKnowledge graphs (KGs) store much structured information on various entities, many of which are not covered by the parallel sentence pairs of neural machine translation (NMT). To improve the translation quality of these entities, in this paper we propose a novel KGs enhanced NMT method. Specifically, we first induce the new translation results of these entities by transforming the source and target KGs into a unified semantic space. We then generate adequate pseudo parallel sentence pairs that contain these induced entity pairs. Finally, NMT model is jointly trained by the original and pseudo sentence pairs. The extensive experiments on Chinese-to-English and Englishto-Japanese translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling the induced entities. Yang Zhao 0007, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IJCAI | 4 |
| 2020 | Non-autoregressive Neural Machine Translation with Distortion Model
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
NLPCC (1) | 4 |
| 2020 | Synchronous bidirectional inference for neural sequence generation
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
Artif. Intell. | 4 |
| 2020 | Fine-grained neural decoding with distributed word representations
Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
Inf. Sci. | 5 |
| 2020 | Conducting Natural Language Inference with Word-Pair-Dependency and Local ContextabstractThis article proposes to conduct natural language inference with novel Enhanced-Relation-Head-Dependent triplets (RHD triplets) , which are constructed via enhancing each word in the RHD triplet with its associated local context. Most previous approaches based on deep neural network (DNN) for this task either perform token alignment without considering syntactic dependency among words, or directly use tree- LSTM to generate passage representation with irrelevant information. To improve token alignment and inference judgment with word-pair-dependency, the RHD triplet structure is first proposed. To avoid incorporating irrelevant information, this proposed approach performs comparison directly on each triplet-pair of the given passage-pair (instead of comparing each triplet in a passage with the content merged from the whole opposite passage). Furthermore, to take local context into consideration while conducting token alignment and inference judgment, we also enhance the words of the triplets with their associated local context to improve the performance. Experimental results show that the proposed approach is better than most previous approaches that adopt tree structures, and its performance is comparable to other state-of-the-art approaches (however, our approach is more human comprehensible). Qianlong Du, Chengqing Zong, Keh-Yih Su |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | Structurally Comparative Hinge Loss for Dependency-Based Neural Text RepresentationabstractDependency-based graph convolutional networks (DepGCNs) are proven helpful for text representation to handle many natural language tasks. Almost all previous models are trained with cross-entropy (CE) loss, which maximizes the posterior likelihood directly. However, the contribution of dependency structures is not well considered by CE loss. As a result, the performance improvement gained by using the structure information can be narrow due to the failure in learning to rely on this structure information. To face the challenge, we propose the novel structurally comparative hinge (SCH) loss function for DepGCNs. SCH loss aims at enlarging the margin gained by structural representations over non-structural ones. From the perspective of information theory, this is equivalent to improving the conditional mutual information of model decision and structure information given text. Our experimental results on both English and Chinese datasets show that by substituting SCH loss for CE loss on various tasks, for both induced structures and structures from an external parser, performance is improved without additional learnable parameters. Furthermore, the extent to which certain types of examples rely on the dependency structure can be measured directly by the learned margin, which results in better interpretability. In addition, through detailed analysis, we show that this structure margin has a positive correlation with task performance and structure induction of DepGCNs, and SCH loss can help model focus more on the shortest dependency path between entities. We achieve the new state-of-the-art results on TACRED, IMDB, and Zh. Literature datasets, even compared with ensemble and BERT baselines. Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2020 | Deep Neural Network-based Machine Translation System CombinationabstractDeep neural networks (DNNs) have provably enhanced the state-of-the-art natural language process (NLP) with their capability of feature learning and representation. As one of the more challenging NLP tasks, neural machine translation (NMT) becomes a new approach to machine translation and generates much more fluent results compared to statistical machine translation (SMT). However, SMT is usually better than NMT in translation adequacy and word coverage. It is therefore a promising direction to combine the advantages of both NMT and SMT. In this article, we propose a deep neural network--based system combination framework leveraging both minimum Bayes-risk decoding and multi-source NMT, which take as input the N-best outputs of NMT and SMT systems and produce the final translation. In particular, we apply the proposed model to both RNN and self-attention networks with different segmentation granularity. We verify our approach empirically through a series of experiments on resource-rich Chinese⇒English and low-resource English⇒Vietnamese translation tasks. Experimental results demonstrate the effectiveness and universality of our proposed approach, which significantly outperforms the conventional system combination methods and the best individual system output. Jiajun Zhang 0001, Xiaomian Kang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2019 | Towards Personalized Review Summarization via User-Aware Sequence NetworkabstractWe address personalized review summarization, which generates a condensed summary for a user’s review, accounting for his preference on different aspects or his writing style. We propose a novel personalized review summarization model named User-aware Sequence Network (USN) to consider the aforementioned users’ characteristics when generating summaries, which contains a user-aware encoder and a useraware decoder. Specifically, the user-aware encoder adopts a user-based selective mechanism to select the important information of a review, and the user-aware decoder incorporates user characteristic and user-specific word-using habits into word prediction process to generate personalized summaries. To validate our model, we collected a new dataset Trip, comprising 536,255 reviews from 19,400 users. With quantitative and human evaluation, we show that USN achieves state-ofthe-art performance on personalized review summarization. Haoran Li 0001, Chengqing Zong |
AAAI | 3 |
| 2019 | Towards Sentence-Level Brain Decoding with Distributed RepresentationsabstractDecoding human brain activities based on linguistic representations has been actively studied in recent years. However, most previous studies exclusively focus on word-level representations, and little is learned about decoding whole sentences from brain activation patterns. This work is our effort to mend the gap. In this paper, we build decoders to associate brain activities with sentence stimulus via distributed representations, the currently dominant sentence representation approach in natural language processing (NLP). We carry out a systematic evaluation, covering both widely-used baselines and state-of-the-art sentence representation models. We demonstrate how well different types of sentence representations decode the brain activation patterns and give empirical explanations of the performance difference. Moreover, to explore how sentences are neurally represented in the brain, we further compare the sentence representation’s correspondence to different brain areas associated with high-level cognitive functions. We find the supervised structured representation models most accurately probe the language atlas of human brain. To the best of our knowledge, this work is the first comprehensive evaluation of distributed sentence representations for brain decoding. We hope this work can contribute to decoding brain activities with NLP representation models, and understanding how linguistic items are neurally represented. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 4 |
| 2019 | Addressing the Under-Translation Problem from the Entropy PerspectiveabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance in recent years. However, the under-translation problem still remains a big challenge. In this paper, we focus on the under-translation problem and attempt to find out what kinds of source words are more likely to be ignored. Through analysis, we observe that a source word with a large translation entropy is more inclined to be dropped. To address this problem, we propose a coarse-to-fine framework. In coarse-grained phase, we introduce a simple strategy to reduce the entropy of highentropy words through constructing the pseudo target sentences. In fine-grained phase, we propose three methods, including pre-training method, multitask method and two-pass method, to encourage the neural model to correctly translate these high-entropy words. Experimental results on various translation tasks show that our method can significantly improve the translation quality and substantially reduce the under-translation cases of high-entropy words. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhongjun He, Hua Wu 0003 |
AAAI | 3 |
| 2019 | Memory Consolidation for Contextual Spoken Language Understanding with Dialogue Logistic InferenceabstractDialogue contexts are proven helpful in the spoken language understanding (SLU) system and they are typically encoded with explicit memory representations.However, most of the previous models learn the context memory with only one objective to maximizing the SLU performance, leaving the context memory under-exploited.In this paper, we propose a new dialogue logistic inference (DLI) task to consolidate the context memory jointly with SLU in the multi-task framework.DLI is defined as sorting a shuffled dialogue session into its original logical order and shares the same memory encoder and retrieval mechanism as the SLU model.Our experimental results show that various popular contextual SLU models can benefit from our approach, and improvements are quite impressive, especially in slot filling. He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 4 |
| 2019 | Incremental Learning from Scratch for Task-Oriented Dialogue SystemsabstractClarifying user needs is essential for existing task-oriented dialogue systems.However, in real-world applications, developers can never guarantee that all possible user demands are taken into account in the design phase.Consequently, existing systems will break down when encountering unconsidered user needs.To address this problem, we propose a novel incremental learning framework to design task-oriented dialogue systems, or for short Incremental Dialogue System (IDS), without pre-defining the exhaustive list of user needs.Specifically, we introduce an uncertainty estimation module to evaluate the confidence of giving correct responses.If there is high confidence, IDS will provide responses to users.Otherwise, humans will be involved in the dialogue process, and IDS can learn from human intervention through an online learning module.To evaluate our method, we propose a new dataset which simulates unanticipated user needs in the deployment stage.Experiments show that IDS is robust to unconsidered user actions, and can update itself online by smartly selecting only the most effective training data, and hence attains better performance with less annotation cost. 1 Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong, Zhifei Li 0001 |
ACL (1) | 5 |
| 2019 | A Compact and Language-Sensitive Multilingual Translation MethodabstractMultilingual neural machine translation (Multi-NMT) with one encoder-decoder model has made remarkable progress due to its simple deployment.However, this multilingual translation paradigm does not make full use of language commonality and parameter sharing between encoder and decoder.Furthermore, this kind of paradigm cannot outperform the individual models trained on bilingual corpus in most cases.In this paper, we propose a compact and language-sensitive method for multilingual translation.To maximize parameter sharing, we first present a universal representor to replace both encoder and decoder models.To make the representor sensitive for specific languages, we further introduce language-sensitive embedding, attention, and discriminator with the ability to enhance model performance.We verify our methods on various translation scenarios, including one-to-many, many-to-many and zero-shot.Extensive experiments demonstrate that our proposed methods remarkably outperform strong standard multilingual translation systems on WMT and IWSLT datasets.Moreover, we find that our model is especially helpful in low-resource and zero-shot translation scenarios. Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong |
ACL (1) | 6 |
| 2019 | Attribute-aware Sequence Network for Review SummarizationabstractJunjie Li, Xuepeng Wang, Dawei Yin, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xuepeng Wang, Dawei Yin 0001, Chengqing Zong |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Are You for Real? Detecting Identity Fraud via Dialogue InteractionsabstractWeikang Wang, Jiajun Zhang, Qian Li, Chengqing Zong, Zhifei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiajun Zhang 0001, Chengqing Zong, Zhifei Li 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Synchronously Generating Two Languages with Interactive DecodingabstractYining Wang, Jiajun Zhang, Long Zhou, Yuchen Liu, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiajun Zhang 0001, Yuchen Liu 0007, Chengqing Zong |
EMNLP/IJCNLP (1) | 5 |
| 2019 | NCLS: Neural Cross-Lingual SummarizationabstractJunnan Zhu, Qian Wang, Yining Wang, Yu Zhou, Jiajun Zhang, Shaonan Wang, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junnan Zhu, Qian Wang 0061, Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong |
EMNLP/IJCNLP (1) | 7 |
| 2019 | CATSLU: The 1st Chinese Audio-Textual Spoken Language Understanding ChallengeabstractSpoken language understanding (SLU) is a key component of conversational dialogue systems, which converts user utterances into semantic representations. The previous works almost focus on parsing semantic from textual inputs (top hypothesis of speech recognition and even manual transcripts) while losing information hidden in the audio. We herein describe the 1st Chinese Audio-Textual Spoken Language Understanding Challenge (CATSLU) which introduces a new dataset with audio-textual information, multiple domains and domain knowledge. We introduce two scenarios of audio-textual SLU in which participants are encouraged to utilize data of other domains or not. In this paper, we will describe the challenge and results. Su Zhu, Tiejun Zhao, Chengqing Zong, Kai Yu 0004 |
ICMI | 4 |
| 2019 | A New Effective Neural Variational Model with Mixture-of-Gaussians Prior for Text ClusteringabstractText clustering is one of the fundamental tasks in natural language processing and text data mining. It remains challenging because texts have complex internal structure besides the sparsity in the high-dimensional representation. In the paper, we propose a new Neural Variational model with mixture-of-Gaussians prior for Text Clustering (abbr. NVTC) to reveal the underlying textual manifold structure and cluster documents effectively. NVTC is a deep latent variable model built on the basis of the neural variational inference. In NVTC, the stochastic latent variable, which is modeled as one obeying a Gaussian mixture distribution, plays an important role in establishing the association of documents and document labels. On the other hand, by joint learning, NVTC simultaneously learns text encoded representations and cluster assignments. Experimental results demonstrate that NVTC is able to learn clustering-friendly representations of texts. It significantly outperforms several baselines including VAE+GMM, VaDE, LCK-NFC, GSDPMM and LDA on four benchmark text datasets in terms of ACC, NMI, and AMI. Furthermore, NVTC learns effective latent embeddings of texts which are interpretable by topics of texts, where each dimension of latent embeddings corresponds to a specific topic. Hongyin Tang, Beihong Jin, Chengqing Zong |
ICTAI | 4 |
| 2019 | Sequence Generation: From Both Sides to the MiddleabstractThe encoder-decoder framework has achieved promising process for many sequence generation tasks, such as neural machine translation and text summarization. Such a framework usually generates a sequence token by token from left to right, hence (1) this autoregressive decoding procedure is time-consuming when the output sentence becomes longer, and (2) it lacks the guidance of future context which is crucial to avoid under-translation. To alleviate these issues, we propose a synchronous bidirectional sequence generation (SBSG) model which predicts its outputs from both sides to the middle simultaneously. In the SBSG model, we enable the left-to-right (L2R) and right-to-left (R2L) generation to help and interact with each other by leveraging interactive bidirectional attention network. Experiments on neural machine translation (En-De, Ch-En, and En-Ro) and text summarization tasks show that the proposed model significantly speeds up decoding while improving the generation quality compared to the autoregressive Transformer. Jiajun Zhang 0001, Chengqing Zong, Heng Yu 0006 |
IJCAI | 3 |
| 2019 | End-to-End Speech Translation with Knowledge DistillationabstractEnd-to-end speech translation (ST), which directly translates from source language speech into target language text, has attracted intensive attentions in recent years.Compared to conventional pipepine systems, end-to-end ST models have advantages of lower latency, smaller model size and less error propagation.However, the combination of speech recognition and text translation in one model is more difficult than each of these two tasks.In this paper, we propose a knowledge distillation approach to improve ST model by transferring the knowledge from text translation model.Specifically, we first train a text translation model, regarded as a teacher model, and then ST model is trained to learn output probabilities from teacher model through knowledge distillation.Experiments on English-French Augmented LibriSpeech and English-Chinese TED corpus show that end-to-end ST is possible to implement on both similar and dissimilar language pairs.In addition, with the instruction of teacher model, end-to-end ST model can gain significant improvements by over 3.5 BLEU points. Yuchen Liu 0007, Hao Xiong 0005, Jiajun Zhang 0001, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
INTERSPEECH | 7 |
| 2019 | A unified framework and models for integrating translation memory into phrase-based statistical machine translation
Yang Liu 0085, Chengqing Zong, Keh-Yih Su |
Comput. Speech Lang. | 3 |
| 2019 | Corrigendum to 'A unified framework and models for integrating translation memory into phrase-based statistical machine translation' [Volume 54, March 2019, Pages 176-206]
Yang Liu 0085, Chengqing Zong, Keh-Yih Su |
Comput. Speech Lang. | 3 |
| 2019 | Synchronous Bidirectional Neural Machine TranslationabstractAbstract Existing approaches to neural machine translation (NMT) generate the target language sequence token-by-token from left to right. However, this kind of unidirectional decoding framework cannot make full use of the target-side future contexts which can be produced in a right-to-left decoding direction, and thus suffers from the issue of unbalanced outputs. In this paper, we introduce a synchronous bidirectional–neural machine translation (SB-NMT) that predicts its outputs using left-to-right and right-to-left decoding simultaneously and interactively, in order to leverage both of the history and future information at the same time. Specifically, we first propose a new algorithm that enables synchronous bidirectional decoding in a single model. Then, we present an interactive decoding model in which left-to-right (right-to-left) generation does not only depend on its previously generated outputs, but also relies on future contexts predicted by right-to-left (left-to-right) decoding. We extensively evaluate the proposed SB-NMT model on large-scale NIST Chinese-English, WMT14 English-German, and WMT18 Russian-English translation tasks. Experimental results demonstrate that our model achieves significant improvements over the strong Transformer model by 3.92, 1.49, and 1.04 BLEU points, respectively, and obtains the state-of-the-art per- formance on Chinese-English and English- German translation tasks.1 Jiajun Zhang 0001, Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Input Method for Human Translators: A Novel Approach to Integrate Machine Translation Effectively and ImperceptiblyabstractComputer-aided translation (CAT) systems are the most popular tool for helping human translators efficiently perform language translation. To further improve the translation efficiency, there is an increasing interest in applying machine translation (MT) technology to upgrade CAT. To thoroughly integrate MT into CAT systems, in this article, we propose a novel approach: a new input method that makes full use of the knowledge adopted by MT systems, such as translation rules, decoding hypotheses, and n-best translation lists. The proposed input method contains two parts: a phrase generation model, allowing human translators to type target sentences quickly, and an n-gram prediction model, helping users choose perfect MT fragments smoothly. In addition, to tune the underlying MT system to generate the input method preferable results, we design a new evaluation metric for the MT system. The proposed input method integrates MT effectively and imperceptibly, and it is particularly suitable for many target languages with complex characters, such as Chinese and Japanese. The extensive experiments demonstrate that our method saves more than 23% in time and over 42% in keystrokes, and it also improves the translation quality by more than 5 absolute BLEU scores compared with the strong baseline, i.e., post-editing using Google Pinyin. Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2019 | A Survey of Discourse Representations for Chinese Discourse AnnotationabstractA key element in computational discourse analysis is the design of a formal representation for the discourse structure of a text. With machine learning being the dominant method, it is important to identify a discourse representation that can be used to perform large-scale annotation. This survey provides a systematic analysis of existing discourse representation theories to evaluate whether they are suitable for annotation of Chinese text. Specifically, the two properties, expressiveness and practicality, are introduced to compare the representations of theories based on rhetorical relations and the representations of theories based on entity relations. The comparison systematically reveals linguistic and computational characteristics of the theories. After that, we conclude that none of the existing theories are quite suitable for scalable Chinese discourse annotation because they are not both expressive and practical. Therefore, a new discourse representation needs to be proposed, which should balance the expressiveness and practicality, and cover rhetorical relations and entity relations. Inspired by the conclusions, this survey discusses some preliminary proposals on how to represent the discourse structure that are worth pursuing. Xiaomian Kang, Chengqing Zong, Nianwen Xue |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Incorporating Multi-Level User Preference into Document-Level Sentiment ClassificationabstractDocument-level sentiment classification aims to predict a user’s sentiment polarity in a document about a product. Most existing methods only focus on review contents and ignore users who post reviews. In fact, when reviewing a product, different users have different word-using habits to express opinions (i.e., word-level user preference), care about different attributes of the product (i.e., aspect-level user preference), and have different characteristics to score the review (i.e., polarity-level user preference). These preferences have great influence on interpreting the sentiment of text. To address this issue, we propose a model called Hierarchical User Attention Network (HUAN), which incorporates multi-level user preference into a hierarchical neural network to perform document-level sentiment classification. Specifically, HUAN encodes different kinds of information (word, sentence, aspect, and document) in a hierarchical structure and imports user embedding and user attention mechanism to model these preferences. Empirical results on two real-world datasets show that HUAN achieves state-of-the-art performance. Furthermore, HUAN can also mine important attributes of products for different users. Haoran Li 0001, Xiaomian Kang, Haitong Yang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2019 | Experience-based Causality Learning for Intelligent AgentsabstractUnderstanding causality in text is crucial for intelligent agents. In this article, inspired by human causality learning, we propose an experience-based causality learning framework. Comparing to traditional approaches, which attempt to handle the causality problem relying on textual clues and linguistic resources, we are the first to use experience information for causality learning. Specifically, we first construct various scenarios for intelligent agents, thus, the agents can gain experience from interaction in these scenarios. Then, human participants build a number of training instances for agents of causality learning based on these scenarios. Each instance contains two sentences and a label. Each sentence describes an event that an agent experienced in a scenario, and the label indicates whether the sentence (event) pair has a causal relation. Accordingly, we propose a model that can infer the causality in text using experience by accessing the corresponding event information based on the input sentence pair. Experiment results show that our method can achieve impressive performance on the grounded causality corpus and significantly outperform the conventional approaches. Our work suggests that experience is very important for intelligent agents to understand causality. Yang Liu 0085, Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2019 | Attention With Sparsity Regularization for Neural Machine Translation and SummarizationabstractThe attention mechanism has become thede factostandard component in neural sequence to sequence tasks, such as machine translation and abstractive summarization. It dynamically determines which parts in the input sentence should be focused on when generating each word in the output sequence. Ideally, only few relevant input words should be attended to at each decoding time step and the attention weight distribution should be sparse and sharp. However, previous methods have no good mechanism to control this attention weight distribution. In this paper, we propose a sparse attention model in which a sparsity regularization term is designed to augment the objective function. We explore two kinds of regularizations:$L_{\infty }$-norm regularization and minimum entropy regularization, both of which aim to sharpen the attention weight distribution. Extensive experiments on both neural machine translation and abstractive summarization demonstrate that our proposed sparse attention model can substantially outperform the strong baselines. And the detailed analyses reveal that the final attention distribution indeed becomes sparse and sharp. Jiajun Zhang 0001, Yang Zhao 0007, Haoran Li 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and VideoabstractAutomatic text summarization is a fundamental natural language processing (NLP) application that aims to condense a source text into a shorter version. The rapid increase in multimedia data transmission over the Internet necessitates multi-modal summarization (MMS) from asynchronous collections of text, image, audio, and video. In this work, we propose an extractive MMS method that unites the techniques of NLP, speech processing, and computer vision to explore the rich information contained in multi-modal data and to improve the quality of multimedia news summarization. The key idea is to bridge the semantic gaps between multi-modal content. Audio and visual are main modalities in the video. For audio information, we design an approach to selectively use its transcription and to infer the salience of the transcription with audio signals. For visual information, we learn the joint representations of text and images using a neural network. Then, we capture the coverage of the generated summary for important visual information through text-image matching or multi-modal topic modeling. Finally, all the multi-modal aspects are considered to generate a textual summary by maximizing the salience, non-redundancy, readability, and coverage through the budgeted optimization of submodular functions. We further introduce a publicly available MMS corpus in English and Chinese.1 The experimental results obtained on our dataset demonstrate that our methods based on image matching and image topic framework outperform other competitive baseline methods. Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Investigating Inner Properties of Multimodal Representation and Semantic Compositionality With Brain-Based Componential SemanticsabstractMultimodal models have been proven to outperform text-based approaches on learning semantic representations. However, it still remains unclear what properties are encoded in multimodal representations, in what aspects do they outperform the single-modality representations, and what happened in the process of semantic compositionality in different input modalities. Considering that multimodal models are originally motivated by human concept representations, we assume that correlating multimodal representations with brain-based semantics would interpret their inner properties to answer the above questions. To that end, we propose simple interpretation methods based on brain-based componential semantics. First we investigate the inner properties of multimodal representations by correlating them with corresponding brain-based property vectors. Then we map the distributed vector space to the interpretable brain-based componential space to explore the inner properties of semantic compositionality. Ultimately, the present paper sheds light on the fundamental questions of natural language understanding, such as how to represent the meaning of words and how to combine word meanings into larger units. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 4 |
| 2018 | Learning Multimodal Word Representation via Dynamic Fusion MethodsabstractMultimodal models have been proven to outperform text-based models on learning semantic word representations. Almost all previous multimodal models typically treat the representations from different modalities equally. However, it is obvious that information from different modalities contributes differently to the meaning of words. This motivates us to build a multimodal model that can dynamically fuse the semantic representations from different modalities according to different types of words. To that end, we propose three novel dynamic fusion methods to assign importance weights to each modality, in which weights are learned under the weak supervision of word association pairs. The extensive experiments have demonstrated that the proposed methods outperform strong unimodal baselines and state-of-the-art multimodal models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 3 |
| 2018 | Source Critical Reinforcement Learning for Transferring Spoken Language Understanding to a New LanguageabstractTo deploy a spoken language understanding (SLU) model to a new language, language transferring is desired to avoid the trouble of acquiring and labeling a new big SLU corpus. An SLU corpus is a monolingual corpus with domain/intent/slot labels. Translating the original SLU corpus into the target language is an attractive strategy. However, SLU corpora consist of plenty of semantic labels (slots), which general-purpose translators cannot handle well, not to mention additional culture differences. This paper focuses on the language transferring task given a small in-domain parallel SLU corpus. The in-domain parallel corpus can be used as the first adaptation on the general translator. But more importantly, we show how to use reinforcement learning (RL) to further adapt the adapted translator, where translated sentences with more proper slot tags receive higher rewards. Our reward is derived from the source input sentence exclusively, unlike reward via actor-critical methods or computing reward with a ground truth target sentence. Hence we can adapt the translator the second time, using the big monolingual SLU corpus from the source language. We evaluate our approach on Chinese to English language transferring for SLU systems. The experimental results show that the generated English SLU corpus via adaptation and reinforcement learning gives us over 97% in the slot F1 score and over 84% accuracy in domain classification. It demonstrates the effectiveness of the proposed language transferring method. Compared with naive translation, our proposed method improves domain classification accuracy by relatively 22%, and the slot filling F1 score by relatively more than 71%. He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong |
COLING | 6 |
| 2018 | Adopting the Word-Pair-Dependency-Triplets with Individual Comparison for Natural Language InferenceabstractThis paper proposes to perform natural language inference with Word-Pair-Dependency-Triplets. Most previous DNN-based approaches either ignore syntactic dependency among words, or directly use tree-LSTM to generate sentence representation with irrelevant information. To overcome the problems mentioned above, we adopt Word-Pair-Dependency-Triplets to improve alignment and inference judgment. To be specific, instead of comparing each triplet from one passage with the merged information of another passage, we first propose to perform comparison directly between the triplets of the given passage-pair to make the judgement more interpretable. Experimental results show that the performance of our approach is better than most of the approaches that use tree structures, and is comparable to other state-of-the-art approaches. Qianlong Du, Chengqing Zong, Keh-Yih Su |
COLING | 2 |
| 2018 | Document-level Multi-aspect Sentiment Classification by Jointly Modeling Users, Aspects, and Overall RatingsabstractDocument-level multi-aspect sentiment classification aims to predict user’s sentiment polarities for different aspects of a product in a review. Existing approaches mainly focus on text information. However, the authors (i.e. users) and overall ratings of reviews are ignored, both of which are proved to be significant on interpreting the sentiments of different aspects in this paper. Therefore, we propose a model called Hierarchical User Aspect Rating Network (HUARN) to consider user preference and overall ratings jointly. Specifically, HUARN adopts a hierarchical architecture to encode word, sentence, and document level information. Then, user attention and aspect attention are introduced into building sentence and document level representation. The document representation is combined with user and overall rating information to predict aspect ratings of a review. Diverse aspects are treated differently and a multi-task framework is adopted. Empirical results on two real-world datasets show that HUARN achieves state-of-the-art performances. Haitong Yang, Chengqing Zong |
COLING | 3 |
| 2018 | Ensure the Correctness of the Summary: Incorporate Entailment Knowledge into Abstractive Sentence SummarizationabstractIn this paper, we investigate the sentence summarization task that produces a summary from a source sentence. Neural sequence-to-sequence models have gained considerable success for this task, while most existing approaches only focus on improving the informativeness of the summary, which ignore the correctness, i.e., the summary should not contain unrelated information with respect to the source sentence. We argue that correctness is an essential requirement for summarization systems. Considering a correct summary is semantically entailed by the source sentence, we incorporate entailment knowledge into abstractive summarization models. We propose an entailment-aware encoder under multi-task framework (i.e., summarization generation and entailment recognition) and an entailment-aware decoder by entailment Reward Augmented Maximum Likelihood (RAML) training. Experiment results demonstrate that our models significantly outperform baselines from the aspects of informativeness and correctness. Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong |
COLING | 4 |
| 2018 | Memory, Show the Way: Memory Based Few Shot Word Representation LearningabstractDistributional semantic models (DSMs) generally require sufficient examples for a word to learn a high quality representation.This is in stark contrast with human who can guess the meaning of a word from one or a few referents only.In this paper, we propose Mem2Vec, a memory based embedding learning method capable of acquiring high quality word representations from fairly limited context.Our method directly adapts the representations produced by a DSM with a longterm memory to guide its guess of a novel word.Based on a pre-trained embedding space, the proposed method delivers impressive performance on two challenging few-shot word similarity tasks.Embeddings learned with our method also lead to considerable improvements over strong baselines on NER and sentiment classification. Shaonan Wang, Chengqing Zong |
EMNLP | 3 |
| 2018 | Associative Multichannel Autoencoder for Multimodal Word RepresentationabstractIn this paper we address the problem of learning multimodal word representations by integrating textual, visual and auditory inputs.Inspired by the re-constructive and associative nature of human memory, we propose a novel associative multichannel autoencoder (AMA).Our model first learns the associations between textual and perceptual modalities, so as to predict the missing perceptual information of concepts.Then the textual and predicted perceptual representations are fused through reconstructing their original and associated embeddings.Using a gating mechanism our model assigns different weights to each modality according to the different concepts.Results on six benchmark concept similarity tests show that the proposed method significantly outperforms strong unimodal baselines and state-of-the-art multimodal models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 3 |
| 2018 | A Teacher-Student Framework for Maintainable Dialog ManagerabstractReinforcement learning (RL) is an attractive solution for task-oriented dialog systems.However, extending RL-based systems to handle new intents and slots requires a system redesign.The high maintenance cost makes it difficult to apply RL methods to practical systems on a large scale.To address this issue, we propose a practical teacherstudent framework to extend RL-based dialog systems without retraining from scratch.Specifically, the "student" is an extended dialog manager based on a new ontology, and the "teacher" is existing resources used for guiding the learning process of the "student".By specifying constraints held in the new dialog manager, we transfer knowledge of the "teacher" to the "student" without additional resources.Experiments show that the performance of the extended system is comparable to the system trained from scratch.More importantly, the proposed framework makes no assumption about the unsupported intents and slots, which makes it possible to improve RL-based systems incrementally.U: I'm looking for a Sichuan restaurant.S: "Spicy Little Girl" Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong, Zhifei Li 0001 |
EMNLP | 5 |
| 2018 | Three Strategies to Improve One-to-Many Multilingual TranslationabstractDue to the benefits of model compactness, multilingual translation (including many-toone, many-to-many and one-to-many) based on a universal encoder-decoder architecture attracts more and more attention.However, previous studies show that one-to-many translation based on this framework cannot perform on par with the individually trained models.In this work, we introduce three strategies to improve one-to-many multilingual translation by balancing the shared and unique features.Within the architecture of one decoder for all target languages, we first exploit the use of unique initial states for different target languages.Then, we employ language-dependent positional embeddings.Finally and especially, we propose to divide the hidden cells of the decoder into shared and language-dependent ones.The extensive experiments demonstrate that our proposed methods can obtain remarkable improvements over the strong baselines.Moreover, our strategies can achieve comparable or even better performance than the individually trained translation models. Jiajun Zhang 0001, Feifei Zhai, Jingfang Xu, Chengqing Zong |
EMNLP | 5 |
| 2018 | Addressing Troublesome Words in Neural Machine TranslationabstractOne of the weaknesses of Neural Machine Translation (NMT) is in handling lowfrequency and ambiguous words, which we refer as troublesome words.To address this problem, we propose a novel memoryenhanced NMT method.First, we investigate different strategies to define and detect the troublesome words.Then, a contextual memory is constructed to memorize which target words should be produced in what situations.Finally, we design a hybrid model to dynamically access the contextual memory so as to correctly translate the troublesome words.The extensive experiments on Chineseto-English and English-to-German translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling troublesome words. Yang Zhao 0007, Jiajun Zhang 0001, Zhongjun He, Chengqing Zong, Hua Wu 0003 |
EMNLP | 4 |
| 2018 | MSMO: Multimodal Summarization with Multimodal OutputabstractMultimodal summarization has drawn much attention due to the rapid growth of multimedia data.The output of the current multimodal summarization systems is usually represented in texts.However, we have found through experiments that multimodal output can significantly improve user satisfaction for informativeness of summaries.In this paper, we propose a novel task, multimodal summarization with multimodal output (MSMO).To handle this task, we first collect a large-scale dataset for MSMO research.We then propose a multimodal attention model to jointly generate text and select the most relevant image from the multimodal input.Finally, to evaluate multimodal outputs, we construct a novel multimodal automatic evaluation (MMAE) method which considers both intramodality salience and intermodality relevance.The experimental results show the effectiveness of MMAE. Junnan Zhu, Haoran Li 0001, Tianshang Liu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 6 |
| 2018 | Multi-modal Sentence Summarization with Modality Attention and Image FilteringabstractIn this paper, we introduce a multi-modal sentence summarization task that produces a short summary from a pair of sentence and image. This task is more challenging than sentence summarization. It not only needs to effectively incorporate visual features into standard text summarization framework, but also requires to avoid noise of image. To this end, we propose a modality-based attention mechanism to pay different attention to image patches and text units, and we design image filters to selectively use visual information to enhance the semantics of the input sentence. We construct a multimodal sentence summarization dataset and extensive experiments on this dataset demonstrate that our models significantly outperform conventional models which only employ text as input. Further analyses suggest that sentence summarization task can benefit from visually grounded representations from a variety of aspects. Haoran Li 0001, Junnan Zhu, Tianshang Liu, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 5 |
| 2018 | Phrase Table as Recommendation Memory for Neural Machine TranslationabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance recently. However, several studies indicate that NMT often generates fluent but unfaithful translations. In this paper, we propose a method to alleviate this problem by using a phrase table as recommendation memory. The main idea is to add bonus to words worthy of recommendation, so that NMT can make correct predictions. Specifically, we first derive a prefix tree to accommodate all the candidate target phrases by searching the phrase translation table according to the source sentence.Then, we construct a recommendation word set by matching between candidate target phrases and previously translated target words by NMT. After that, we determine the specific bonus value for each recommendable word by using the attention vector and phrase translation probability. Finally,we integrate this bonus value into NMT to improve the translation results. The extensive experiments demonstrate that the proposed methods obtain remarkable improvements over the strong attention based NMT. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 4 |
| 2018 | One Sentence One Model for Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
LREC | 3 |
| 2018 | Exploiting Pre-Ordering for Neural Machine Translation
Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
LREC | 3 |
| 2018 | A Comparable Study on Model Averaging, Ensembling and Reranking in NMT
Yuchen Liu 0007, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
NLPCC (2) | 6 |
| 2018 | Empirical Exploring Word-Character Relationship for Chinese Sentence RepresentationabstractThis article addresses the problem of learning compositional Chinese sentence representations, which represent the meaning of a sentence by composing the meanings of its constituent words. In contrast to English, a Chinese word is composed of characters, which contain rich semantic information. However, this information has not been fully exploited by existing methods. In this work, we introduce a novel, mixed character-word architecture to improve the Chinese sentence representations by utilizing rich semantic information of inner-word characters. We propose two novel strategies to reach this purpose. The first one is to use a mask gate on characters, learning the relation among characters in a word. The second one is to use a max-pooling operation on words to adaptively find the optimal mixture of the atomic and compositional word representations. Finally, the proposed architecture is applied to various sentence composition models, which achieves substantial performance gains over baseline models on sentence similarity task. To further verify the generalization ability of our model, we employ the learned sentence representations as features in sentence classification task, question classification task, and sentence entailment task. Results have shown that the proposed mixed character-word sentence representation models outperform both the character-based and word-based models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2017 | A Dynamic Window Neural Network for CCG SupertaggingabstractCombinatory Category Grammar (CCG) supertagging is a task to assign lexical categories to each word in a sentence. Almost all previous methods use fixed context window sizes to encode input tokens. However, it is obvious that different tags usually rely on different context window sizes. This motivates us to build a supertagger with a dynamic window approach, which can be treated as an attention mechanism on the local contexts. We find that applying dropout on the dynamic filters is superior to the regular dropout on word embeddings. We use this approach to demonstrate the state-of-the-art CCG supertagging performance on the standard test set. Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 3 |
| 2017 | Multi-modal Summarization for Asynchronous Collection of Text, Image, Audio and VideoabstractThe rapid increase in multimedia data transmission over the Internet necessitates the multi-modal summarization (MMS) from collections of text, image, audio and video.In this work, we propose an extractive multi-modal summarization method that can automatically generate a textual summary given a set of documents, images, audios and videos related to a specific topic.The key idea is to bridge the semantic gaps between multi-modal content.For audio information, we design an approach to selectively use its transcription.For visual information, we learn the joint representations of text and images using a neural network.Finally, all of the multimodal aspects are considered to generate the textual summary by maximizing the salience, non-redundancy, readability and coverage through the budgeted optimization of submodular functions.We further introduce an MMS corpus in English and Chinese, which is released to the public 1 .The experimental results obtained on this dataset demonstrate that our method outperforms other competitive baseline methods. Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 5 |
| 2017 | Exploiting Word Internal Structures for Generic Chinese Sentence RepresentationabstractWe introduce a novel mixed characterword architecture to improve Chinese sentence representations, by utilizing rich semantic information of word internal structures.Our architecture uses two key strategies.The first is a mask gate on characters, learning the relation among characters in a word.The second is a maxpooling operation on words, adaptively finding the optimal mixture of the atomic and compositional word representations.Finally, the proposed architecture is applied to various sentence composition models, which achieves substantial performance gains over baseline models on sentence similarity task. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 3 |
| 2017 | Learning Sentence Representation with Guidance of Human AttentionabstractRecently, much progress has been made in learning general-purpose sentence representations that can be used across domains. However, most of the existing models typically treat each word in a sentence equally. In contrast, extensive studies have proven that human read sentences efficiently by making a sequence of fixation and saccades. This motivates us to improve sentence representations by assigning different weights to the vectors of the component words, which can be treated as an attention mechanism on single sentences. To that end, we propose two novel attention models, in which the attention weights are derived using significant predictors of human reading time, i.e., Surprisal, POS tags and CCG supertags. The extensive experiments demonstrate that the proposed methods significantly improve upon the state-of-the-art sentence representation models. Shaonan Wang, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 3 |
| 2017 | Towards Neural Machine Translation with Partially Aligned CorporaabstractWhile neural machine translation (NMT) has become the new paradigm, the parameter optimization requires large-scale parallel data which is scarce in many domains and language pairs. In this paper, we address a new translation scenario in which there only exists monolingual corpora and phrase pairs. We propose a new method towards translation with partially aligned sentence pairs which are derived from the phrase pairs and monolingual corpora. To make full use of the partially aligned corpora, we adapt the conventional NMT training method in two aspects. On one hand, different generation strategies are designed for aligned and unaligned target words. On the other hand, a different objective function is designed to model the partially aligned parts. The experiments demonstrate that our method can achieve a relatively good result in such a translation scenario, and tiny bitexts can boost translation quality to a large extent. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhengshan Xue |
IJCNLP(1) | 4 |
| 2017 | Shortcut Sequence Tagging
Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
NLPCC | 3 |
| 2017 | Look-Ahead Attention for Generation in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
NLPCC | 3 |
| 2017 | Augmenting Neural Sentence Summarization Through Extractive Summarization
Junnan Zhu, Haoran Li 0001, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
NLPCC | 6 |
| 2017 | Implicit Discourse Relation Recognition for English and Chinese with Multiview Modeling and Effective Representation LearningabstractDiscourse relations between two text segments play an important role in many Natural Language Processing (NLP) tasks. The connectives strongly indicate the sense of discourse relations, while in fact, there are no connectives in a large proportion of discourse relations, that is, implicit discourse relations. Compared with explicit relations, implicit relations are much harder to detect and have drawn significant attention. Until now, there have been many studies focusing on English implicit discourse relations, and few studies address implicit relation recognition in Chinese even though the implicit discourse relations in Chinese are more common than those in English. In our work, both the English and Chinese languages are our focus. The key to implicit relation prediction is to properly model the semantics of the two discourse arguments, as well as the contextual interaction between them. To achieve this goal, we propose a neural network based framework that consists of two hierarchies. The first one is the model hierarchy, in which we propose a max-margin learning method to explore the implicit discourse relation from multiple views. The second one is the feature hierarchy, in which we learn multilevel distributed representations from words, arguments, and syntactic structures to sentences. We have conducted experiments on the standard benchmarks of English and Chinese, and the results show that compared with several methods our proposed method can achieve the best performance in most cases. Haoran Li 0001, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2017 | Comparison Study on Critical Components in Composition Model for Phrase RepresentationabstractPhrase representation, an important step in many NLP tasks, involves representing phrases as continuous-valued vectors. This article presents detailed comparisons concerning the effects of word vectors, training data, and the composition and objective function used in a composition model for phrase representation. Specifically, we first discuss how the augmented word representations affect the performance of the composition model. Then, we investigate whether different types of training data influence the performance of the composition model and, if so, how they influence it. Finally, we evaluate combinations of different composition and objective functions and discuss the factors related to composition model performance. All evaluations were conducted in both English and Chinese. Our main findings are as follows: (1) The Additive model with semantic enhanced word vectors performs comparably to the state-of-the-art model; (2) The Additive model which updates augmented word vectors and the Matrix model with semantic enhanced word vectors systematically outperforms the state-of-the-art model in bigram and multi-word phrase similarity task, respectively; (3) Representing the high frequency phrases by estimating their surrounding contexts is a good training objective for bigram phrase similarity tasks; and (4) The performance gain of composition model with semantic enhanced word vectors is due to the composition function and the greater weight attached to important words. Previous works focus on the composition function; however, our findings indicate that other components in the composition model (especially word representation) make a critical difference in phrase representation. Shaonan Wang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | An Empirical Exploration of Skip Connections for Sequential TaggingabstractIn this paper, we empirically explore the effects of various kinds of skip connections in stacked bidirectional LSTMs for sequential tagging. We investigate three kinds of skip connections connecting to LSTM cells: (a) skip connections to the gates, (b) skip connections to the internal states and (c) skip connections to the cell outputs. We present comprehensive experiments showing that skip connections to cell outputs outperform the remaining two. Furthermore, we observe that using gated identity functions as skip mappings works pretty well. Based on this novel skip connections, we successfully train deep stacked bidirectional LSTM models and obtain state-of-the-art results on CCG supertagging and comparable results on POS tagging. Huijia Wu, Jiajun Zhang 0001, Chengqing Zong |
COLING | 3 |
| 2016 | Exploiting Source-side Monolingual Data in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
EMNLP | 2 |
| 2016 | Towards Zero Unknown Word in Neural Machine Translation
Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 3 |
| 2016 | A Bilingual Discourse Corpus and Its Applications
Yang Liu 0085, Jiajun Zhang 0001, Chengqing Zong, Yating Yang, Xi Zhou 0007 |
LREC | 3 |
| 2016 | Learning Generalized Features for Semantic Role LabelingabstractThis article makes an effort to improve Semantic Role Labeling (SRL) through learning generalized features. The SRL task is usually treated as a supervised problem. Therefore, a huge set of features are crucial to the performance of SRL systems. But these features often lack generalization powers when predicting an unseen argument. This article proposes a simple approach to relieve the issue. A strong intuition is that arguments occurring in similar syntactic positions are likely to bear the same semantic role, and, analogously, arguments that are lexically similar are likely to represent the same semantic role. Therefore, it will be informative to SRL if syntactic or lexical similar arguments can activate the same feature. Inspired by this, we embed the information of lexicalization and syntax into a feature vector for each argument and then use K -means to make clustering for all feature vectors of training set. For an unseen argument to be predicted, it will belong to the same cluster as its similar arguments of training set. Therefore, the clusters can be thought of as a kind of generalized feature. We evaluate our method on several benchmarks. The experimental results show that our approach can significantly improve the SRL performance. Haitong Yang, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Bilingual Semantic Role Labeling Inference via Dual DecompositionabstractThis article focuses on bilingual Semantic Role Labeling (SRL); its goal is to annotate semantic roles on both sides of the parallel bilingual texts (bi-texts). Since rich bilingual information is encoded, bilingual SRL has been applied in many natural-language processing (NLP) tasks such as machine translation (MT), cross-lingual information retrieval (IR), and the like. A feasible way of performing bilingual SRL is using monolingual SRL systems to perform SRL on each side of bi-texts separately. However, it is difficult to obtain consistent SRL results on both sides of bi-texts in this way. Some works have tried to jointly infer bilingual SRL because there are many complementary language cues on both sides of bi-texts and they reported better performance than monolingual systems. However, there are two limits in the existing methods. First, the existing methods often require high inference costs due to the complex objective function. Second, the existing methods fully adopt the candidates generated by monolingual SRL systems, but many candidates are discarded in the argument pruning or identification stage of monolingual systems. In this article, we propose two strategies to overcome these limits. We utilize a simple but efficient technique: Dual Decomposition to search for consistent results for both sides of bi-texts. On the other hand, we propose a method called Bi-Directional Projection (BDP) to recover arguments discarded in monolingual SRL systems. We evaluate our method on a standard parallel benchmark: the OntoNotes dataset. The experimental results show that our method yields significant improvements over the state-of-the-art monolingual systems. In addition, our approach is also better and faster than existing methods due to BDP and Dual Decomposition. Haitong Yang, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | Abstractive Cross-Language Summarization via Translation Model Enhanced Predicate Argument Structure FusingabstractCross-language multidocument summarization is the task to generate a summary in a target language (e.g., Chinese) from a collection of documents in a different source language (e.g., English). Previous methods such as the extractive and compressive algorithms focus only on single sentence selection and compression, which cannot make full use of the similar sentences containing complementary information. Furthermore, the translation model knowledge is not fully explored in previous approaches. To address these two problems, we propose in this paper an abstractive cross-language summarization framework. First, the source language documents are translated into target language with a machine translation system. Then, the method constructs a pool of bilingual concepts and facts represented by the bilingual elements of the source-side predicate-argument structures (PAS) and their target-side counterparts. Finally, new summary sentences are produced by fusing bilingual PAS elements with the integer linear programming algorithm to maximize both of the salience and translation quality of the PAS elements. The experimental results on English-to-Chinese cross-language summarization demonstrate that our proposed method outperforms the state-of-the-art extractive systems in both automatic and manual evaluations. Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | A New Input Method for Human Translators: Integrating Machine Translation Effectively and Imperceptibly
Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IJCAI | 4 |
| 2015 | Feature Ensemble Plus Sample Selection: Domain Adaptation for Sentiment Classification (Extended Abstract)
Chengqing Zong, Xuelei Hu, Erik Cambria |
IJCAI | 2 |
| 2015 | Domain Adaptation for Syntactic and Semantic Dependency Parsing Using Deep Belief NetworksabstractIn current systems for syntactic and semantic dependency parsing, people usually define a very high-dimensional feature space to achieve good performance. But these systems often suffer severe performance drops on out-of-domain test data due to the diversity of features of different domains. This paper focuses on how to relieve this domain adaptation problem with the help of unlabeled target domain data. We propose a deep learning method to adapt both syntactic and semantic parsers. With additional unlabeled target domain data, our method can learn a latent feature representation (LFR) that is beneficial to both domains. Experiments on English data in the CoNLL 2009 shared task show that our method largely reduced the performance drop on out-of-domain test data. Moreover, we get a Macro F1 score that is 2.32 points higher than the best system in the CoNLL 2009 shared task in out-of-domain tests. Haitong Yang, Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 3 |
| 2015 | A Unified Model for Solving the OOV Problem of Chinese Word SegmentationabstractThis article proposes a unified, character-based, generative model to incorporate additional resources for solving the out-of-vocabulary (OOV) problem of Chinese word segmentation, within which different types of additional information can be utilized independently in corresponding submodels. This article mainly addresses the following three types of OOV: unseen dictionary words, named entities, and suffix-derived words, none of which are handled well by current approaches. The results show that our approach can effectively improve the performance of the first two types with positive interaction in F-score. Additionally, we also analyze reason that suffix information is not helpful. After integrating the proposed generative model with the corresponding discriminative approach, our evaluation on various corpora---including SIGHAN-2005, CIPS-SIGHAN-2010, and the Chinese Treebank (CTB)---shows that our integrated approach achieves the best performance reported in the literature on all testing sets when additional information and resources are allowed. Chengqing Zong, Keh-Yih Su |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Towards Machine Translation in Semantic Vector SpaceabstractMeasuring the quality of the translation rules and their composition is an essential issue in the conventional statistical machine translation (SMT) framework. To express the translation quality, the previous lexical and phrasal probabilities are calculated only according to the co-occurrence statistics in the bilingual corpus and may be not reliable due to the data sparseness problem. To address this issue, we propose measuring the quality of the translation rules and their composition in the semantic vector embedding space (VES). We present a recursive neural network (RNN)-based translation framework, which includes two submodels. One is the bilingually-constrained recursive auto-encoder, which is proposed to convert the lexical translation rules into compact real-valued vectors in the semantic VES. The other is a type-dependent recursive neural network, which is proposed to perform the decoding process by minimizing the semantic gap (meaning distance) between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function that maximizes the margin between the reference translation and the n-best translations in forced decoding. In the experiments, we first show that the proposed vector representations for the translation rules are very reliable for application in translation modeling. We further show that the proposed type-dependent, RNN-based model can significantly improve the translation quality in the large-scale, end-to-end Chinese-to-English translation evaluation. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2015 | Exploring Diverse Features for Statistical Machine Translation Model PruningabstractIn phrase-based and hierarchical phrase-based statistical machine translation systems, translation performance depends heavily on the size and quality of the translation table. To meet the requirements of making a real-time response, some research has been performed to filter the translation table. However, most existing methods are always based on one or two constraints that act as hard rules, such as not allowing phrase-pairs with low translation probabilities. These approaches sometimes make constraints rigid because they consider only a single factor instead of composite factors. Based on the considerations above, in this paper, we propose a machine learning-based framework that integrates multiple features for translation model pruning. Experimental results show that our framework is effective by pruning 80% of the phrase-pairs and 70% of the hierarchical rules, while retaining the quality of the translation models when using the BLEU evaluation metric. Our study further shows that our method can select the most useful phrase-pairs and rules, including those that are low in frequency but still very useful. Mei Tu, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Dual Sentiment Analysis: Considering Two Sides of One ReviewabstractBag-of-words (BOW) is now the most popular way to model text in statistical machine learning approaches in sentiment analysis. However, the performance of BOW sometimes remains limited due to some fundamental deficiencies in handling the polarity shift problem. We propose a model called dual sentiment analysis (DSA), to address this problem for sentiment classification. We first propose a novel data expansion technique by creating a sentiment-reversed review for each training and test review. On this basis, we propose a dual training algorithm to make use of original and reversed training reviews in pairs for learning a sentiment classifier, and a dual prediction algorithm to classify the test reviews by considering two sides of one review. We also extend the DSA framework from polarity (positive-negative) classification to 3-class (positive-negative-neutral) classification, by taking the neutral reviews into consideration. Finally, we develop a corpus-based method to construct a pseudo-antonym dictionary, which removes DSA's dependency on an external antonym dictionary for review reversion. We conduct a wide range of experiments including two tasks, nine datasets, two antonym dictionaries, three classification algorithms, and two types of features. The results demonstrate the effectiveness of DSA in supervised sentiment classification. Chengqing Zong, Qianmu Li, Yong Qi 0002, Tao Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Mind the Gap: Machine Translation by Minimizing the Semantic Gap in Embedding SpaceabstractThe conventional statistical machine translation (SMT) methods perform the decoding process by compositing a set of the translation rules which are associated with high probabilities. However, the probabilities of the translation rules are calculated only according to the cooccurrence statistics in the bilingual corpus rather than the semantic meaning similarity. In this paper, we propose a Recursive Neural Network (RNN) based model that converts each translation rule into a compact real-valued vector in the semantic embedding space and performs the decoding process by minimizing the semantic gap between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function. Extensive experiments on Chinese-to-English translation show that our RNN-based model can significantly improve the translation quality by up to 1.68 BLEU score. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
AAAI | 5 |
| 2014 | Enhancing Grammatical Cohesion: Generating Transitional Expressions for SMTabstractTransitional expressions provide glue that holds ideas together in a text and enhance the logical organization, which together help improve readability of a text. However, in most current statistical machine translation (SMT) systems, the outputs of compound-complex sentences still lack proper transitional expressions. As a result, the translations are often hard to read and understand. To address this issue, we propose two novel models to encourage generating such transitional expressions by introducing the source compoundcomplex sentence structure (CSS). Our models include a CSS-based translation model, which generates new CSS-based translation rules, and a generative transfer model, which encourages producing transitional expressions during decoding. The two models are integrated into a hierarchical phrase-based translation system to evaluate their effectiveness. The experimental results show that significant improvements are achieved on various test data meanwhile the translations are more cohesive and smooth. Mei Tu, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 3 |
| 2014 | Bilingually-constrained Phrase Embeddings for Machine TranslationabstractWe propose Bilingually-constrained Recursive Auto-encoders (BRAE) to learn semantic phrase embeddings (compact vector representations for phrases), which can distinguish the phrases with different semantic meanings.The BRAE is trained in a way that minimizes the semantic distance of translation equivalents and maximizes the semantic distance of nontranslation pairs simultaneously.After training, the model learns how to embed each phrase semantically in two languages and also learns how to transform semantic embedding space in one language to the other.We evaluate our proposed method on two end-to-end SMT tasks (phrase table pruning and decoding with phrasal semantic similarities) which need to measure semantic similarity between a source phrase and its translation candidates.Extensive experiments show that the BRAE is remarkably effective in these two tasks. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACL (1) | 5 |
| 2014 | Dynamically Integrating Cross-Domain Translation Memory into Phrase-Based Machine Translation during Decoding
Chengqing Zong, Keh-Yih Su |
COLING | 2 |
| 2014 | Multi-Predicate Semantic Role LabelingabstractThe current approaches to Semantic Role Labeling (SRL) usually perform role classification for each predicate separately and the interaction among individual predicate's role labeling is ignored if there is more than one predicate in a sentence.In this paper, we prove that different predicates in a sentence could help each other during SRL.In multi-predicate role labeling, there are mainly two key points: argument identification and role labeling of the arguments shared by multiple predicates.To address these issues, in the stage of argument identification, we propose novel predicate-related features which help remove many argument identification errors; in the stage of argument classification, we adopt a discriminative reranking approach to perform role classification of the shared arguments, in which a large set of global features are proposed.We conducted experiments on two standard benchmarks: Chinese PropBank and English PropBank.The experimental results show that our approach can significantly improve SRL performance, especially in Chinese Prop-Bank. Haitong Yang, Chengqing Zong |
EMNLP | 2 |
| 2014 | A Global Generative Model for Chinese Semantic Role Labeling
Haitong Yang, Chengqing Zong |
NLPCC | 2 |
| 2013 | Integrating Translation Memory into Phrase-Based Machine Translation during Decoding
Chengqing Zong, Keh-Yih Su |
ACL (1) | 2 |
| 2013 | Handling Ambiguities of Bilingual Predicate-Argument Structures for Statistical Machine Translation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 4 |
| 2013 | Learning a Phrase-based Translation Model from Monolingual Data with Application to Domain Adaptation
Jiajun Zhang 0001, Chengqing Zong |
ACL (1) | 2 |
| 2013 | Instance Selection and Instance Weighting for Cross-Domain Sentiment Classification via PU Learning
Xuelei Hu, Jianfeng Lu 0003, Jian Yang 0003, Chengqing Zong |
IJCAI | 5 |
| 2013 | An Efficient Framework to Extract Parallel Units from Comparable Data
Lu Xiang, Yu Zhou 0001, Chengqing Zong |
NLPCC | 3 |
| 2013 | A Study of the Effectiveness of Suffixes for Chinese Word Segmentation
Chengqing Zong, Keh-Yih Su |
PACLIC | 2 |
| 2013 | A Joint Model to Identify and Align Bilingual Named EntitiesabstractIn this article, an integrated model is derived that jointly identifies and aligns bilingual named entities (NEs) between Chinese and English. The model is motivated by the following observations: (1) whether an NE is translated semantically or phonetically depends greatly on its entity type, (2) entities within an aligned pair should share the same type, and (3) the initially detected NEs can act as anchors and provide further information while selecting NE candidates. Based on these observations, this article proposes a translation mode ratio feature (defined as the proportion of NE internal tokens that are semantically translated), enforces an entity type consistency constraint, and utilizes additional new NE likelihoods (based on the initially detected NE anchors). Experiments show that this novel method significantly outperforms the baseline. The type-insensitive F-score of identified NE pairs increases from 78.4% to 88.0% (12.2% relative improvement) in our Chinese–English NE alignment task, and the type-sensitive F-score increases from 68.4% to 83.0% (21.3% relative improvement). Furthermore, the proposed model demonstrates its robustness when it is tested across different domains. Finally, when semi-supervised learning is conducted to train the adopted English NE recognition model, the proposed model also significantly boosts the English NE recognition type-sensitive F-score. Chengqing Zong, Keh-Yih Su |
Comput. Linguistics | 2 |
| 2013 | A Substitution-Translation-Restoration Framework for Handling Unknown Words in Statistical Machine Translation
Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
J. Comput. Sci. Technol. | 3 |
| 2013 | Large-scale Word Alignment Using Soft Dependency Cohesion ConstraintsabstractDependency cohesion refers to the observation that phrases dominated by disjoint dependency subtrees in the source language generally do not overlap in the target language. It has been verified to be a useful constraint for word alignment. However, previous work either treats this as a hard constraint or uses it as a feature in discriminative models, which is ineffective for large-scale tasks. In this paper, we take dependency cohesion as a soft constraint, and integrate it into a generative model for large-scale word alignment experiments. We also propose an approximate EM algorithm and a Gibbs sampling algorithm to estimate model parameters in an unsupervised manner. Experiments on large-scale Chinese-English translation tasks demonstrate that our model achieves improvements in both alignment quality and translation quality. Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Unsupervised Tree Induction for Tree-based TranslationabstractIn current research, most tree-based translation models are built directly from parse trees. In this study, we go in another direction and build a translation model with an unsupervised tree structure derived from a novel non-parametric Bayesian model. In the model, we utilize synchronous tree substitution grammars (STSG) to capture the bilingual mapping between language pairs. To train the model efficiently, we develop a Gibbs sampler with three novel Gibbs operators. The sampler is capable of exploring the infinite space of tree structures by performing local changes on the tree nodes. Experimental results show that the string-to-tree translation system using our Bayesian tree structures significantly outperforms the strong baseline string-to-tree system using parse trees. Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
Trans. Assoc. Comput. Linguistics | 4 |
| 2013 | Syntax-Based Translation With Bilingually Lexicalized Synchronous Tree Substitution GrammarsabstractSyntax-based models can significantly improve the translation performance due to their grammatical modeling on one or both language side(s). However, the translation rules such as the non-lexical rule “ VP→(x0x1,VP:x1PP:x0)” in string-to-tree models do not consider any lexicalized information on the source or target side. The rule is so generalized that any subtree rooted at VP can substitute for the nonterminal VP:x1. Because rules containing nonterminals are frequently used when generating the target-side tree structures, there is a risk that rules of this type will potentially be severely misused in decoding due to a lack of lexicalization guidance. In this article, inspired by lexicalized PCFG, which is widely used in monolingual parsing, we propose to upgrade the STSG (synchronous tree substitution grammars)-based syntax translation model with bilingually lexicalized STSG. Using the string-to-tree translation model as a case study, we present generative and discriminative models to integrate lexicalized STSG into the translation model. Both small- and large-scale experiments on Chinese-to-English translation demonstrate that the proposed lexicalized STSG can provide superior rule selection in decoding and substantially improve the translation quality. Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Integrating Surface and Abstract Features for Robust Cross-Domain Chinese Word Segmentation
Chengqing Zong, Keh-Yih Su |
COLING | 3 |
| 2012 | Machine Translation by Modeling Predicate-Argument Structure Transformation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 4 |
| 2012 | Tree-based Translation without using Parse Trees
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 4 |
| 2012 | Integrating Generative and Discriminative Character-Based Models for Chinese Word SegmentationabstractAmong statistical approaches to Chinese word segmentation, the word-based n-gram ( generative ) model and the character-based tagging ( discriminative ) model are two dominant approaches in the literature. The former gives excellent performance for the in-vocabulary (IV) words; however, it handles out-of-vocabulary (OOV) words poorly. On the other hand, though the latter is more robust for OOV words, it fails to deliver satisfactory performance for IV words. These two approaches behave differently due to the unit they use (word vs. character) and the model form they adopt (generative vs. discriminative). In general, character-based approaches are more robust than word-based ones, as the vocabulary of characters is a closed set; and discriminative models are more robust than generative ones, since they can flexibly include all kinds of available information, such as future context. This article first proposes a character-based n -gram model to enhance the robustness of the generative approach. Then the proposed generative model is further integrated with the character-based discriminative model to take advantage of both approaches. Our experiments show that this integrated approach outperforms all the existing approaches reported in the literature. Afterwards, a complete and detailed error analysis is conducted. Since a significant portion of the critical errors is related to numerical/foreign strings, character-type information is then incorporated into the model to further improve its performance. Last, the proposed integrated approach is tested on cross-domain corpora, and a semi-supervised domain adaptation algorithm is proposed and shown to be effective in our experiments. Chengqing Zong, Keh-Yih Su |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2011 | Augmenting String-to-Tree Translation Models with Fuzzy Use of Source-side Syntax
Jiajun Zhang 0001, Feifei Zhai, Chengqing Zong |
EMNLP | 3 |
| 2011 | A Semantic-Specific Model for Chinese Named Entity Translation
Chengqing Zong |
IJCNLP | 2 |
| 2011 | Parse Reranking Based on Higher-Order Lexical Dependencies
Chengqing Zong |
IJCNLP | 2 |
| 2011 | A POS-based Ensemble Model for Cross-domain Sentiment Classification
Chengqing Zong |
IJCNLP | 2 |
| 2011 | Simple but Effective Approaches to Improving Tree-to-tree Model
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
MTSummit | 4 |
| 2011 | Ensemble of feature sets and classification algorithms for sentiment classification
Chengqing Zong, Shoushan Li |
Inf. Sci. | 2 |
| 2011 | Multi-Domain Sentiment Classification with Classifier Combination
Shoushan Li, Chu-Ren Huang, Chengqing Zong |
J. Comput. Sci. Technol. | 3 |
| 2011 | Preface
Chengqing Zong, Hans Uszkoreit |
J. Comput. Sci. Technol. | 1 |
| 2010 | On Jointly Recognizing and Aligning Bilingual Named Entities
Chengqing Zong, Keh-Yih Su |
ACL | 2 |
| 2010 | A Novel Reordering Model Based on Multi-layer Phrase for Statistical Machine Translation
Yanqing He, Yu Zhou 0001, Chengqing Zong, Huilin Wang |
COLING | 3 |
| 2010 | A Character-Based Joint Model for Chinese Word Segmentation
Chengqing Zong, Keh-Yih Su |
COLING | 2 |
| 2010 | A Minimum Error Weighting Combination Strategy for Chinese Semantic Role Labeling
Chengqing Zong |
COLING | 2 |
| 2010 | Joint Inference for Bilingual Semantic Role Labeling
Chengqing Zong |
EMNLP | 2 |
| 2010 | CASIA-CASSIL: a Chinese Telephone Conversation Corpus in Real Scenarios with Multi-leveled Annotation
Keyan Zhou, Zhigang Yin, Chengqing Zong |
LREC | 4 |
| 2009 | A Framework of Feature Selection Methods for Text Categorization
Shoushan Li, Chengqing Zong, Chu-Ren Huang |
ACL/IJCNLP | 3 |
| 2009 | Layer-Based Dependency Parsing
Ping Jian, Chengqing Zong |
PACLIC | 2 |
| 2009 | Approach to Selecting Best Development Set for Phrase-Based Statistical Machine Translation
Yu Zhou 0001, Chengqing Zong |
PACLIC | 3 |
| 2009 | Which is More Suitable for Chinese Word Segmentation, the Generative Model or the Discriminative One?
Chengqing Zong, Keh-Yih Su |
PACLIC | 2 |
| 2009 | A Framework for Effectively Integrating Hard and Soft Syntactic Rules into Phrase Based Translation
Jiajun Zhang 0001, Chengqing Zong |
PACLIC | 2 |
| 2008 | Domain Adaptation for Statistical Machine Translation with Domain Dictionary and Monolingual Corpora
Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
COLING | 3 |
| 2008 | Sentence Type Based Reordering Model for Statistical Machine Translation
Jiajun Zhang 0001, Chengqing Zong, Shoushan Li |
COLING | 2 |
| 2008 | A New Approach to Automatic Document Summarization
Chengqing Zong |
IJCNLP | 2 |
| 2008 | A Structure-Based Model for Chinese Organization Name TranslationabstractNamed entity (NE) translation is a fundamental task in multilingual natural language processing. The performance of a machine translation system depends heavily on precise translation of the inclusive NEs. Furthermore, organization name (ON) is the most complex NE for translation among all the NEs. In this article, the structure formulation of ONs is investigated and a hierarchical structure-based ON translation model for Chinese-to-English translation system is presented. First, the model performs ON chunking; then both the translation of words within chunks and the process of chunk-reordering are achieved by synchronous context-free grammar (CFG). The CFG rules are extracted from bilingual ON pairs in a training program. The main contributions of this article are: (1) defining appropriate chunk-units for analyzing the internal structure of Chinese ONs; (2) making the chunk-based ON translation feasible and flexible via a hierarchical CFG derivation; and (3) proposing a training architecture to automatically learn the synchronous CFG for constructing ONs with chunk-units from aligned bilingual ON pairs. The experiments show that the proposed approach translates the Chinese ONs into English with an accuracy of 93.75% and significantly improves the performance of a baseline statistical machine translation (SMT) system. Chengqing Zong |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2006 | Concept Features Extraction and Text Clustering Analysis of Neural Networks Based on Cognitive Mechanism
Lin Wang 0006, Minghu Jiang, Shasha Liao, Beixing Deng, Chengqing Zong, Yinghua Lu |
ICIC (1) | 5 |
| 2006 | An approach to automatic acquisition of translation templates based on phrase structure extraction and alignmentabstractIn this paper, we propose a new approach for automatically acquiring translation templates from unannotated bilingual spoken language corpora. Two basic algorithms are adopted: a grammar induction algorithm, and an alignment algorithm using bracketing transduction grammar. The approach is unsupervised, statistical, and data-driven, and employs no parsing procedure. The acquisition procedure consists of two steps. First, semantic groups and phrase structure groups are extracted from both the source language and the target language. Second, an alignment algorithm based on bracketing transduction grammar aligns the phrase structure groups. The aligned phrase structure groups are post-processed, yielding translation templates. Preliminary experimental results show that the algorithm is effective. Rile Hu, Chengqing Zong, Bo Xu 0002 |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Investigation of Emotive Expressions of Spoken Sentences
Wenjie Cao, Chengqing Zong, Bo Xu 0002 |
ACII | 2 |
| 2005 | Self-organizing Map Analysis of Conceptual and Semantic Relations for Noun
Minghu Jiang, Chengqing Zong, Beixing Deng |
ISNN (3) | 2 |
| 2005 | Toward Practical Spoken Language Translation
Chengqing Zong, Mark Seligman |
Mach. Transl. | 1 |
| 2004 | Approach to interchange-format based Chinese generationabstractInterlingua-based machine translation is an important approach to implement multi-lingual speech-to-speech (S2S) translation. The natural language generation (NLG) is one of the key components in the interlingua-based machine translation systems. This paper introduces our approach to Chinese generation based on the Interchange Format (IF) developed by the C-STAR organization. In our approach, the hybrid method of feature-based deep generation method and template-based method are employed. The deep generator ensures that the generation component possesses the merits of flexibility and domain portability. The template-based generator makes the system more efficient. We also introduce another simplified Chinese generator applied in specific domain. The experimented results show that our approach is effective and practical for the natural language generation in the Interchange-Format (IF) based S2S translation system. 1. Wenjie Cao, Chengqing Zong, Bo Xu 0002 |
INTERSPEECH | 2 |
| 2004 | Worldwide ongoing activities on multilingual speech to speech translationabstractThis paper presents an overview of worldwide going on activities on Speech-to-Speech Translation. After a short introduction of the field, including the major projects and milestones, activities and projects going on in Asia, Europe and US are presented and described. Gianni Lazzari, Alex Waibel, Chengqing Zong |
INTERSPEECH | 3 |
| 2004 | Collecting and Sharing Bilingual Spontaneous Speech Corpora: the ChinFaDial Experiment
Georges Fafiotte, Christian Boitet, Mark Seligman, Chengqing Zong |
LREC | 4 |
| 2003 | An Estimate Method of the Minimum Entropy of Natural Languages
Fuji Ren, Shunji Mitsuyoshi, Kang Yen, Chengqing Zong, Hongbing Zhu |
CICLing | 4 |
| 2003 | A Maximum Entropy Approach for Spoken Chinese Understanding
Guodong Xie, Chengqing Zong, Bo Xu 0002 |
CICLing | 2 |
| 2003 | Chinese Utterance Segmentation in Spoken Language Translation
Chengqing Zong, Fuji Ren |
CICLing | 1 |
| 2003 | Statistical speech-to-speech translation with multilingual speech recognition and bilingual-chunk parsing
Bo Xu 0002, Shuwu Zhang, Chengqing Zong |
INTERSPEECH | 3 |
| 2003 | Rule base combined linguistics knowledge with corpusabstractThis paper proposes a new approach to construction of rule bases for the transferred-based machine translation. In our approach, the rule bases are constructed in combination of the linguistics knowledge and large scale of corpora. On the one hand, the lexical knowledge, the syntactic knowledge and the semantic knowledge are all used in the rules. On the other hand, the knowledge is used for the statistics and self-learning rules. In each rule base, all rules are scored and ranked. Thus, an impersonal choice for the sentence can be made. The preliminary experimental results show that the approach may increase the speed to build the rule base and improve the quality of rules. Chengqing Zong |
SMC | 2 |
| 2003 | Automatic evaluation of sentence fluencyabstractIn the machine translation (MT) system, how to evaluate the sentence fluency of the translation results is an important research topic. Most of the current methods are based on the similarity of output words compared with the reference translations, which don't specially address the evaluation of the sentence fluency according to the syntactic structure. This paper proposes a statistical approach to the problem, which is based on the n-gram language model and reference-independent. Our approach is a beneficial compensation to the reference-dependent methods and has better robustness than those methods based on the syntactic analysis. The preliminary experimental results indicate that the approach basically reflects the reality of human's judgment. Yu Zhou 0001, Chengqing Zong, Fuji Ren |
SMC | 3 |
| 2002 | Chinese Syntactic Parsing Based on Extended GLR Parsing Algorithm with PCFG*
Bo Xu 0002, Chengqing Zong |
COLING | 3 |
| 2002 | Chinese spoken language analyzing based on combination of statistical and rule methodsabstractA combination of statistical and rule methods has been developed for Chinese spoken language analyzing. The analyzing result is a middle semantic frame, which can be converted to different language according people’s needs. We adopt the statistical method in the stage of extracting semantic meaning and the rule method in the stage of mapping the semantic units to middle semantic frame. Experiment shows this method has high robustness and can analyzing Chinese spoken language effectively. 1. Guodong Xie, Chengqing Zong, Bo Xu 0002 |
INTERSPEECH | 2 |
| 2000 | Chinese Generation in a Spoken Dialogue Translation System
Taiyi Huang 0001, Chengqing Zong |
COLING | 3 |
| 2000 | Approach to Recognition and Understanding of the Time Constituents in the Spoken Chinese Language Translation
Chengqing Zong, Taiyi Huang 0001, Bo Xu 0002 |
ICMI | 1 |
| 2000 | An improved template-based approach to spoken language translationabstractIn this paper, we describe an improved template-based approach to Chinese-to-English Spoken Language Translation (SLT) and present experimental results. The improved template-based translation approach uses flexible expression format to describe the template condition. The condition of a template may consist of keywords, parts-of-speech and also semantic features, so the input may be matched with a template from shallow level to deep level. In the condition of a template, the distance between two fixed keywords is stretchable, thus some needless words in the input utterances may be skipped in matching operation. And also the translation results of the same template are alterable. The proper results are finally generated according to the specific context. That is, the relation between a template and translated utterance is one-to-n (where, n is an integer and n≥1). The experiments were performed with input of both text transcription and results of speech recognition. The preliminary experimental results have proven the approach is practical. 1. Chengqing Zong, Taiyi Huang 0001, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2000 | Japanese-to-Chinese spoken language translation based on the simple expressionabstractThis paper describes a Japanese-to-Chinese spoken language translation (SLT) method based on simple expression and presents the experimental results. The method is aimed at developing a compact speech translation system, which is robust for spontaneous spoken language phenomena, including the recognition errors and different expression from various speakers. The idea of translation method based on simple expression is that the mechanism interprets speech-act rather than the direct translation of the speaker’s words. The method is realized by mapping the simple expression instead of deep Chengqing Zong, Yumi Wakita, Bo Xu 0002, Zhenbiao Chen, Kenji Matsui |
INTERSPEECH | 1 |
| 1997 | Parsing with dynamic rule selection
Chengqing Zong, Zhaoxiong Chen, Heyan Huang |
J. Comput. Sci. Technol. | 1 |