Yu Zhou 0001

dblp:36/2728-1 · DBLP profile ↗
← Back
60ranked-venue papers
0as first author
33since 2021 · last 2025
0000-0002-4911-4717ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
abstract
Yupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yupu Liang, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001
ACL (1)7
2025 From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image Translation
abstract
Document Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in complex-layout scenarios, frequently loses the crucial document coherence, leading to chaotic text. To overcome this problem, we introduce a novel end-to-end network, named Zoom-out DIT (ZoomDIT), inspired by human translation procedures. It jointly accomplishes the multi-level tasks including word positioning, sentence recognition & translation, and document organization, based on a fine-to-coarse zoom-out framework, to progressively realize “chaotic words to coherent document” and improve translation. We further contribute a new large-scale DIT dataset with multi-level fine-grained labels. Extensive experiments on public and our new dataset demonstrate significant improvements in translation quality towards complex-layout document images, offering a robust solution for reorganizing the chaotic OCR outputs to a coherent document translation.
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
COLING6
2025 ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ICDAR (5)7
2025 SimulPL: Aligning Human Preferences in Simultaneous Machine Translation
abstract
Simultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios have distinct preferences, such as accurate translations, simpler syntax, and no unnecessary latency. Aligning SiMT models with these human preferences is crucial to improve their performances. However, this issue still remains unexplored. Additionally, preference optimization for SiMT task is also challenging. Existing methods focus solely on optimizing the generated responses, ignoring human preferences related to latency and the optimization of read/write policy during the preference optimization phase. To address these challenges, we propose Simultaneous Preference Learning (SimulPL), a preference learning framework tailored for the SiMT task. In the SimulPL framework, we categorize SiMT human preferences into five aspects: **translation quality preference**, **monotonicity preference**, **key point preference**, **simplicity preference**, and **latency preference**. By leveraging the first four preferences, we construct human preference prompts to efficiently guide GPT-4/4o in generating preference data for the SiMT task. In the preference optimization phase, SimulPL integrates **latency preference** into the optimization objective and enables SiMT models to improve the read/write policy, thereby aligning with human preferences more effectively. Experimental results indicate that SimulPL exhibits better alignment with human preferences across all latency levels in Zh$\rightarrow$En, De$\rightarrow$En and En$\rightarrow$Zh SiMT tasks. Our data and code will be available at https://github.com/EurekaForNLP/SimulPL.
Donglei Yu, Yang Zhao 0007, Yangyifan Xu, Yu Zhou 0001, Chengqing Zong
ICLR5
2025 Pay More Attention to Images: Numerous Images-Oriented Multimodal Summarization
abstract
Min Xiao, Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou 0001
NAACL (Long Papers)5
2025 Investigating Hallucinations in Simultaneous Machine Translation: Knowledge Distillation Solution and Components Analysis
abstract
Donglei Yu, Xiaomian Kang, Yuchen Liu, Feifei Zhai, Nanchang Cheng, Yu Zhou, Chengqing Zong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Donglei Yu, Xiaomian Kang, Yuchen Liu 0007, Feifei Zhai, Nanchang Cheng, Yu Zhou 0001, Chengqing Zong
NAACL (Long Papers)6
2025 Boosting Document Image Translation via Layout-Aware Semantic Paragraph Clustering
Yupu Liang, Yunfei Lu, Dandan Tu, Chengqing Zong, Yu Zhou 0001
PRCV (7)9
2025 Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image Translation
abstract
Document Image Translation (DIT) aims to translate texts on document images from one language to another. It is a multi-modal task involving cooperation of text and layout. Current approaches either handle layout and translation as separate processes, risking accumulative errors, or use vanilla end-to-end encoder-decoder models to capture layout implicitly, often suffering inadequate layout incorporation. We argue that a favorable framework should explicitly engage layout-specific modules and properly organize them toward translation. For this, we first revisit two key layouts: the geometric layout reflecting word's spatial positions, and the logical layout depicting word's logical order. Then, a novel pipeline (understand layout $\rightarrow$→ translate text) is determined to prioritize layouts such that preceding layouts contribute to translation. Following this pipeline, we introduce Unified Document Image Translation (UniDIT), a comprehensive framework that unifies layout with translation in one network. It is devised to leverage each module's advantage, and provide an elaborate feature-conductive flow for module communication globally. A novel bridging mechanism is also introduced to adapt layout features conducive to translation. We further contribute DITransv2, a large-scale fine-grained benchmark that includes heterogeneous and complex document layouts. Extensive experiments on DITransv2 and additional established benchmarks demonstrate UniDIT outperforms previous state-of-the-arts in all aspects.
Yupu Liang, Cong Ma 0002, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 DIUSum: Dynamic Image Utilization for Multimodal Summarization
abstract
Existing multimodal summarization approaches focus on fusing image features in the encoding process, ignoring the individualized needs for images when generating different summaries. However, whether intuitively or empirically, not all images can improve summary quality. Therefore, we propose a novel Dynamic Image Utilization framework for multimodal Summarization (DIUSum) to select and utilize valuable images for summarization. First, to predict whether an image helps produce a high-quality summary, we propose an image selector to score the usefulness of each image. Second, to dynamically utilize the multimodal information, we incorporate the hard and soft guidance from the image selector. Under the guidance, the image information is plugged into the decoder to generate a summary. Experimental results have shown that DIUSum outperforms multiple strong baselines and achieves SOTA on two public multimodal summarization datasets. Further analysis demonstrates that the image selector can reflect the improved level of summary quality brought by the images.
Junnan Zhu, Feifei Zhai, Yu Zhou 0001, Chengqing Zong
AAAI4
2024 Self-Modifying State Modeling for Simultaneous Machine Translation
abstract
Simultaneous Machine Translation (SiMT) generates target outputs while receiving stream source inputs and requires a read/write policy to decide whether to wait for the next source token or generate a new target token, whose decisions form a decision path.Existing SiMT methods, which learn the policy by exploring various decision paths in training, face inherent limitations.These methods not only fail to precisely optimize the policy due to the inability to accurately assess the individual impact of each decision on SiMT performance, but also cannot sufficiently explore all potential paths because of their vast number.Besides, building decision paths requires unidirectional encoders to simulate streaming source inputs, which impairs the translation quality of SiMT models.To solve these issues, we propose Self-Modifying State Modeling (SM 2 ), a novel training paradigm for SiMT task.Without building decision paths, SM 2 individually optimizes decisions at each state during training.To precisely optimize the policy, SM 2 introduces Self-Modifying process to independently assess and adjust decisions at each state.For sufficient exploration, SM 2 proposes Prefix Sampling to efficiently traverse all potential states.Moreover, SM 2 ensures compatibility with bidirectional encoders, thus achieving higher translation quality.Experiments show that SM 2 outperforms strong baselines.Furthermore, SM 2 allows offline machine translation models to acquire SiMT ability with fine-tuning 1 .
Donglei Yu, Xiaomian Kang, Yuchen Liu 0007, Yu Zhou 0001, Chengqing Zong
ACL (1)4
2024 Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine Translation
abstract
Text image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained knowledge supervision to make it consistent between recognition and translation modules. In this paper, we propose a novel TIMT method named as BabyNet, which is optimized with hierarchical parental supervision to improve translation performance. Inspired by genetic recombination and variation in the field of genetics, the proposed BabyNet is inherited from the recognition and translation parent models with a variation module of which parameters can be updated when training on the TIMT task. Meanwhile, hierarchical and multi-granularity supervision from parent models is introduced to bridge the gap between inherited modules in BabyNet. Extensive experiments on both synthetic and real-world TIMT tests show that our proposed method significantly outperforms existing methods. Further analyses of various parent model combinations show the good generalization of our method.
Cong Ma 0002, Yupu Liang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
LREC/COLING6
2024 A Hybrid Approach towards Chinese Spelling and Splitting Error Correction
abstract
Existing Chinese spelling check (CSC) methods have limitations in correcting variable-length error characters, requiring the input and output to be the same length. They mainly focus on modelling Chinese characters’ phonetic information and generating candidates for each position. In contrast, few approaches delve into the intricacies of splitting Chinese characters to address glyph errors and splitting variable-length corrections. We define the Chinese Splitting Error Correction (CSEC) task and develop CSEC datasets in news and social media domains to address this issue. We then propose Soft-Masked Multi-feature Error Correction (SoMu) model, which first generates semantic, phonetic, graphic, and unique Chinese Wubi embeddings, then integrates those features through selective gating fusion, followed by a soft-mask strategy to filter incorrect tokens and finally use transformer layers to predict the correct ones. This model effectively addresses both spelling and splitting errors. Extensive analysis shows that our model significantly improves character-splitting information modelling for CSEC. Our dataset is available at https://github.com/Skywalker-Harrison/SoMu.
Junhong Liang, Junnan Zhu, Feifei Zhai, Nanchang Cheng, Chengqing Zong, Yu Zhou 0001
ECAI6
2024 Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation
abstract
End-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning continuous features from related tasks of text image recognition (TIR) or machine translation (MT) to alleviate the problem of data limitation, but the alignment in continuous vector space is extremely difficult and it inevitably introduces fitting errors resulting in significant performance degradation. To better align TIMT features with MT semantic features, we propose a novel Vector Quantization Knowledge Transfer (VQKT) method that employs a trainable codebook to quantize continuous features into discrete space. The quantization distribution of the MT feature is utilized as the teacher distribution to guide the TIMT model to generate similar discrete codes. Through alignment and knowledge transfer based on probability distribution, the TIMT model can better imitate the feature representation of the MT teacher model and generate high-quality target language translation. Extensive experiments demonstrate VQKT significantly outperforms the existing end-to-end TIMT performance.
Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ICASSP4
2024 Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling
abstract
Yupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yupu Liang, Cong Ma 0002, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001
NAACL-HLT8
2024 🚀 TableRocket: An Efficient and Effective Framework for Table Reconstruction
Liucheng Pang, Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
PRCV (7)5
2024 Multi-Modal Attention Based on 2D Structured Sequence for Table Recognition
Lu Xiang, Yu Zhou 0001
PRCV (7)4
2024 Knowledge Graph Guided Neural Machine Translation with Dynamic Reinforce-selected Triples
abstract
Previous methods incorporating knowledge graphs (KGs) into neural machine translation (NMT) adopt a static knowledge utilization strategy, that introduces many useless knowledge triples and makes the useful triples difficult to be utilized by NMT. To address this problem, we propose a KG guided NMT model with dynamic reinforce-selected triples. The proposed methods could dynamically select the different useful knowledge triples for different source sentences. Specifically, the proposed model contains two components: (1) knowledge selector, that dynamically selects useful knowledge triples for a source sentence, and (2) knowledge guided NMT (KgNMT), that utilizes the selected triples to guide the translation of NMT. Meanwhile, to overcome the non-differentiable problem and guide the training procedure, we propose a policy gradient strategy to encourage the model to select useful triples and improve the generation probability of gold target sentence. Various experimental results show that the proposed method can significantly outperform the baseline models in both translation quality and handling the entities.
Yang Zhao 0007, Xiaomian Kang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2024 Modal Contrastive Learning Based End-to-End Text Image Machine Translation
abstract
Text image machine translation (TIMT) aims at directly translating text in the source language embedded in images into the target language. Most existing systems follow the cascaded pipeline diagram from recognition to translation, which suffers from the problem of error propagation, parameter redundancy, and information reduction. The end-to-end model has the potential to alleviate these issues via bridging the recognition and translation models. However, the challenge is the data limitation and modality gap between text and image. In this paper, we propose a novel end-to-end model, namely Modal contrastive learning based End-to-end Text Image Machine Translation (METIMT), which alleviates these issues through end-to-end text image machine translation architecture and modal contrastive learning. Specifically, an image encoder is designed to encode images into the same feature space of corresponding text sentences, with the guidance of an intramodal and inter-modal contrastive learning module. To further promote the research of text image machine translation, we have constructed one synthetic and two real-world datasets. Extensive experiments show that our lighter, faster model outperforms not only existing pipeline methods but also state-of-the-art end-to-end models on both synthetic and real-world evaluation sets. Our code and dataset will be released to the public.
Cong Ma 0002, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Multilingual Knowledge Graph Completion with Language-Sensitive Multi-Graph Attention
abstract
Multilingual Knowledge Graph Completion (KGC) aims to predict missing links with multilingual knowledge graphs.However, existing approaches suffer from two main drawbacks: (a) alignment dependency: the multilingual KGC is always realized with joint entity or relation alignment, which introduces additional alignment models and increases the complexity of the whole framework; (b) training inefficiency: the trained model will only be used for the completion of one target KG, although the data from all KGs are used simultaneously.To address these drawbacks, we propose a novel multilingual KGC framework with languagesensitive multi-graph attention such that the missing links on all given KGs can be inferred by a universal knowledge completion model.Specifically, we first build a relational graph neural network by sharing the embeddings of aligned nodes to transfer language-independent knowledge.Meanwhile, a language-sensitive multi-graph attention (LSMGA) is proposed to deal with the information inconsistency among different KGs.Experimental results show that our model achieves significant improvements on the DBP-5L and E-PKG datasets.1
Rongchuan Tang, Yang Zhao 0007, Chengqing Zong, Yu Zhou 0001
ACL (1)4
2023 CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization
abstract
Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear.Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are useful.Therefore, we propose a novel Coarse-to-Fine contribution network for multimodal Summarization (CFSum) to consider different contributions of images for summarization.First, to eliminate the interference of useless images, we propose a pre-filter module to abandon useless images.Second, to make accurate use of useful images, we propose two levels of visual complement modules, word level and phrase level.Specifically, image contributions are calculated and are adopted to guide the attention of both textual and visual modalities.Experimental results have shown that CFSum significantly outperforms multiple strong baselines on the standard benchmark.Furthermore, the analysis verifies that useful images can even help generate nonvisual words which are implicitly represented in the image 1 .
Junnan Zhu, Haitao Lin 0001, Yu Zhou 0001, Chengqing Zong
ACL (1)4
2023 Multi-teacher Knowledge Distillation for End-to-End Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ICDAR (1)5
2023 E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ICDAR (6)5
2023 Zero-shot language extension for dialogue state tracking via pre-trained models and multi-auxiliary-tasks fine-tuning
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong
Knowl. Based Syst.4
2023 Topic-Oriented Dialogue Summarization
abstract
A multi-turn dialogue often contains multiple discussion topics. In several scenarios (e.g., customer service dispute, public opinion monitoring), people are only interested in the gist of a specific topic in the dialogue. Therefore, we propose a novel summarization task, i.e., Topic-Oriented Dialogue Summarization (TODS). Given a dialogue with a topic label, TODS aims to produce a summary covering the main content of the given topic in the dialogue. To model the relationship between dialogues and topics, three key abilities are needed for TODS: (1) Learning the semantic information of different topics. (2) Locating the topic-related content in the dialogue. (3) Distinguishing summaries for different topics in the same dialogue. Thus, we propose three topic-related auxiliary tasks to make the summarization model learn the three abilities above. First, the topic identification task aims at generating all the topics in the dialogue. Second, the topic attention restriction task tries to constrain the attention distribution on topic-related utterances. Third, the topic summary distinguishing task focuses on increasing the difference of summaries for different topics in the same dialogue. Experimental results on two public TODS datasets show that all auxiliary tasks are critical for TODS and help generate high-quality summaries. We also point out the expansions and challenges in TODS for future research.
Haitao Lin 0001, Junnan Zhu, Lu Xiang, Feifei Zhai, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role Interactions
abstract
Role-oriented dialogue summarization is to generate summaries for different roles in the dialogue, e.g., merchants and consumers.Existing methods handle this task by summarizing each role's content separately and thus are prone to ignore the information from other roles.However, we believe that other roles' content could benefit the quality of summaries, such as the omitted information mentioned by other roles.Therefore, we propose a novel role interaction enhanced method for role-oriented dialogue summarization.It adopts cross attention and decoder self-attention interactions to interactively acquire other roles' critical information.The cross attention interaction aims to select other roles' critical dialogue utterances, while the decoder self-attention interaction aims to obtain key information from other roles' summaries.Experimental results have shown that our proposed method significantly outperforms strong baselines on two public role-oriented dialogue summarization datasets.Extensive analyses have demonstrated that other roles' content could help generate summaries with more complete semantics and correct topic structures. 1
Haitao Lin 0001, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACL (1)4
2022 Improving End-to-End Text Image Translation From the Auxiliary Text Translation Task
abstract
End-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent research. However, data sparsity limits the performance of end-to-end text image translation. Multi-task learning is a nontrivial way to alleviate this problem via exploring knowledge from complementary related tasks. In this paper, we propose a novel text translation enhanced text image translation, which trains the end-to-end model with text translation as an auxiliary task. By sharing model parameters and multi-task training, our model is able to take full advantage of easily-available large-scale text parallel corpus. Extensive experimental results show our proposed method outperforms existing end-to-end methods, and the joint multi-task learning with both text translation and recognition tasks achieves better results, proving translation and recognition auxiliary tasks are complementary.1
Cong Ma 0002, Mei Tu, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001
ICPR7
2022 One-Shot Relation Learning for Knowledge Graphs via Neighborhood Aggregation and Paths Encoding
abstract
The relation learning between two entities is an essential task in knowledge graph (KG) completion that has received much attention recently. Previous work almost exclusively focused on relations widely seen in the original KGs, which means that enough training data are available for modeling. However, long-tail relations that only show in a few triples are actually much more common in practical KGs. Without sufficiently large training data, the performance of existing models on predicting long-tail relations drops impressively. This work aims to predict the relation under a challenging setting where only one instance is available for training. We propose a path-based one-shot relation prediction framework, which can extract neighborhood information of an entity based on the relation query attention mechanism to learn transferable knowledge among the same relation. Simultaneously, to reduce the impact of long-tail entities on relation prediction, we selectively fuse path information between entity pairs as auxiliary information of relation features. Experiments in three one-shot relation learning datasets show that our proposed framework substantially outperforms existing models on one-shot link prediction and relation prediction.
Jian Sun 0035, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue Summarization
abstract
Dialogue summarization has drawn much attention recently.Especially in the customer service domain, agents could use dialogue summaries to help boost their works by quickly knowing customer's issues and service progress.These applications require summaries to contain the perspective of a single speaker and have a clear topic flow structure, while neither are available in existing datasets.Therefore, in this paper, we introduce a novel Chinese dataset for Customer Service Dialogue Summarization (CSDS).CSDS improves the abstractive summaries in two aspects: (1) In addition to the overall summary for the whole dialogue, role-oriented summaries are also provided to acquire different speakers' viewpoints.(2) All the summaries sum up each topic separately, thus containing the topic-level structure of the dialogue.We define tasks in CSDS as generating the overall summary and different role-oriented summaries for a given dialogue.Next, we compare various summarization methods on CSDS, and experiment results show that existing methods are prone to generate redundant and incoherent summaries.Besides, the performance becomes much worse when analyzing the performance on role-oriented summaries and topic structures.We hope that this study could benchmark Chinese dialogue summarization and benefit further studies.
Haitao Lin 0001, Liqun Ma, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
EMNLP (1)5
2021 Augmenting Slot Values and Contexts for Spoken Language Understanding with Pretrained Models
abstract
Spoken Language Understanding (SLU) is one essential step in building a dialogue system. Due to the expensive cost of obtaining the labeled data, SLU suffers from the data scarcity problem. Therefore, in this paper, we focus on data augmentation for slot filling task in SLU. To achieve that, we aim at generating more diverse data based on existing data. Specifically, we try to exploit the latent language knowledge from pretrained language models by finetuning them. We propose two strategies for finetuning process: value-based and context-based augmentation. Experimental results on two public SLU datasets have shown that compared with existing data augmentation methods, our proposed method can generate more diverse sentences and significantly improve the performance on SLU.
Haitao Lin 0001, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
Interspeech3
2021 Zero-Shot Deployment for Cross-Lingual Dialogue System
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong
NLPCC (2)4
2021 Robust Cross-lingual Task-oriented Dialogue
abstract
Cross-lingual dialogue systems are increasingly important in e-commerce and customer service due to the rapid progress of globalization. In real-world system deployment, machine translation (MT) services are often used before and after the dialogue system to bridge different languages. However, noises and errors introduced in the MT process will result in the dialogue system's low robustness, making the system's performance far from satisfactory. In this article, we propose a novel MT-oriented noise enhanced framework that exploits multi-granularity MT noises and injects such noises into the dialogue system to improve the dialogue system's robustness. Specifically, we first design a method to automatically construct multi-granularity MT-oriented noises and multi-granularity adversarial examples, which contain abundant noise knowledge oriented to MT. Then, we propose two strategies to incorporate the noise knowledge: (i) Utterance-level adversarial learning and (ii) Knowledge-level guided method. The former adopts adversarial learning to learn a perturbation-invariant encoder, guiding the dialogue system to learn noise-independent hidden representations. The latter explicitly incorporates the multi-granularity noises, which contain the noise tokens and their possible correct forms, into the training and inference process, thus improving the dialogue system's robustness. Experimental results on three dialogue models, two dialogue datasets, and two language pairs have shown that the proposed framework significantly improves the performance of the cross-lingual dialogue system.
Lu Xiang, Junnan Zhu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2021 Graph-based Multimodal Ranking Models for Multimodal Summarization
abstract
Multimodal summarization aims to extract the most important information from the multimedia input. It is becoming increasingly popular due to the rapid growth of multimedia data in recent years. There are various researches focusing on different multimodal summarization tasks. However, the existing methods can only generate single-modal output or multimodal output. In addition, most of them need a lot of annotated samples for training, which makes it difficult to be generalized to other tasks or domains. Motivated by this, we propose a unified framework for multimodal summarization that can cover both single-modal output summarization and multimodal output summarization. In our framework, we consider three different scenarios and propose the respective unsupervised graph-based multimodal summarization models without the requirement of any manually annotated document-summary pairs for training: (1) generic multimodal ranking, (2) modal-dominated multimodal ranking, and (3) non-redundant text-image multimodal ranking. Furthermore, an image-text similarity estimation model is introduced to measure the semantic similarity between image and text. Experiments show that our proposed models outperform the single-modal summarization methods on both automatic and human evaluation metrics. Besides, our models can also improve the single-modal summarization with the guidance of the multimedia information. This study can be applied as the benchmark for further study on multimodal summarization task.
Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Medical Term and Status Generation From Chinese Clinical Dialogue With Multi-Granularity Transformer
abstract
This paper describes a generative model for extracting medical terms and their status from Chinese medical dialogues. Notably, the extracted semantic information plays an essential role in downstream tasks such as automatic medical scribe and automatic diagnosis system. However, how to effectively leverage dialogue context to generate medical terms and their corresponding status accurately remains less explored. Existing generative methods treat dialogue text as concentrated long text without considering the characteristics of conversation, such as colloquialism, redundancy, interactions, etc. Various colloquial medical information is frequently discussed between doctor and patient. Each of the speakers (doctor and patient) plays a specific role in the goals of interaction. Thus the role information and interactions between utterances are vital. Besides, current generative methods only utilize character-level tokens ignoring the word-level tokens, which is the smallest meaningful utterance in Chinese. In this paper, we propose a Multi-granularity Transformer (MGT) model to enhance the dialogue context understanding from multi-granularity features. We introduce word-level information by adapting a Lattice-based encoder with our proposed relative position encoding method. We further introduce utterance-level interaction information by proposing a Role Access Controlled Attention (RaCa) mechanism. Experimental results on two benchmark datasets illustrate our model's validity and effectiveness, achieving state-of-the-art performance on both datasets.
Lu Xiang, Xiaomian Kang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Multimodal Summarization with Guidance of Multimodal Reference
abstract
Multimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Order-ranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments.
Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Haoran Li 0001, Chengqing Zong, Changliang Li
AAAI2
2020 Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual Summarization
abstract
Cross-lingual summarization aims at summarizing a document in one language (e.g., Chinese) into another language (e.g., English).In this paper, we propose a novel method inspired by the translation pattern in the process of obtaining a cross-lingual summary.We first attend to some words in the source text, then translate them into the target language, and summarize to get the final summary.Specifically, we first employ the encoder-decoder attention distribution to attend to the source words.Second, we present three strategies to acquire the translation probability, which helps obtain the translation candidates for each source word.Finally, each summary word is generated either from the neural distribution or from the translation candidates of source words.Experimental results on Chinese-to-English and English-to-Chinese summarization tasks have shown that our proposed method can significantly outperform the baselines, achieving comparable performance with the state-of-the-art.
Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACL2
2020 Dual Attention Network for Cross-lingual Entity Alignment
abstract
Cross-lingual Entity alignment is an essential part of building a knowledge graph, which can help integrate knowledge among different language knowledge graphs.In the real KGs, there exists an imbalance among the information in the same hierarchy of corresponding entities, which results in the heterogeneity of neighborhood structure, making this task challenging.To tackle this problem, we propose a dual attention network for cross-lingual entity alignment (DAEA).Specifically, our dual attention consists of relation-aware graph attention and hierarchical attention.The relation-aware graph attention aims at selectively aggregating multi-hierarchy neighborhood information to alleviate the difference of heterogeneity among counterpart entities.The hierarchical attention adaptively aggregates the low-hierarchy and the high-hierarchy information, which is beneficial to balance the neighborhood information of counterpart entities and distinguish noncounterpart entities with similar structures.Finally, we treat cross-lingual entity alignment as a process of linking prediction.Experimental results on three real-world cross-lingual entity alignment datasets have shown the effectiveness of DAEA.
Jian Sun 0035, Yu Zhou 0001, Chengqing Zong
COLING2
2020 Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity Granularity
abstract
Previous studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized.ii) Granularity mismatch: the current KG methods utilize the entity as the basic granularity, while NMT utilizes the sub-word as the granularity, making the KG different to be utilized in NMT.To alleviate above problems, we propose a multi-task learning method on sub-entity granularity.Specifically, we first split the entities in KG and sentence pairs into sub-entity granularity by using joint BPE.Then we utilize the multi-task learning to combine the machine translation task and knowledge reasoning task.The extensive experiments on various translation tasks have demonstrated that our method significantly outperforms the baseline models in both translation quality and handling the entities.
Yang Zhao 0007, Lu Xiang, Junnan Zhu, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
COLING5
2020 A Knowledge-driven Generative Model for Multi-implication Chinese Medical Procedure Entity Normalization
abstract
Medical entity normalization, which links medical mentions in the text to entities in knowledge bases, is an important research topic in medical natural language processing.In this paper, we focus on Chinese medical procedure entity normalization.However, nonstandard Chinese expressions and combined procedures present challenges in our problem.The existing strategies relying on the discriminative model are poorly to cope with normalizing combined procedure mentions.We propose a sequence generative framework to directly generate all the corresponding medical procedure entities.we adopt two strategies: category-based constraint decoding and category-based model refining to avoid unrealistic results.The method is capable of linking entities when a mention contains multiple procedure concepts and our comprehensive experiments demonstrate that the proposed model can achieve remarkable improvements over existing baselines, particularly significant in the case of multi-implication Chinese medical procedures.
Jinghui Yan, Lu Xiang, Yu Zhou 0001, Chengqing Zong
EMNLP (1)4
2020 Knowledge Graphs Enhanced Neural Machine Translation
abstract
Knowledge graphs (KGs) store much structured information on various entities, many of which are not covered by the parallel sentence pairs of neural machine translation (NMT). To improve the translation quality of these entities, in this paper we propose a novel KGs enhanced NMT method. Specifically, we first induce the new translation results of these entities by transforming the source and target KGs into a unified semantic space. We then generate adequate pseudo parallel sentence pairs that contain these induced entity pairs. Finally, NMT model is jointly trained by the original and pseudo sentence pairs. The extensive experiments on Chinese-to-English and Englishto-Japanese translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling the induced entities.
Yang Zhao 0007, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
IJCAI3
2020 Structurally Comparative Hinge Loss for Dependency-Based Neural Text Representation
abstract
Dependency-based graph convolutional networks (DepGCNs) are proven helpful for text representation to handle many natural language tasks. Almost all previous models are trained with cross-entropy (CE) loss, which maximizes the posterior likelihood directly. However, the contribution of dependency structures is not well considered by CE loss. As a result, the performance improvement gained by using the structure information can be narrow due to the failure in learning to rely on this structure information. To face the challenge, we propose the novel structurally comparative hinge (SCH) loss function for DepGCNs. SCH loss aims at enlarging the margin gained by structural representations over non-structural ones. From the perspective of information theory, this is equivalent to improving the conditional mutual information of model decision and structure information given text. Our experimental results on both English and Chinese datasets show that by substituting SCH loss for CE loss on various tasks, for both induced structures and structures from an external parser, performance is improved without additional learnable parameters. Furthermore, the extent to which certain types of examples rely on the dependency structure can be measured directly by the learned margin, which results in better interpretability. In addition, through detailed analysis, we show that this structure margin has a positive correlation with task performance and structure induction of DepGCNs, and SCH loss can help model focus more on the shortest dependency path between entities. We achieve the new state-of-the-art results on TACRED, IMDB, and Zh. Literature datasets, even compared with ensemble and BERT baselines.
Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Memory Consolidation for Contextual Spoken Language Understanding with Dialogue Logistic Inference
abstract
Dialogue contexts are proven helpful in the spoken language understanding (SLU) system and they are typically encoded with explicit memory representations.However, most of the previous models learn the context memory with only one objective to maximizing the SLU performance, leaving the context memory under-exploited.In this paper, we propose a new dialogue logistic inference (DLI) task to consolidate the context memory jointly with SLU in the multi-task framework.DLI is defined as sorting a shuffled dialogue session into its original logical order and shares the same memory encoder and retrieval mechanism as the SLU model.Our experimental results show that various popular contextual SLU models can benefit from our approach, and improvements are quite impressive, especially in slot filling.
He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACL (1)2
2019 NCLS: Neural Cross-Lingual Summarization
abstract
Junnan Zhu, Qian Wang, Yining Wang, Yu Zhou, Jiajun Zhang, Shaonan Wang, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junnan Zhu, Qian Wang 0061, Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong
EMNLP/IJCNLP (1)4
2019 Input Method for Human Translators: A Novel Approach to Integrate Machine Translation Effectively and Imperceptibly
abstract
Computer-aided translation (CAT) systems are the most popular tool for helping human translators efficiently perform language translation. To further improve the translation efficiency, there is an increasing interest in applying machine translation (MT) technology to upgrade CAT. To thoroughly integrate MT into CAT systems, in this article, we propose a novel approach: a new input method that makes full use of the knowledge adopted by MT systems, such as translation rules, decoding hypotheses, and n-best translation lists. The proposed input method contains two parts: a phrase generation model, allowing human translators to type target sentences quickly, and an n-gram prediction model, helping users choose perfect MT fragments smoothly. In addition, to tune the underlying MT system to generate the input method preferable results, we design a new evaluation metric for the MT system. The proposed input method integrates MT effectively and imperceptibly, and it is particularly suitable for many target languages with complex characters, such as Chinese and Japanese. The extensive experiments demonstrate that our method saves more than 23% in time and over 42% in keystrokes, and it also improves the translation quality by more than 5 absolute BLEU scores compared with the strong baseline, i.e., post-editing using Google Pinyin.
Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Source Critical Reinforcement Learning for Transferring Spoken Language Understanding to a New Language
abstract
To deploy a spoken language understanding (SLU) model to a new language, language transferring is desired to avoid the trouble of acquiring and labeling a new big SLU corpus. An SLU corpus is a monolingual corpus with domain/intent/slot labels. Translating the original SLU corpus into the target language is an attractive strategy. However, SLU corpora consist of plenty of semantic labels (slots), which general-purpose translators cannot handle well, not to mention additional culture differences. This paper focuses on the language transferring task given a small in-domain parallel SLU corpus. The in-domain parallel corpus can be used as the first adaptation on the general translator. But more importantly, we show how to use reinforcement learning (RL) to further adapt the adapted translator, where translated sentences with more proper slot tags receive higher rewards. Our reward is derived from the source input sentence exclusively, unlike reward via actor-critical methods or computing reward with a ground truth target sentence. Hence we can adapt the translator the second time, using the big monolingual SLU corpus from the source language. We evaluate our approach on Chinese to English language transferring for SLU systems. The experimental results show that the generated English SLU corpus via adaptation and reinforcement learning gives us over 97% in the slot F1 score and over 84% accuracy in domain classification. It demonstrates the effectiveness of the proposed language transferring method. Compared with naive translation, our proposed method improves domain classification accuracy by relatively 22%, and the slot filling F1 score by relatively more than 71%.
He Bai 0002, Yu Zhou 0001, Jiajun Zhang 0001, Mei-Yuh Hwang, Chengqing Zong
COLING2
2018 MSMO: Multimodal Summarization with Multimodal Output
abstract
Multimodal summarization has drawn much attention due to the rapid growth of multimedia data.The output of the current multimodal summarization systems is usually represented in texts.However, we have found through experiments that multimodal output can significantly improve user satisfaction for informativeness of summaries.In this paper, we propose a novel task, multimodal summarization with multimodal output (MSMO).To handle this task, we first collect a large-scale dataset for MSMO research.We then propose a multimodal attention model to jointly generate text and select the most relevant image from the multimodal input.Finally, to evaluate multimodal outputs, we construct a novel multimodal automatic evaluation (MMAE) method which considers both intramodality salience and intermodality relevance.The experimental results show the effectiveness of MMAE.
Junnan Zhu, Haoran Li 0001, Tianshang Liu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
EMNLP4
2017 Augmenting Neural Sentence Summarization Through Extractive Summarization
Junnan Zhu, Haoran Li 0001, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
NLPCC5
2016 Bilingual Semantic Role Labeling Inference via Dual Decomposition
abstract
This article focuses on bilingual Semantic Role Labeling (SRL); its goal is to annotate semantic roles on both sides of the parallel bilingual texts (bi-texts). Since rich bilingual information is encoded, bilingual SRL has been applied in many natural-language processing (NLP) tasks such as machine translation (MT), cross-lingual information retrieval (IR), and the like. A feasible way of performing bilingual SRL is using monolingual SRL systems to perform SRL on each side of bi-texts separately. However, it is difficult to obtain consistent SRL results on both sides of bi-texts in this way. Some works have tried to jointly infer bilingual SRL because there are many complementary language cues on both sides of bi-texts and they reported better performance than monolingual systems. However, there are two limits in the existing methods. First, the existing methods often require high inference costs due to the complex objective function. Second, the existing methods fully adopt the candidates generated by monolingual SRL systems, but many candidates are discarded in the argument pruning or identification stage of monolingual systems. In this article, we propose two strategies to overcome these limits. We utilize a simple but efficient technique: Dual Decomposition to search for consistent results for both sides of bi-texts. On the other hand, we propose a method called Bi-Directional Projection (BDP) to recover arguments discarded in monolingual SRL systems. We evaluate our method on a standard parallel benchmark: the OntoNotes dataset. The experimental results show that our method yields significant improvements over the state-of-the-art monolingual systems. In addition, our approach is also better and faster than existing methods due to BDP and Dual Decomposition.
Haitong Yang, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2016 Abstractive Cross-Language Summarization via Translation Model Enhanced Predicate Argument Structure Fusing
abstract
Cross-language multidocument summarization is the task to generate a summary in a target language (e.g., Chinese) from a collection of documents in a different source language (e.g., English). Previous methods such as the extractive and compressive algorithms focus only on single sentence selection and compression, which cannot make full use of the similar sentences containing complementary information. Furthermore, the translation model knowledge is not fully explored in previous approaches. To address these two problems, we propose in this paper an abstractive cross-language summarization framework. First, the source language documents are translated into target language with a machine translation system. Then, the method constructs a pool of bilingual concepts and facts represented by the bilingual elements of the source-side predicate-argument structures (PAS) and their target-side counterparts. Finally, new summary sentences are produced by fusing bilingual PAS elements with the integer linear programming algorithm to maximize both of the salience and translation quality of the PAS elements. The experimental results on English-to-Chinese cross-language summarization demonstrate that our proposed method outperforms the state-of-the-art extractive systems in both automatic and manual evaluations.
Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 A New Input Method for Human Translators: Integrating Machine Translation Effectively and Imperceptibly
Guoping Huang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
IJCAI3
2015 Exploring Diverse Features for Statistical Machine Translation Model Pruning
abstract
In phrase-based and hierarchical phrase-based statistical machine translation systems, translation performance depends heavily on the size and quality of the translation table. To meet the requirements of making a real-time response, some research has been performed to filter the translation table. However, most existing methods are always based on one or two constraints that act as hard rules, such as not allowing phrase-pairs with low translation probabilities. These approaches sometimes make constraints rigid because they consider only a single factor instead of composite factors. Based on the considerations above, in this paper, we propose a machine learning-based framework that integrates multiple features for translation model pruning. Experimental results show that our framework is effective by pruning 80% of the phrase-pairs and 70% of the hierarchical rules, while retaining the quality of the translation models when using the BLEU evaluation metric. Our study further shows that our method can select the most useful phrase-pairs and rules, including those that are low in frequency but still very useful.
Mei Tu, Yu Zhou 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Enhancing Grammatical Cohesion: Generating Transitional Expressions for SMT
abstract
Transitional expressions provide glue that holds ideas together in a text and enhance the logical organization, which together help improve readability of a text. However, in most current statistical machine translation (SMT) systems, the outputs of compound-complex sentences still lack proper transitional expressions. As a result, the translations are often hard to read and understand. To address this issue, we propose two novel models to encourage generating such transitional expressions by introducing the source compoundcomplex sentence structure (CSS). Our models include a CSS-based translation model, which generates new CSS-based translation rules, and a generative transfer model, which encourages producing transitional expressions during decoding. The two models are integrated into a hierarchical phrase-based translation system to evaluate their effectiveness. The experimental results show that significant improvements are achieved on various test data meanwhile the translations are more cohesive and smooth.
Mei Tu, Yu Zhou 0001, Chengqing Zong
ACL (1)2
2013 Handling Ambiguities of Bilingual Predicate-Argument Structures for Statistical Machine Translation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
ACL (1)3
2013 An Efficient Framework to Extract Parallel Units from Comparable Data
Lu Xiang, Yu Zhou 0001, Chengqing Zong
NLPCC2
2013 Unsupervised Tree Induction for Tree-based Translation
abstract
In current research, most tree-based translation models are built directly from parse trees. In this study, we go in another direction and build a translation model with an unsupervised tree structure derived from a novel non-parametric Bayesian model. In the model, we utilize synchronous tree substitution grammars (STSG) to capture the bilingual mapping between language pairs. To train the model efficiently, we develop a Gibbs sampler with three novel Gibbs operators. The sampler is capable of exploring the infinite space of tree structures by performing local changes on the tree nodes. Experimental results show that the string-to-tree translation system using our Bayesian tree structures significantly outperforms the strong baseline string-to-tree system using parse trees.
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
Trans. Assoc. Comput. Linguistics3
2012 Machine Translation by Modeling Predicate-Argument Structure Transformation
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
COLING3
2012 Tree-based Translation without using Parse Trees
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
COLING3
2011 Simple but Effective Approaches to Improving Tree-to-tree Model
Feifei Zhai, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
MTSummit3
2010 A Novel Reordering Model Based on Multi-layer Phrase for Statistical Machine Translation
Yanqing He, Yu Zhou 0001, Chengqing Zong, Huilin Wang
COLING2
2009 Approach to Selecting Best Development Set for Phrase-Based Statistical Machine Translation
Yu Zhou 0001, Chengqing Zong
PACLIC2
2003 Automatic evaluation of sentence fluency
abstract
In the machine translation (MT) system, how to evaluate the sentence fluency of the translation results is an important research topic. Most of the current methods are based on the similarity of output words compared with the reference translations, which don't specially address the evaluation of the sentence fluency according to the syntactic structure. This paper proposes a statistical approach to the problem, which is based on the n-gram language model and reference-independent. Our approach is a beneficial compensation to the reference-dependent methods and has better robustness than those methods based on the syntactic analysis. The preliminary experimental results indicate that the approach basically reflects the reality of human's judgment.
Yu Zhou 0001, Chengqing Zong, Fuji Ren
SMC2