VLDB 2026 Research / reviewers in the wild / expert
Yang Zhao 0007
dblp:50/2082-7
· DBLP profile ↗
41ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-6599-3600ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DART: Disambiguation-Aware Reasoning for Video-guided Machine TranslationabstractVideo-guided Machine Translation (VMT) seeks to enhance translation quality by incorporating contextual information derived from paired short video clips.However, many VMT samples are text-sufficient; even when visual information is needed, only minimal cues are required.Aiming to tackle these issues, we propose a novel framework DART (Disambiguation-Aware Reasoning for Videoguided Machine Translation).Reinforcement learning is used to incorporate multimodal large language models' multimodal reasoning into VMT.The model dynamically switches between text-only processing and multimodal integration, contingent on the necessity of visual disambiguation.Furthermore, we present TVRF (Translation-oriented Video Relevance Filtering), a systematic pipeline for constructing training data based on multimodal relevance to translation.This pipeline filters samples where video information is translationrelevant, mitigating training collapse caused by video-irrelevant data in conventional VMT.Experimental results show that our approach improves multimodal information utilization in VMT, yielding gains in both translation quality and computational efficiency.* Equal corresponding authors.Reasoning: Okay, I need to translate the input sentence ... So the answer is "哦!".(809 tokens in total) DA T Existing LMRMs … "It's hitting his yacht."Translation: 哦!(Oh!) Reasoning: Okay, so I need to translate the sentence .. Boyu Guan, Chuang Han, Yang Zhao 0007, Chengqing Zong |
ACL (1) | 3 |
| 2025 | Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine TranslationabstractYupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yupu Liang, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
ACL (1) | 4 |
| 2025 | TriFine: A Large-Scale Dataset of Vision-Audio-Subtitle for Tri-Modal Machine Translation and Benchmark with Fine-Grained Annotated TagsabstractCurrent video-guided machine translation (VMT) approaches primarily use coarse-grained visual information, resulting in information redundancy, high computational overhead, and neglect of audio content. Our research demonstrates the significance of fine-grained visual and audio information in VMT from both data and methodological perspectives. From the data perspective, we have developed a large-scale dataset TriFine, the first vision-audio-subtitle tri-modal VMT dataset with annotated multimodal fine-grained tags. Each entry in this dataset not only includes the triples found in traditional VMT datasets but also encompasses seven fine-grained annotation tags derived from visual and audio modalities. From the methodological perspective, we propose a Fine-grained Information-enhanced Approach for Translation (FIAT). Experimental results have shown that, in comparison to traditional coarse-grained methods and text-only models, our fine-grained approach achieves superior performance with lower computational overhead. These findings underscore the pivotal role of fine-grained annotated information in advancing the field of VMT. Boyu Guan, Yang Zhao 0007, Chengqing Zong |
COLING | 3 |
| 2025 | From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image TranslationabstractDocument Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in complex-layout scenarios, frequently loses the crucial document coherence, leading to chaotic text. To overcome this problem, we introduce a novel end-to-end network, named Zoom-out DIT (ZoomDIT), inspired by human translation procedures. It jointly accomplishes the multi-level tasks including word positioning, sentence recognition & translation, and document organization, based on a fine-to-coarse zoom-out framework, to progressively realize “chaotic words to coherent document” and improve translation. We further contribute a new large-scale DIT dataset with multi-level fine-grained labels. Extensive experiments on public and our new dataset demonstrate significant improvements in translation quality towards complex-layout document images, offering a robust solution for reorganizing the chaotic OCR outputs to a coherent document translation. Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
COLING | 5 |
| 2025 | SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationabstractVideo-guided Machine Translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips.Mainstream VMT approaches typically incorporate multimodal information by uniformly sampling frames from the input videos.However, this paradigm frequently incurs significant computational overhead and introduces redundant multimodal content, which degrades both efficiency and translation quality.To tackle these challenges, we propose SHIFT (Selected Helpful Informative Frame for Translation).It is a lightweight, plug-andplay framework designed for VMT with Multimodal Large Language Models (MLLMs).SHIFT adaptively selects a single informative key frame when visual context is necessary; otherwise, it relies solely on textual input.This process is guided by a dedicated clustering module and a selector module.Experimental results demonstrate that SHIFT enhances the performance of MLLMs on the VMT task while simultaneously reducing computational cost, without sacrificing generalization ability. Boyu Guan, Chuang Han, Yupu Liang, Yang Zhao 0007, Chengqing Zong |
EMNLP | 6 |
| 2025 | ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (5) | 6 |
| 2025 | SimulPL: Aligning Human Preferences in Simultaneous Machine TranslationabstractSimultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios have distinct preferences, such as accurate translations, simpler syntax, and no unnecessary latency. Aligning SiMT models with these human preferences is crucial to improve their performances. However, this issue still remains unexplored. Additionally, preference optimization for SiMT task is also challenging. Existing methods focus solely on optimizing the generated responses, ignoring human preferences related to latency and the optimization of read/write policy during the preference optimization phase. To address these challenges, we propose Simultaneous Preference Learning (SimulPL), a preference learning framework tailored for the SiMT task. In the SimulPL framework, we categorize SiMT human preferences into five aspects: **translation quality preference**, **monotonicity preference**, **key point preference**, **simplicity preference**, and **latency preference**. By leveraging the first four preferences, we construct human preference prompts to efficiently guide GPT-4/4o in generating preference data for the SiMT task. In the preference optimization phase, SimulPL integrates **latency preference** into the optimization objective and enables SiMT models to improve the read/write policy, thereby aligning with human preferences more effectively. Experimental results indicate that SimulPL exhibits better alignment with human preferences across all latency levels in Zh$\rightarrow$En, De$\rightarrow$En and En$\rightarrow$Zh SiMT tasks. Our data and code will be available at https://github.com/EurekaForNLP/SimulPL. Donglei Yu, Yang Zhao 0007, Yangyifan Xu, Yu Zhou 0001, Chengqing Zong |
ICLR | 2 |
| 2025 | Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image TranslationabstractDocument Image Translation (DIT) aims to translate texts on document images from one language to another. It is a multi-modal task involving cooperation of text and layout. Current approaches either handle layout and translation as separate processes, risking accumulative errors, or use vanilla end-to-end encoder-decoder models to capture layout implicitly, often suffering inadequate layout incorporation. We argue that a favorable framework should explicitly engage layout-specific modules and properly organize them toward translation. For this, we first revisit two key layouts: the geometric layout reflecting word's spatial positions, and the logical layout depicting word's logical order. Then, a novel pipeline (understand layout $\rightarrow$→ translate text) is determined to prioritize layouts such that preceding layouts contribute to translation. Following this pipeline, we introduce Unified Document Image Translation (UniDIT), a comprehensive framework that unifies layout with translation in one network. It is devised to leverage each module's advantage, and provide an elaborate feature-conductive flow for module communication globally. A novel bridging mechanism is also introduced to adapt layout features conducive to translation. We further contribute DITransv2, a large-scale fine-grained benchmark that includes heterogeneous and complex document layouts. Extensive experiments on DITransv2 and additional established benchmarks demonstrate UniDIT outperforms previous state-of-the-arts in all aspects. Yupu Liang, Cong Ma 0002, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine TranslationabstractText image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained knowledge supervision to make it consistent between recognition and translation modules. In this paper, we propose a novel TIMT method named as BabyNet, which is optimized with hierarchical parental supervision to improve translation performance. Inspired by genetic recombination and variation in the field of genetics, the proposed BabyNet is inherited from the recognition and translation parent models with a variation module of which parameters can be updated when training on the TIMT task. Meanwhile, hierarchical and multi-granularity supervision from parent models is introduced to bridge the gap between inherited modules in BabyNet. Extensive experiments on both synthetic and real-world TIMT tests show that our proposed method significantly outperforms existing methods. Further analyses of various parent model combinations show the good generalization of our method. Cong Ma 0002, Yupu Liang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
LREC/COLING | 5 |
| 2024 | Vector Quantization Knowledge Transfer for End-to-End Text Image Machine TranslationabstractEnd-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning continuous features from related tasks of text image recognition (TIR) or machine translation (MT) to alleviate the problem of data limitation, but the alignment in continuous vector space is extremely difficult and it inevitably introduces fitting errors resulting in significant performance degradation. To better align TIMT features with MT semantic features, we propose a novel Vector Quantization Knowledge Transfer (VQKT) method that employs a trainable codebook to quantize continuous features into discrete space. The quantization distribution of the MT feature is utilized as the teacher distribution to guide the TIMT model to generate similar discrete codes. Through alignment and knowledge transfer based on probability distribution, the TIMT model can better imitate the feature representation of the MT teacher model and generate high-quality target language translation. Extensive experiments demonstrate VQKT significantly outperforms the existing end-to-end TIMT performance. Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICASSP | 3 |
| 2024 | A 2.5 kHz 50.57 dB Linearized VCO ADC Using 6 µm LTPS TFTsabstractThis paper presents a VCO-based ADC design for use in flexible electronics, specifically leveraging Low-Temperature Polysilicon Thin-Film Transistor (LTPS TFT) technology. The design achieves a resolution of 8.11 bits at a 20 MHz sampling rate with a bandwidth (BW) of 2.5kHz. A linearity compensation technique combining a resistive input stage with a frequency-dependent resistor (FDR)-based feedback loop significantly enhances the VCO linearity. Simulation results validate the design, exhibiting an R2value of 0.9999 for the VCO tuning curve, a Signal-to-Noise and Distortion Ratio (SNDR) of 50.57 dB, and a Spurious-Free Dynamic Range (SFDR) of 52.34 dB at a supply voltage of 10V. The proposed design also features the lowest Figure of Merit (FoM) (0.73 nJ/conversion-step) compared with the state-of-the-art. Wangzilu Lu, Chao Wang 0101, Yang Zhao 0007, Jian Zhao 0004, Yongfu Li 0002 |
ISCAS | 5 |
| 2024 | Document Image Machine Translation with Dynamic Multi-pre-trained Models AssemblingabstractYupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yupu Liang, Cong Ma 0002, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
NAACL-HLT | 5 |
| 2024 | 🚀 TableRocket: An Efficient and Effective Framework for Table Reconstruction
Liucheng Pang, Cong Ma 0002, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
PRCV (7) | 4 |
| 2024 | Knowledge Graph Guided Neural Machine Translation with Dynamic Reinforce-selected TriplesabstractPrevious methods incorporating knowledge graphs (KGs) into neural machine translation (NMT) adopt a static knowledge utilization strategy, that introduces many useless knowledge triples and makes the useful triples difficult to be utilized by NMT. To address this problem, we propose a KG guided NMT model with dynamic reinforce-selected triples. The proposed methods could dynamically select the different useful knowledge triples for different source sentences. Specifically, the proposed model contains two components: (1) knowledge selector, that dynamically selects useful knowledge triples for a source sentence, and (2) knowledge guided NMT (KgNMT), that utilizes the selected triples to guide the translation of NMT. Meanwhile, to overcome the non-differentiable problem and guide the training procedure, we propose a policy gradient strategy to encourage the model to select useful triples and improve the generation probability of gold target sentence. Various experimental results show that the proposed method can significantly outperform the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Xiaomian Kang, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | Modal Contrastive Learning Based End-to-End Text Image Machine TranslationabstractText image machine translation (TIMT) aims at directly translating text in the source language embedded in images into the target language. Most existing systems follow the cascaded pipeline diagram from recognition to translation, which suffers from the problem of error propagation, parameter redundancy, and information reduction. The end-to-end model has the potential to alleviate these issues via bridging the recognition and translation models. However, the challenge is the data limitation and modality gap between text and image. In this paper, we propose a novel end-to-end model, namely Modal contrastive learning based End-to-end Text Image Machine Translation (METIMT), which alleviates these issues through end-to-end text image machine translation architecture and modal contrastive learning. Specifically, an image encoder is designed to encode images into the same feature space of corresponding text sentences, with the guidance of an intramodal and inter-modal contrastive learning module. To further promote the research of text image machine translation, we have constructed one synthetic and two real-world datasets. Extensive experiments show that our lighter, faster model outperforms not only existing pipeline methods but also state-of-the-art end-to-end models on both synthetic and real-world evaluation sets. Our code and dataset will be released to the public. Cong Ma 0002, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Multilingual Knowledge Graph Completion with Language-Sensitive Multi-Graph AttentionabstractMultilingual Knowledge Graph Completion (KGC) aims to predict missing links with multilingual knowledge graphs.However, existing approaches suffer from two main drawbacks: (a) alignment dependency: the multilingual KGC is always realized with joint entity or relation alignment, which introduces additional alignment models and increases the complexity of the whole framework; (b) training inefficiency: the trained model will only be used for the completion of one target KG, although the data from all KGs are used simultaneously.To address these drawbacks, we propose a novel multilingual KGC framework with languagesensitive multi-graph attention such that the missing links on all given KGs can be inferred by a universal knowledge completion model.Specifically, we first build a relational graph neural network by sharing the embeddings of aligned nodes to transfer language-independent knowledge.Meanwhile, a language-sensitive multi-graph attention (LSMGA) is proposed to deal with the information inconsistency among different KGs.Experimental results show that our model achieves significant improvements on the DBP-5L and E-PKG datasets.1 Rongchuan Tang, Yang Zhao 0007, Chengqing Zong, Yu Zhou 0001 |
ACL (1) | 2 |
| 2023 | Multi-teacher Knowledge Distillation for End-to-End Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (1) | 4 |
| 2023 | E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (6) | 4 |
| 2023 | A Gain and Bandwidth Individually Tunable ExG Analog Frontend with 516nVrms Noise for Flexible Biomedical SensorsabstractThis paper presents a low-power, low-noise, gain and bandwidth individually tunable analog front-end (AFE) for ExG signals. The proposed three-stage AFE with a bandwidth programmable amplifier enables individually tuning the bandpass cutoff frequencies in the range of 0.4 to 1.9kHz as well as the gain from 40 to 63dB. Designed in a$0.35 \mu\mathrm{m}$CMOS process with an area of 0.3mm2, the AFE achieves over 120dB CMRR with input referred noise of 516nVrms and a noise efficiency factor of 2.57. The chip consumes$2.5 \mu\mathrm{A}$at 1.8V supply. Yanxing Suo, Yang Zhao 0007, Yongfu Li 0002, Yan Liu 0016, Yong Lian 0001 |
ISCAS | 3 |
| 2023 | Zero-shot language extension for dialogue state tracking via pre-trained models and multi-auxiliary-tasks fine-tuning
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong |
Knowl. Based Syst. | 2 |
| 2022 | Improving End-to-End Text Image Translation From the Auxiliary Text Translation TaskabstractEnd-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent research. However, data sparsity limits the performance of end-to-end text image translation. Multi-task learning is a nontrivial way to alleviate this problem via exploring knowledge from complementary related tasks. In this paper, we propose a novel text translation enhanced text image translation, which trains the end-to-end model with text translation as an auxiliary task. By sharing model parameters and multi-task training, our model is able to take full advantage of easily-available large-scale text parallel corpus. Extensive experimental results show our proposed method outperforms existing end-to-end methods, and the joint multi-task learning with both text translation and recognition tasks achieves better results, proving translation and recognition auxiliary tasks are complementary.1 Cong Ma 0002, Mei Tu, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001 |
ICPR | 6 |
| 2022 | An Input-Output Regulated Adaptive Ramp for Fast Load Transition of PWM Buck ConvertorabstractAn input-output regulated adaptive ramp ($\mathrm{IOR}^{2})$ for fast load transition of pulse width modulation (PWM) buck convertor is presented. The scheme employs an adaptive ramp regulated by input and the voltage from the error amplifier to achieve fast load response and low line and load regulation rate. Simulation shows that an under/overshoot voltage of −32 mV and 35 mV, with $8 \mu$ and $8.3 \mu \mathrm{s}$ recovery time are respectively obtained for the load current stepping between 1 A and 1.7 A. The line/load regulation rate is respectively $12.4 \mu \mathrm{V} / \mathrm{V}$ and $1.04 \mathrm{mV} / \mathbf{A}$. Implemented in a $0.25 \mu \mathrm{m}$ BCD process, the proposed IOR2PWM regulator is capable of converting input voltage of 5 V to 65 V to output range of 3.3 V to 60 V with adjustable switching frequency up to 2.2 MHz, showing a peak efficiency of 94.5% at 1 A load current. Bingbing He, Haoran Li 0001, Yongfu Li 0002, Yan Liu 0016, Yang Zhao 0007 |
ISCAS | 6 |
| 2022 | Enhancing Lexical Translation Consistency for Document-Level Neural Machine TranslationabstractDocument-level neural machine translation (DocNMT) has yielded attractive improvements. In this article, we systematically analyze the discourse phenomena in Chinese-to-English translation, and focus on the most obvious ones, namely lexical translation consistency. To alleviate the lexical inconsistency, we propose an effective approach that is aware of the words which need to be translated consistently and constrains the model to produce more consistent translations. Specifically, we first introduce a global context extractor to extract the document context and consistency context, respectively. Then, the two types of global context are integrated into a encoder enhancer and a decoder enhancer to improve the lexical translation consistency. We create a test set to evaluate the lexical consistency automatically. Experiments demonstrate that our approach can significantly alleviate the lexical translation inconsistency. In addition, our approach can also substantially improve the translation quality compared to sentence-level Transformer. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2021 | Synchronous Interactive Decoding for Multilingual Neural Machine TranslationabstractTo simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future information, and therefore may suffer from some issues like the unbalanced outputs. In this paper, we present a new approach for synchronous interactive multilingual neural machine translation (SimNMT), which predicts each target language output simultaneously and interactively using historical and future information of all target languages. Specifically, we first propose a synchronous cross-interactive decoder in which generation of each target output does not only depend on its generated sequences, but also relies on its future information, as well as history and future contexts of other target languages. Then, we present a new interactive multilingual beam search algorithm that enables synchronous interactive decoding of all target languages in a single model. We take two target languages as an example to illustrate and evaluate the proposed SimNMT model on IWSLT datasets. The experimental results demonstrate that our method achieves significant improvements over several advanced NMT and MNMT models. Qian Wang 0061, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
AAAI | 4 |
| 2021 | A Low-Power Heart Rate Sensor with Adaptive Heartbeat Locked LoopabstractPhotoplethysmography (PPG) is one of the widely used noninvasive heart rate (HR) monitoring techniques in wearable devices. Lighting of the LED dominates the power consumption of a PPG sensor. Lowering the LED lighting duration is an effectively approach to save power. This paper proposes an adaptive heartbeat locked loop (AHBLL) technique, which can dynamically adjust the dividing ratio N according to the heart rate derivative (HRD). In this way, the LED pulse duty cycle can be significantly reduced to save power. To improve the robustness of the AHBLL system, the relationship between HRD and the dividing ratio N is theoretically analyzed, which provides an optimal design guideline. To verify the proposed technique, an HR sensing circuit including an HRD detector is designed and simulated. The results show that the LED power consumption is reduced by 2.2~3.3× compared with the state-of-the-art heartbeat locked loop (HBLL) technique. Zhouchen Ma, Cheng Chen 0054, Min Wang 0014, Yang Zhao 0007, Liang Ying, Guoxing Wang, Jian Zhao 0004 |
ISCAS | 4 |
| 2021 | A Resource-Efficient, Robust QRS Detector Using Data Compression and Time-Sharing ArchitectureabstractIn this paper, we proposed a resource-efficient 'QRS' detector with superior detection accuracy. Inspired by the strategy of the folded architecture, we adopted a reconfigurable time-sharing computation unit with a pipeline schedule. To further precisely locate the position of the 'R' peak and minimize the extra hardware cost, we designed the position calibration unit (PCU) based on the data compression technique. The proposed architecture was implemented on Xilinx Zynq-7000 with Verilog programming language. The proposed architecture achieves a sensitivity, Se of 99.76%, a precision, +P of 99.85%, and a detection error rate, DER of 0.40% on MIT-BIH database, which attains the best performance compared to state-of-the-art designs. Furthermore, the proposed architecture achieves a better hardware efficiency with 13×, 1.28×, and 4.35× reductions in computing resources, storage memory, and power consumption, respectively. Weihong Yan, Yuxin Ji, Lining Hu, Yang Zhao 0007, Yan Liu 0016, Yongfu Li 0002 |
ISCAS | 5 |
| 2021 | Zero-Shot Deployment for Cross-Lingual Dialogue System
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong |
NLPCC (2) | 2 |
| 2021 | Robust Cross-lingual Task-oriented DialogueabstractCross-lingual dialogue systems are increasingly important in e-commerce and customer service due to the rapid progress of globalization. In real-world system deployment, machine translation (MT) services are often used before and after the dialogue system to bridge different languages. However, noises and errors introduced in the MT process will result in the dialogue system's low robustness, making the system's performance far from satisfactory. In this article, we propose a novel MT-oriented noise enhanced framework that exploits multi-granularity MT noises and injects such noises into the dialogue system to improve the dialogue system's robustness. Specifically, we first design a method to automatically construct multi-granularity MT-oriented noises and multi-granularity adversarial examples, which contain abundant noise knowledge oriented to MT. Then, we propose two strategies to incorporate the noise knowledge: (i) Utterance-level adversarial learning and (ii) Knowledge-level guided method. The former adopts adversarial learning to learn a perturbation-invariant encoder, guiding the dialogue system to learn noise-independent hidden representations. The latter explicitly incorporates the multi-granularity noises, which contain the noise tokens and their possible correct forms, into the training and inference process, thus improving the dialogue system's robustness. Experimental results on three dialogue models, two dialogue datasets, and two language pairs have shown that the proposed framework significantly improves the performance of the cross-lingual dialogue system. Lu Xiang, Junnan Zhu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Medical Term and Status Generation From Chinese Clinical Dialogue With Multi-Granularity TransformerabstractThis paper describes a generative model for extracting medical terms and their status from Chinese medical dialogues. Notably, the extracted semantic information plays an essential role in downstream tasks such as automatic medical scribe and automatic diagnosis system. However, how to effectively leverage dialogue context to generate medical terms and their corresponding status accurately remains less explored. Existing generative methods treat dialogue text as concentrated long text without considering the characteristics of conversation, such as colloquialism, redundancy, interactions, etc. Various colloquial medical information is frequently discussed between doctor and patient. Each of the speakers (doctor and patient) plays a specific role in the goals of interaction. Thus the role information and interactions between utterances are vital. Besides, current generative methods only utilize character-level tokens ignoring the word-level tokens, which is the smallest meaningful utterance in Chinese. In this paper, we propose a Multi-granularity Transformer (MGT) model to enhance the dialogue context understanding from multi-granularity features. We introduce word-level information by adapting a Lattice-based encoder with our proposed relative position encoding method. We further introduce utterance-level interaction information by proposing a Role Access Controlled Attention (RaCa) mechanism. Experimental results on two benchmark datasets illustrate our model's validity and effectiveness, achieving state-of-the-art performance on both datasets. Lu Xiang, Xiaomian Kang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity GranularityabstractPrevious studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized.ii) Granularity mismatch: the current KG methods utilize the entity as the basic granularity, while NMT utilizes the sub-word as the granularity, making the KG different to be utilized in NMT.To alleviate above problems, we propose a multi-task learning method on sub-entity granularity.Specifically, we first split the entities in KG and sentence pairs into sub-entity granularity by using joint BPE.Then we utilize the multi-task learning to combine the machine translation task and knowledge reasoning task.The extensive experiments on various translation tasks have demonstrated that our method significantly outperforms the baseline models in both translation quality and handling the entities. Yang Zhao 0007, Lu Xiang, Junnan Zhu, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
COLING | 1 |
| 2020 | Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningabstractDocument-level neural machine translation has yielded attractive improvements.However, majority of existing methods roughly use all context sentences in a fixed scope.They neglect the fact that different source sentences need different sizes of context.To address this problem, we propose an effective approach to select dynamic context so that the document-level translation model can utilize the more useful selected context sentences to produce better translations.Specifically, we introduce a selection module that is independent of the translation module to score each candidate context sentence.Then, we propose two strategies to explicitly select a variable number of context sentences and feed them into the translation module.We train the two modules end-to-end via reinforcement learning.A novel reward is proposed to encourage the selection and utilization of dynamic context sentences.Experiments demonstrate that our approach can select adaptive context sentences for different source sentences, and significantly improves the performance of document-level translation methods. Xiaomian Kang, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
EMNLP (1) | 2 |
| 2020 | Knowledge Graphs Enhanced Neural Machine TranslationabstractKnowledge graphs (KGs) store much structured information on various entities, many of which are not covered by the parallel sentence pairs of neural machine translation (NMT). To improve the translation quality of these entities, in this paper we propose a novel KGs enhanced NMT method. Specifically, we first induce the new translation results of these entities by transforming the source and target KGs into a unified semantic space. We then generate adequate pseudo parallel sentence pairs that contain these induced entity pairs. Finally, NMT model is jointly trained by the original and pseudo sentence pairs. The extensive experiments on Chinese-to-English and Englishto-Japanese translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling the induced entities. Yang Zhao 0007, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong |
IJCAI | 1 |
| 2020 | Non-autoregressive Neural Machine Translation with Distortion Model
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
NLPCC (1) | 3 |
| 2020 | Synchronous bidirectional inference for neural sequence generation
Jiajun Zhang 0001, Yang Zhao 0007, Chengqing Zong |
Artif. Intell. | 3 |
| 2019 | Addressing the Under-Translation Problem from the Entropy PerspectiveabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance in recent years. However, the under-translation problem still remains a big challenge. In this paper, we focus on the under-translation problem and attempt to find out what kinds of source words are more likely to be ignored. Through analysis, we observe that a source word with a large translation entropy is more inclined to be dropped. To address this problem, we propose a coarse-to-fine framework. In coarse-grained phase, we introduce a simple strategy to reduce the entropy of highentropy words through constructing the pseudo target sentences. In fine-grained phase, we propose three methods, including pre-training method, multitask method and two-pass method, to encourage the neural model to correctly translate these high-entropy words. Experimental results on various translation tasks show that our method can significantly improve the translation quality and substantially reduce the under-translation cases of high-entropy words. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhongjun He, Hua Wu 0003 |
AAAI | 1 |
| 2019 | Attention With Sparsity Regularization for Neural Machine Translation and SummarizationabstractThe attention mechanism has become thede factostandard component in neural sequence to sequence tasks, such as machine translation and abstractive summarization. It dynamically determines which parts in the input sentence should be focused on when generating each word in the output sequence. Ideally, only few relevant input words should be attended to at each decoding time step and the attention weight distribution should be sparse and sharp. However, previous methods have no good mechanism to control this attention weight distribution. In this paper, we propose a sparse attention model in which a sparsity regularization term is designed to augment the objective function. We explore two kinds of regularizations:$L_{\infty }$-norm regularization and minimum entropy regularization, both of which aim to sharpen the attention weight distribution. Extensive experiments on both neural machine translation and abstractive summarization demonstrate that our proposed sparse attention model can substantially outperform the strong baselines. And the detailed analyses reveal that the final attention distribution indeed becomes sparse and sharp. Jiajun Zhang 0001, Yang Zhao 0007, Haoran Li 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Addressing Troublesome Words in Neural Machine TranslationabstractOne of the weaknesses of Neural Machine Translation (NMT) is in handling lowfrequency and ambiguous words, which we refer as troublesome words.To address this problem, we propose a novel memoryenhanced NMT method.First, we investigate different strategies to define and detect the troublesome words.Then, a contextual memory is constructed to memorize which target words should be produced in what situations.Finally, we design a hybrid model to dynamically access the contextual memory so as to correctly translate the troublesome words.The extensive experiments on Chineseto-English and English-to-German translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling troublesome words. Yang Zhao 0007, Jiajun Zhang 0001, Zhongjun He, Chengqing Zong, Hua Wu 0003 |
EMNLP | 1 |
| 2018 | Phrase Table as Recommendation Memory for Neural Machine TranslationabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance recently. However, several studies indicate that NMT often generates fluent but unfaithful translations. In this paper, we propose a method to alleviate this problem by using a phrase table as recommendation memory. The main idea is to add bonus to words worthy of recommendation, so that NMT can make correct predictions. Specifically, we first derive a prefix tree to accommodate all the candidate target phrases by searching the phrase translation table according to the source sentence.Then, we construct a recommendation word set by matching between candidate target phrases and previously translated target words by NMT. After that, we determine the specific bonus value for each recommendable word by using the attention vector and phrase translation probability. Finally,we integrate this bonus value into NMT to improve the translation results. The extensive experiments demonstrate that the proposed methods obtain remarkable improvements over the strong attention based NMT. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
IJCAI | 1 |
| 2018 | Exploiting Pre-Ordering for Neural Machine Translation
Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
LREC | 1 |
| 2018 | A Comparable Study on Model Averaging, Ensembling and Reranking in NMT
Yuchen Liu 0007, Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong |
NLPCC (2) | 4 |
| 2017 | Towards Neural Machine Translation with Partially Aligned CorporaabstractWhile neural machine translation (NMT) has become the new paradigm, the parameter optimization requires large-scale parallel data which is scarce in many domains and language pairs. In this paper, we address a new translation scenario in which there only exists monolingual corpora and phrase pairs. We propose a new method towards translation with partially aligned sentence pairs which are derived from the phrase pairs and monolingual corpora. To make full use of the partially aligned corpora, we adapt the conventional NMT training method in two aspects. On one hand, different generation strategies are designed for aligned and unaligned target words. On the other hand, a different objective function is designed to model the partially aligned parts. The experiments demonstrate that our method can achieve a relatively good result in such a translation scenario, and tiny bitexts can boost translation quality to a large extent. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhengshan Xue |
IJCNLP(1) | 2 |