EDBT 2026 Demo / reviewers in the wild / expert
Mingxuan Wang
dblp:43/11214
· DBLP profile ↗
61ranked-venue papers
11as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 8 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper ImagesabstractJiayi Tuo, Cheng Tang, Zihan Wang, Chenyue Zhou, Yao Li, Yanbiao Ma, Chao Wang, Wei Dai, Mingxuan Wang, Shitong Qin, Ziwei Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiayi Tuo, Cheng Tang 0004, Chenyue Zhou, Yanbiao Ma, Chao Wang 0003, Wei Dai 0015, Mingxuan Wang, Shitong Qin |
ACL (1) | 9 |
| 2026 | Enhancing Grounded Multimodal Named Entity Recognition with Dual-Level Representation Alignment
Xinzhi Wang 0001, Mingxuan Wang, Ruishen Liu, Xiangfeng Luo |
KSEM (7) | 2 |
| 2026 | Understanding the effects of emotional and cognitive factors on trust in human-robot interaction using interpretable machine learning
Mingxuan Wang, Ningning Zhao, Hongtao Zheng, Susana Liu |
Adv. Eng. Informatics | 1 |
| 2026 | MSTNet: Multi-Scale Contextual Analysis Network for Semantic Segmentation of Remote Sensing ImagesabstractABSTRACT With the rapid development of deep learning, research on semantic segmentation of remote sensing images has made significant progress. However, there are common problems in remote sensing images, such as large‐scale differences between different types of objects and unbalanced sample numbers, which leads to poor semantic segmentation results, especially for small targets and rare types of objects. To address these challenges, a remote sensing image semantic segmentation method based on multi‐scale contextual information analysis named MSTNet is innovatively proposed. Its core design includes the semantic information enhancement module (SIE) of feature adaptive clustering, which strengthens the feature expression of different categories through adaptive clustering to alleviate sample imbalance; the weighted feature fusion module (WFF) adaptively aggregates cross‐level features, cooperates with the multi‐scale context enhancement module (MSCE), combines convolution and Transformer operations, and deeply mines local and global contexts to cope with scale changes; in addition, the network also contains a pixel space feature optimisation module (SFEM) to enhance spatial details. Experiments on the UAVid, LoveDA, Potsdam and Vaihingen datasets show that MSTNet significantly improves the ability to handle scale changes and imbalance problems and reaches advanced levels in key indicators such as OA, mIoU, and mF1, proving that MSTNet achieves competitive or even better performance. Longbao Wang, Mingxuan Wang, Xiaoliang Luo, Lvchun Wang, Chong Long |
IET Image Process. | 2 |
| 2026 | Dynamic Re-Energized RLS for Adaptive Fault-Tolerant Tracking Control of Discrete-Time SystemsabstractThis paper addresses the critical challenge of adaptive fault-tolerant state tracking for discrete-time systems in the presence of controller failures and external disturbances. Traditional gradient-based adaptive control methods suffer from slow convergence and degraded transient performance, while conventional recursive least squares (RLS) algorithms exhibit reduced fault tolerance after convergence. To overcome these limitations, a novel closed-loop architecture is proposed, integrating a dynamic re-energized RLS mechanism. The framework employs simultaneous estimation of controller faults and adaptive parameters via function approximation, ensuring robust fault compensation. A key innovation is the error-driven re-energization mechanism, which resets the RLS gain matrix to isolate historical estimation biases while preserving fast convergence. Rigorous Lyapunov analysis proves global exponential stability of the closed-loop system. Comparative simulations demonstrate superior performance in handling abrupt and concurrent faults, validating significant improvements in convergence speed and sustained fault tolerance over baseline approaches. Mingxuan Wang, Jiuxiang Dong |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Horizon-Greedy Q-Ensembles Regularized Decision Transformer: An Offline Reinforcement Learning Approach for Robotic TasksabstractOffline reinforcement learning (RL) offers a promising paradigm for learning policies from precollected datasets. Nonetheless, applying it to robotic control poses significant challenges, including non-Markovian dynamics and the scarcity of high-quality demonstrations, both of which can undermine the performance of existing methods. To address these issues, this work introduces the horizon-GreedyQ-ensemblesRegularized decisionTransformer (GQRT), an offline RL algorithm tailored for robotic tasks. GQRT leverages a transformer-based policy for trajectory modeling, thereby enabling effective long-horizon decision-making in non-Markovian settings. To alleviate the lack of expert demonstrations, we develop a multistep horizon-greedy policy evaluation mechanism that stitches suboptimal sequences into improved trajectories. Furthermore, to cope with mixed-quality demonstrations collected from diverse sources, we incorporate aQ-ensemble with lower confidence bound regularization, which ensures more stable and reliable value estimation. Extensive experiments on robotic locomotion and manipulation benchmarks demonstrate that GQRT achieves state-of-the-art performance, validating its robustness in complex robotic scenarios. Botao Dong, Xin Dong 0021, Mingxuan Wang, Hongtian Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Empowering Self-Learning of LLMs: Inner Knowledge Explicitation as a CatalystabstractSelf-learning of Large Language Models (LLMs) facilitates their advancement towards super-intelligence by training with self-synthesized experiences. However, a critical challenge is the amplification of hallucinations in generated data during iterative self-learning, underscoring the need for reliable data selection. To address this, we investigate the mechanism of Inner Knowledge Explicitation, which involves explicitly extracting the inner knowledge from memory of LLMs, to concurrently improves reasoning, and enables reliable self-learning data selection. This paper introduces a Self Knowledge Explicitation Learning (SKE-Learn) framework, which equips the LLMs with meta-skills to explicitly extract, verify and utilize inner knowledge for reasoning. By leveraging these meta-skills, SKE-Learn establishes a self-learning approach that ensures reliable selection of self-synthetic data. This approach enhances performance through iterative self-learning while mitigating the problem of hallucinations. Empirical results from six benchmarks demonstrate that Inner Knowledge Explicitation improves reasoning by serving as a more effective prompting method. Additionally, SKE-Learn, based on the verifiability of explicit knowledge, shows consistent performance improvements over multiple self-training iterations, with an average performance increase from 52.79% to 56.54% across all benchmarks. Furthermore, Inner Knowledge Explicitation provides explanation and intervention space during LLM's generation process. Shijue Huang, Wanjun Zhong, Deng Cai 0002, Fanqi Wan, Mingxuan Wang, Ruifeng Xu 0001 |
AAAI | 6 |
| 2025 | Automated Coding Utterances Toward Chinese Course Core Competence with Large Language Models
Mingyang Yue 0001, Tengda Qi, Mingxuan Wang, Wang Ruan, Jun He 0009, Bo Sun 0006, Guomin Zheng |
ICIC (23) | 3 |
| 2025 | CDE-CCL: A Classroom Dialogue Evaluation Framework for Chinese Core Literacy via LLMabstractClassroom dialogue evaluation is a crucial component of the teaching process assessment. It not only enhances the quality of classroom interaction and student engagement but also fosters the development of students’ core literacy. By leveraging classroom dialogue evaluation, teachers can implement personalized teaching strategies, overcoming the problem of unequal distribution of educational resources caused by regional disparities. However, current classroom dialogue evaluation frameworks are manually implemented, which results in low automation, high time consumption, and a lack of systematic assessment related to core literacy. To address these limitations, we propose the Classroom Dialogue Evaluation Framework for Chinese Core Literacy (CDE-CCL). This framework integrates Chinese Core Literacy, dialogue round segmentation, prompt engineering, and LoRA fine-tuning techniques to encode complete classroom dialogues, thereby enabling automated evaluation of classroom dialogues. Experimental results demonstrate the superiority of CDE-CCL in both tasks of classroom dialogue segmentation and encoding over existing methods across various aspects and educational levels. Mingxuan Wang, Mingyang Yue 0001, Guomin Zheng, Jun He 0009, Bo Sun 0006 |
IJCNN | 1 |
| 2025 | Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesabstractLarge Language Models (LLMs), such as OpenAI’s o1 and DeepSeek’s R1, excel at advanced reasoning tasks like math and coding via Reinforcement Learning with Verifiable Rewards (RLVR), but still struggle with puzzles solvable by humans without domain knowledge. We introduce ENIGMATA, the first comprehensive suite tailored for improving LLMs with puzzle reasoning skills. It includes 36 tasks across 7 categories, each with: 1) a generator that produces unlimited examples with controllable difficulty, and 2) a rule-based verifier for automatic evaluation. This generator-verifier design supports scalable, multi-task RL training, fine-grained analysis, and seamless RLVR integration. We further propose ENIGMATA-Eval, a rigorous benchmark, and develop optimized multi-task RLVR strategies. Our trained model, Qwen2.5-32B-ENIGMATA, consistently surpasses o3-mini-high and o1 on the puzzle reasoning benchmarks like ENIGMATA-Eval, ARC-AGI (32.8%), and ARC-AGI 2 (0.6%). It also generalizes well to out-of-domain puzzle benchmarks and mathematical reasoning, with little multi-tasking trade-off. When trained on larger models like Seed1.5-Thinking (20B activated parameters and 200B total parameters), puzzle data from ENIGMATA further boosts SoTA performance on advanced math and STEM reasoning tasks such as AIME (2024-2025), BeyondAIME and GPQA (Diamond), showing nice generalization benefits of ENIGMATA. This work offers a unified, controllable framework for advancing logical reasoning in LLMs. Project page: https://seed-enigmata.github.io. Jiangjie Chen, Qianyu He, Aili Chen, Zhicheng Cai, Weinan Dai, Hongli Yu, Jiaze Chen, Qiying Yu, Hao Zhou 0012, Mingxuan Wang |
NeurIPS | 12 |
| 2025 | ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable GenerationabstractGenerative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex centroids, reducing generation complexity and enhancing scalability. Additionally, SLM incorporates a flexible implementation of classifier-free guidance, enhancing unconditional generation performance. Extensive experiments on DNA promoter and enhancer design, protein design, character-level and large-vocabulary language modeling demonstrate the competitive performance and strong potential of SLM. Our code can be found at https://github.com/GenSI-THUAIR/SLM. Yuxuan Song 0002, Jingjing Gong, Qiying Yu, Zheng Zhang 0001, Mingxuan Wang, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 7 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 36 |
| 2025 | SCLResNet and DSAF: A self-supervised contrastive learning and deep self-attention fusion-based multimodal network for predicting central lymph node metastasis in papillary thyroid carcinoma
Wenjuan Huang, Mengzhuo Sun, Mingxuan Wang, Hongzhuo Qi, Zengyao Liu, Qiujun Wang, Ruitao Wang, Xuemei Ding |
Artif. Intell. Medicine | 6 |
| 2025 | Beyond looks: the effect of voice personality traits and gestures in virtual agent interactionsabstractAbstract With the expansion of extended reality (XR), virtual agents engage more with humans across fields. While the influence of anthropomorphic traits on user perceptions is studied, the combined effects of voice traits and gestures need further investigation. We explored the impact of a virtual agent's voice personality and gestures on subjective perception and visual behaviors in a hospital guidance setting. Using a 2 × 2 within-subjects design, we assessed the influence of voice traits (calm vs. lively) and gestures (with vs. without) through subjective reports and eye tracking. Results show that the impact of gestures on content comprehension ease varied depending on the voice personality traits. Specifically, a lively voice with gestures significantly increased comprehension while a calm voice with gestures did not. Additionally, a calm voice led to longer average fixation durations on body-related areas of interest (AOIs), shorter time to first fixation, and fewer fixation counts on dialog content-related AOIs. Gestures decreased fixation counts on face-related AOIs. For some eye metrics, the impact of gestures varied depending on voice personality traits, whereas for others, the influence of voice traits depended on the presence of gestures. Thus, voice traits and gestures influence user perceptions and visual attention differently. Mingxuan Wang, Nasi Wang, Zhanxun Dong |
Interact. Comput. | 1 |
| 2024 | PolyVoice: Language Models for Speech to Speech TranslationabstractWith the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks.
Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate the impact of language modeling approaches in this area.
In this study, we introduce PolyVoice, a language model-based framework designed for S2ST systems. Our framework comprises three decoder-only language models: a translation language model, a duration language model, and a speech synthesis language model.
These language models employ different types of prompts to extract learned information effectively. By utilizing unsupervised semantic units, our framework can transfer semantic information across these models, making it applicable even to unwritten languages.
We evaluate our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish language pairs. Experimental results demonstrate that \method outperforms the state-of-the-art encoder-decoder model, producing voice-cloned speech with high translation and audio quality.
Speech samples are available at https://polyvoice.github.io. Qianqian Dong, Zhiying Huang, Qi Tian 0001, Chen Xu 0008, Tom Ko, Yunlong Zhao 0004, Tang Li 0001, Xuxin Cheng, Fengpeng Yue, Ye Bai 0001, Lu Lu 0015, Zejun Ma 0001, Yuping Wang 0005, Mingxuan Wang, Yuxuan Wang 0002 |
ICLR | 17 |
| 2024 | Diffusion Glancing Transformer for Parallel Sequence-to-Sequence LearningabstractLihua Qian, Mingxuan Wang, Yang Liu, Hao Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lihua Qian, Mingxuan Wang |
NAACL-HLT | 2 |
| 2024 | Amphion: an Open-Source Audio, Music, and Speech Generation ToolkitabstractAmphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion. Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li 0030, Haorui He, Chaoren Wang, Songting Liu, Junan Zhang, Zihao Fang, Haopeng Chen, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Kai Chen 0026, Haizhou Li 0001, Zhizheng Wu 0001 |
SLT | 15 |
| 2024 | SingVisio: Visual analytics of diffusion model for singing voice conversion
Liumeng Xue, Chaoren Wang, Mingxuan Wang, Xueyao Zhang, Zhizheng Wu 0001 |
Comput. Graph. | 3 |
| 2024 | Residual adaptive sparse hybrid attention transformer for image super resolution
Hai Huan, Mingxuan Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | CTC-based Non-autoregressive Speech TranslationabstractChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chen Xu 0008, Qingxuan Sun, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao 0001, Anxiang Ma |
ACL (1) | 9 |
| 2023 | SESCORE2: Learning Text Generation Evaluation via Synthesizing Realistic MistakesabstractIs it possible to train a general metric for evaluating text generation quality without humanannotated ratings?Existing learned metrics either perform unsatisfactorily across text generation tasks or require human ratings for training on specific tasks.In this paper, we propose SESCORE2, a self-supervised approach for training a model-based metric for text generation evaluation.The key concept is to synthesize realistic model mistakes by perturbing sentences retrieved from a corpus.The primary advantage of the SESCORE2 is its ease of extension to many other languages while providing reliable severity estimation.We evaluate SESCORE2 and previous methods on four text generation tasks across three languages.SESCORE2 outperforms unsupervised metric PRISM on four text generation evaluation benchmarks, with a Kendall improvement of 0.078.Surprisingly, SESCORE2 even outperforms the supervised BLEURT and COMET on multiple text generation tasks.The code and data are available at https://github.com/ xu1998hz/SEScore2 1 . Wenda Xu, Xian Qian, Mingxuan Wang, Lei Li 0005, William Yang Wang |
ACL (1) | 3 |
| 2023 | BLEURT Has Universal Translations: An Analysis of Automatic Metrics by Minimum Risk TrainingabstractAutomatic metrics play a crucial role in machine translation.Despite the widespread use of n-gram-based metrics, there has been a recent surge in the development of pre-trained model-based metrics that focus on measuring sentence semantics.However, these neural metrics, while achieving higher correlations with human evaluations, are often considered to be black boxes with potential biases that are difficult to detect.In this study, we systematically analyze and compare various mainstream and cutting-edge automatic metrics from the perspective of their guidance for training machine translation systems.Through Minimum Risk Training (MRT), we find that certain metrics exhibit robustness defects, such as the presence of universal adversarial translations in BLEURT and BARTScore.In-depth analysis suggests two main causes of these robustness deficits: distribution biases in the training datasets, and the tendency of the metric paradigm.By incorporating token-level constraints, we enhance the robustness of evaluation metrics, which in turn leads to an improvement in the performance of machine translation systems.Codes are available at https://github.com/ powerpuffpomelo/fairseq_mrt. Tao Wang 0086, Chengqi Zhao, Shujian Huang, Jiajun Chen 0001, Mingxuan Wang |
ACL (1) | 6 |
| 2023 | Leveraging per Image-Token Consistency for Vision-Language Pre-trainingabstractMost existing vision-language pre-training (VLP) approaches adopt cross-modal masked language modeling (CMLM) to learn vision-language associations. However, we find that CMLM is insufficient for this purpose according to our observations: (1) Modality bias: a considerable amount of masked tokens in CMLM can be recovered with only the language information, ignoring the visual inputs. (2) Underutilization of the unmasked tokens: CMLM primarily focuses on the masked token but it cannot simultaneously leverage other tokens to learn vision-language associations. To handle those limitations, we propose EPIC (lEveraging Per Image-Token Consistency for vision-language pre-training). In EPIC, for each image-sentence pair, we mask tokens that are salient to the image (i.e., Saliency-based Masking Strategy) and replace them with alternatives sampled from a language model (i.e., Inconsistent Token Generation Procedure), and then the model is required to determine for each token in the sentence whether it is consistent with the image (i.e., Image-Token Consistency Task). The proposed EPIC method is easily combined with pre-training methods. Extensive experiments show that the combination of the EPIC method and state-of-the-art pre-training approaches, including ViLT, ALBEF, METER, and X-VLM, leads to significant improvements on downstream tasks. Our coude is released at https://github.com/gyhdog99/epic Yunhao Gou, Tom Ko, Hansi Yang, James T. Kwok, Yu Zhang 0006, Mingxuan Wang |
CVPR | 6 |
| 2023 | M3ST: Mix at Three Levels for Speech TranslationabstractHow to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It’s well known that data augmentation is an efficient method to improve performance for many tasks by enlarging the dataset. In this paper, we propose Mix at three levels for Speech Translation (M3ST) method to increase the diversity of the augmented training corpus. Specifically, we conduct two phases of fine-tuning based on a pre-trained model using external machine translation (MT) data. In the first stage of fine-tuning, we mix the training corpus at three levels, including word level, sentence level and frame level, and fine-tune the entire model with mixed data. At the second stage of fine-tuning, we take both original speech sequences and original text sequences in parallel into the model to fine-tune the network, and use Jensen-Shannon divergence to regularize their outputs. Experiments on MuST-C speech translation benchmark and analysis show M3ST outperforms current strong baselines and achieves state-of-the-art results on eight directions with an average BLEU of 29.9. Xuxin Cheng, Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Yuexian Zou |
ICASSP | 5 |
| 2023 | Recent Advances in Direct Speech-to-text TranslationabstractRecently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges --- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work. Chen Xu 0008, Rong Ye, Qianqian Dong, Chengqi Zhao, Tom Ko, Mingxuan Wang, Tong Xiao 0001 |
IJCAI | 6 |
| 2023 | Only 5% Attention Is All You Need: Efficient Long-range Document-level Neural Machine TranslationabstractZihan Liu, Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang |
IJCNLP (1) | 5 |
| 2023 | CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
Chutong Meng, Junyi Ao, Tom Ko, Mingxuan Wang, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | GigaST: A 10, 000-hour Pseudo Speech Translation Corpus
Rong Ye, Chengqi Zhao, Tom Ko, Chutong Meng, Tao Wang 0086, Mingxuan Wang |
INTERSPEECH | 6 |
| 2023 | Bilingual attention based neural machine translation
Liyan Kang, Shaojie He, Mingxuan Wang, Jinsong Su |
Appl. Intell. | 3 |
| 2022 | Learning When to Translate for Streaming SpeechabstractHow to find proper moments to generate partial sentence translation given a streaming speech input?Existing approaches waitingand-translating for a fixed duration often break the acoustic units in speech, since the boundaries between acoustic units in speech are not even.In this paper, we propose MoSST, a simple yet effective method for translating streaming speech content.Given a usually long speech sequence, we develop an efficient monotonic segmentation module inside an encoder-decoder model to accumulate acoustic information incrementally and detect proper speech unit boundaries for the input in speech translation task.Experiments on multiple translation directions of the MuST-C dataset show that MoSST outperforms existing methods and achieves the best trade-off between translation quality (BLEU) and latency.Our code is available at https://github. com/dqqcasia/mosst. Yaoming Zhu, Mingxuan Wang, Lei Li 0005 |
ACL (1) | 3 |
| 2022 | STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationabstractHow to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data?Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities.In this paper, we propose the Speech-TExt Manifold Mixup (STEMM) method to calibrate such discrepancy.Specifically, we mix up the representation sequences of different modalities, and take both unimodal speech sequences and multimodal mixed sequences as input to the translation model in parallel, and regularize their output predictions with a selflearning framework.Experiments on MuST-C speech translation benchmark and further analysis show that our method effectively alleviates the cross-modal representation discrepancy, and achieves significant improvements over a strong baseline on eight translation directions.* indicates corresponding authors. Qingkai Fang, Rong Ye, Lei Li 0005, Yang Feng 0004, Mingxuan Wang |
ACL (1) | 5 |
| 2022 | Unified Multimodal Punctuation Restoration Framework for Mixed-Modality CorpusabstractThe punctuation restoration task aims to correctly punctuate the output transcriptions of automatic speech recognition systems. Previous punctuation models, either using text only or demanding the corresponding audio, tend to be constrained by real scenes, where unpunctuated sentences are a mixture of those with and without audio. This paper proposes a unified multimodal punctuation restoration framework, named UniPunc, to punctuate the mixed sentences with a single model. UniPunc jointly represents audio and non-audio samples in a shared latent space, based on which the model learns a hybrid representation and punctuates both kinds of samples. We validate the effectiveness of the UniPunc on real-world datasets, which outperforms various strong baselines (e.g. BERT, MuSe) by at least 0.8 overall F1 scores, making a new state-of-the-art. Extensive experiments show that UniPunc’s design is a pervasive solution: by grafting onto previous models, UniPunc enables them to punctuate on the mixed corpus. Our code is available at github.com/Yaoming95/UniPunc Yaoming Zhu, Shanbo Cheng, Mingxuan Wang |
ICASSP | 4 |
| 2022 | switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch Decoder
Zhenqiao Song, Hao Zhou 0012, Lihua Qian, Jingjing Xu 0001, Shanbo Cheng, Mingxuan Wang, Lei Li 0005 |
ICLR | 6 |
| 2022 | Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech TranslationabstractDirect Speech-to-speech translation (S2ST) has drawn more and more attention recently. The task is very challenging due to data scarcity and complex speech-to-speech mapping. In this paper, we report our recent achievements in S2ST. Firstly, we build a S2ST Transformer baseline which outperforms the original Translatotron. Secondly, we utilize the external data by pseudo-labeling and obtain a new state-of-the-art result on the Fisher English-to-Spanish test set. Indeed, we exploit the pseudo data with a combination of popular techniques which are not trivial when applied to S2ST. Moreover, we evaluate our approach on both syntactically similar (Spanish-English) and distant (English-Chinese) language pairs. Our implementation is available at https://github.com/fengpeng-yue/speech-to-speech-translation. Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, Yu Zhang 0006 |
INTERSPEECH | 4 |
| 2022 | Cross-modal Contrastive Learning for Speech TranslationabstractHow can we learn unified representations for spoken utterances and their written text?Learning similar representations for semantically similar speech and text is important for speech translation.To this end, we propose ConST, a cross-modal contrastive learning method for end-to-end speech-to-text translation.We evaluate ConST and a variety of previous baselines on a popular benchmark MuST-C.Experiments show that the proposed ConST consistently outperforms the previous methods, and achieves an average BLEU of 29.4.The analysis further verifies that ConST indeed closes the representation gap of different modalities -its learned representation improves the accuracy of cross-modal speechtext retrieval from 4% to 88%.Code and models are available at https://github. com/ReneeYe/ConST. Rong Ye, Mingxuan Wang, Lei Li 0005 |
NAACL-HLT | 2 |
| 2022 | LightSeq2: Accelerated Training for Transformer-Based Models on GPUsabstractTransformer-based neural models are used in many AI applications. Training these models is expensive, as it takes huge GPU resources and long duration. It is challenging because typical data like sentences have variable lengths, and Transformer's computation patterns are more complex than convolutional neural networks. Existing systems either only focus on model inference or optimization for only BERT-like encoder models. In this paper, we present LightSeq2, a system to accelerate training for a general family of Transformer models on GPUs. We propose a series of GPU optimization techniques tailored to the specific computation flow and memory access patterns of Transformer models. LightSeq2 supports many model architectures, including BERT (encoder-only), GPT (decoder-only), Transformer (encoder-decoder), and vision Transformer. Our experiments for a variety of models and benchmarks show that LightSeq2 is consistently faster (1.4-3.5 x) than previous systems on different GPUs. In particular, it gains 308 % training speedup compared with existing systems on a large public machine translation benchmark (WMTI4 English-German). Guyue Huang, Xian Qian, Yufei Ding 0001, Mingxuan Wang, Lei Li 0005 |
SC | 7 |
| 2021 | Consecutive Decoding for Speech-to-text TranslationabstractSpeech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal cross-lingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, TED English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms the previous state-of-the-art methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 2 |
| 2021 | Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationabstractAn end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding system which is composed of auditory perception and cognitive processing. In this paper, we propose Listen-Understand-Translate, (LUT), a unified framework with triple supervision signals to decouple the end-to-end speech-to-text translation task. LUT is able to guide the acoustic encoder to extract as much information from the auditory input. In addition, LUT utilizes a pre-trained BERT model to enforce the upper encoder to produce as much semantic information as possible, without extra data. We perform experiments on a diverse set of speech translation benchmarks, including Librispeech English-French, IWSLT English-German and TED English-Chinese. Our results demonstrate LUT achieves the state-of-the-art performance, outperforming previous methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 3 |
| 2021 | Finding Sparse Structures for Domain Specific Neural Machine TranslationabstractNeural machine translation often adopts the fine-tuning approach to adapt to specific domains. However, nonrestricted fine-tuning can easily degrade on the general domain and over-fit to the target domain. To mitigate the issue, we propose Prune-Tune, a novel domain adaptation method via gradual pruning. It learns tiny domain-specific sub-networks during fine-tuning on new domains. Prune-Tune alleviates the over-fitting and the degradation problem without model modification. Furthermore, Prune-Tune is able to sequentially learn a single network with multiple disjoint domain-specific sub-networks for multiple domains. Empirical experiment results show that Prune-Tune outperforms several strong competitors in the target domain test set without sacrificing the quality on the general domain in both single and multi-domain settings. The source code and data are available at https://github.com/ohlionel/Prune-Tune. Jianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu, Lei Li 0005 |
AAAI | 3 |
| 2021 | Learning Language Specific Sub-network for Multilingual Machine TranslationabstractZehui Lin, Liwei Wu, Mingxuan Wang, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mingxuan Wang, Lei Li 0005 |
ACL/IJCNLP (1) | 3 |
| 2021 | Contrastive Learning for Many-to-many Multilingual Neural Machine TranslationabstractXiao Pan, Mingxuan Wang, Liwei Wu, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mingxuan Wang, Lei Li 0005 |
ACL/IJCNLP (1) | 2 |
| 2021 | Glancing Transformer for Non-Autoregressive Neural Machine TranslationabstractLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lihua Qian, Hao Zhou 0012, Mingxuan Wang, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
ACL/IJCNLP (1) | 4 |
| 2021 | Learning Kernel-Smoothed Machine Translation with Retrieved ExamplesabstractHow to effectively adapt neural machine translation (NMT) models according to emerging cases without retraining?Despite the great success of neural machine translation, updating the deployed models online remains a challenge.Existing non-parametric approaches that retrieve similar examples from a database to guide the translation process are promising but are prone to overfit the retrieved examples.However, non-parametric methods are prone to overfit the retrieved examples.In this work, we propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER), an effective approach to adapt neural machine translation models online.Experiments on domain adaptation and multi-domain machine translation datasets show that even without expensive retraining, KSTER is able to achieve improvement of 1.1 to 1.5 BLEU scores over the best existing online adaptation methods.The code and trained models are released at https://github.com/jiangqn/KSTER. Qingnan Jiang, Mingxuan Wang, Shanbo Cheng, Shujian Huang, Lei Li 0005 |
EMNLP (1) | 2 |
| 2021 | End-to-End Speech Translation via Cross-Modal Progressive TrainingabstractEnd-to-end speech translation models have become a new trend in research due to their potential of reducing error propagation.However, these models still suffer from the challenge of data scarcity.How to effectively use unlabeled or other parallel corpora from machine translation is promising but still an open problem.In this paper, we propose Cross Speech-Text Network (XSTNet), an end-to-end model for speech-to-text translation.XSTNet takes both speech and text as input and outputs both transcription and translation text.The model benefits from its three key design aspects: a self-supervised pretrained sub-network as the audio encoder, a multi-task training objective to exploit additional parallel bilingual text, and a progressive training procedure.We evaluate the performance of XSTNet and baselines on the MuST-C En-X and LibriSpeech En-Fr datasets.In particular, XSTNet achieves state-of-the-art results on all language directions with an average BLEU of 28.8, outperforming the previous best method by 3.2 BLEU.Code, models, cases, and more detailed analysis are available at https://github.com/ReneeYe/XSTNet. Rong Ye, Mingxuan Wang, Lei Li 0005 |
Interspeech | 2 |
| 2021 | Generative Imagination Elevates Machine TranslationabstractThere are common semantics shared across text and images.Given a sentence in a source language, whether depicting the visual scene helps translation into a target language?Existing multimodal neural machine translation methods (MNMT) require triplets of bilingual sentence -image for training and tuples of source sentence -image for inference.In this paper, we propose ImagiT, a novel machine translation method via visual imagination.ImagiT first learns to generate visual representation from the source sentence, and then utilizes both source sentence and the "imagined representation" to produce a target translation.Unlike previous methods, it only needs the source sentence at the inference time.Experiments demonstrate that ImagiT benefits from visual imagination and significantly outperforms the text-only neural machine translation baselines.Further analysis reveals that the imagination process in ImagiT helps fill in missing information when performing the degradation strategy. Quanyu Long, Mingxuan Wang, Lei Li 0005 |
NAACL-HLT | 2 |
| 2020 | Towards Making the Most of BERT in Neural Machine TranslationabstractGPT-2 and BERT demonstrate the effectiveness of using pre-trained language models (LMs) on various natural language processing tasks. However, LM fine-tuning often suffers from catastrophic forgetting when applied to resource-rich tasks. In this work, we introduce a concerted training framework (CTnmt) that is the key to integrate the pre-trained LMs to neural machine translation (NMT). Our proposed CTnmt} consists of three techniques: a) asymptotic distillation to ensure that the NMT model can retain the previous pre-trained knowledge; b) a dynamic switching gate to avoid catastrophic forgetting of pre-trained knowledge; and c) a strategy to adjust the learning paces according to a scheduled policy. Our experiments in machine translation show CTnmt gains of up to 3 BLEU score on the WMT14 English-German language pair which even surpasses the previous state-of-the-art pre-training aided NMT by 1.4 BLEU score. While for the large WMT14 English-French task with 40 millions of sentence-pairs, our base model still significantly improves upon the state-of-the-art Transformer big model by more than 1 BLEU score. Mingxuan Wang, Hao Zhou 0012, Chengqi Zhao, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
AAAI | 2 |
| 2020 | Improving Maximum Likelihood Training for Text Generation with Density Ratio EstimationabstractAutoregressive neural sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is usually limited and the input data distributions are different at training and inference stages. Many methods have been proposed to solve the above problem, which relies on sampling from the non-stationary model distribution and suffers from high variance or biased estimations. In this paper, we propose $\psi$-MLE, a new training scheme for autoregressive sequence generative models, which is effective and stable when operating at large sample space encountered in text generation. We derive our algorithm from a new perspective of self-augmentation and introduce bias correction with density ratio estimation. Extensive experimental results on synthetic data and real-world text generation tasks demonstrate that our method stably outperforms Maximum Likelihood Estimation and other state-of-the-art sequence generative models in terms of both quality and diversity. Yuxuan Song 0002, Ning Miao, Hao Zhou 0012, Lantao Yu, Mingxuan Wang, Lei Li 0005 |
AISTATS | 5 |
| 2020 | On the Sentence Embeddings from Pre-trained Language ModelsabstractPre-trained contextual representations like BERT have achieved great success in natural language processing.However, the sentence embeddings from the pre-trained language models without fine-tuning have been found to poorly capture semantic meaning of sentences.In this paper, we argue that the semantic information in the BERT embeddings is not fully exploited.We first reveal the theoretical connection between the masked language model pre-training objective and the semantic similarity task theoretically, and then analyze the BERT sentence embeddings empirically.We find that BERT always induces a non-smooth anisotropic semantic space of sentences, which harms its performance of semantic similarity.To address this issue, we propose to transform the anisotropic sentence embedding distribution to a smooth and isotropic Gaussian distribution through normalizing flows that are learned with an unsupervised objective.Experimental results show that our proposed BERT-flow method obtains significant performance gains over the state-of-the-art sentence embeddings on a variety of semantic textual similarity tasks.The code is available at https://github.com/ bohanli/BERT-flow. Hao Zhou 0012, Junxian He, Mingxuan Wang, Yiming Yang 0002, Lei Li 0005 |
EMNLP (1) | 4 |
| 2020 | Pre-training Multilingual Neural Machine Translation by Leveraging Alignment InformationabstractWe investigate the following question for machine translation (MT): can we develop a single universal MT model to serve as the common seed and obtain derivative and improved models on arbitrary language pairs?We propose mRASP, an approach to pre-train a universal multilingual neural machine translation model.Our key idea in mRASP is its novel technique of random aligned substitution, which brings words and phrases with similar meanings across multiple languages closer in the representation space.We pre-train a mRASP model on 32 language pairs jointly with only public datasets.The model is then fine-tuned on downstream language pairs to obtain specialized MT models.We carry out extensive experiments on 42 translation directions across a diverse settings, including low, medium, rich resource, and as well as transferring to exotic language pairs.Experimental results demonstrate that mRASP achieves significant performance improvement compared to directly training on those target pairs.It is the first time to verify that multiple lowresource language pairs can be utilized to improve rich resource MT.Surprisingly, mRASP is even able to improve the translation quality on exotic languages that never occur in the pretraining corpus.Code, data, and pre-trained models are available at https://github. com/linzehui/mRASP. Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou 0012, Lei Li 0005 |
EMNLP (1) | 3 |
| 2019 | Imitation Learning for Non-Autoregressive Neural Machine TranslationabstractNon-autoregressive translation models (NAT) have achieved impressive inference speedup.A potential issue of the existing NAT algorithms, however, is that the decoding is conducted in parallel, without directly considering previous context.In this paper, we propose an imitation learning framework for nonautoregressive machine translation, which still enjoys the fast translation speed but gives comparable translation performance compared to its auto-regressive counterpart.We conduct experiments on the IWSLT16, WMT14 and WMT16 datasets.Our proposed model achieves a significant speedup over the autoregressive models, while keeping the translation quality comparable to the autoregressive models.By sampling sentence length in parallel at inference time, we achieve the performance of 31.85BLEU on WMT16 Ro→En and 30.68 BLEU on IWSLT16 En→De. Bingzhen Wei, Mingxuan Wang, Hao Zhou 0012, Junyang Lin, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Towards Linear Time Neural Machine Translation with Capsule NetworksabstractMingxuan Wang, Jun Xie, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li 0005 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Deep Semantic Role Labeling With Self-AttentionabstractSemantic Role Labeling (SRL) is believed to be a crucial step towards natural language understanding and has been widely studied. Recent years, end-to-end SRL with recurrent neural networks (RNN) has gained increasing attention. However, it remains a major challenge for RNNs to handle structural information and long range dependencies. In this paper, we present a simple and effective architecture for SRL which aims to address these problems. Our model is based on self-attention which can directly capture the relationships between two tokens regardless of their distance. Our single model achieves F1=83.4 on the CoNLL-2005 shared task dataset and F1=82.7 on the CoNLL-2012 shared task dataset, which outperforms the previous state-of-the-art results by 1.8 and 1.0 F1 score respectively. Besides, our model is computationally efficient, and the parsing speed is 50K tokens per second on a single Titan X GPU. Zhixing Tan, Mingxuan Wang, Yidong Chen 0001, Xiaodong Shi |
AAAI | 2 |
| 2018 | Neural Machine Translation with Decoding History Enhanced AttentionabstractNeural machine translation with source-side attention have achieved remarkable performance. however, there has been little work exploring to attend to the target-side which can potentially enhance the memory capbility of NMT. We reformulate a Decoding History Enhanced Attention mechanism (DHEA) to render NMT model better at selecting both source-side and target-side information. DHA enables dynamic control of the ratios at which source and target contexts contribute to the generation of target words, offering a way to weakly induce structure relations among both source and target tokens. It also allows training errors to be directly back-propagated through short-cut connections and effectively alleviates the gradient vanishing problem. The empirical study on Chinese-English translation shows that our model with proper configuration can improve by 0:9 BLEU upon Transformer and the best reported results in the dataset. On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Chao Bian 0005 |
COLING | 1 |
| 2018 | A Hierarchy-to-Sequence Attentional Neural Machine Translation ModelabstractAlthough sequence-to-sequence attentional neural machine translation (NMT) has achieved great progress recently, it is confronted with two challenges: learning optimal model parameters for long parallel sentences and well exploiting different scopes of contexts. In this paper, partially inspired by the idea of segmenting a long sentence into short clauses, each of which can be easily translated by NMT, we propose a hierarchy-to-sequence attentional NMT model to handle these two challenges. Our encoder takes the segmented clause sequence as input and explores a hierarchical neural network structure to model words, clauses, and sentences at different levels, particularly with two layers of recurrent neural networks modeling semantic compositionality at the word and clause level. Correspondingly, the decoder sequentially translates segmented clauses and simultaneously applies two types of attention models to capture contexts of interclause and intraclause for translation prediction. In this way, we can not only improve parameter learning, but also well explore different scopes of contexts for translation. Experimental results on Chinese-English and English-German translation demonstrate the superiorities of the proposed model over the conventional NMT model. Jinsong Su, Jiali Zeng, Deyi Xiong, Yang Liu 0005, Mingxuan Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Deep Neural Machine Translation with Linear Associative UnitabstractDeep Neural Networks (DNNs) have provably enhanced the state-of-the-art Neural Machine Translation (NMT) with their capability in modeling complex functions and capturing complex linguistic structures.However NMT systems with deep architecture in their encoder or decoder RNNs often suffer from severe gradient diffusion due to the non-linear recurrent activations, which often make the optimization much more difficult.To address this problem we propose novel linear associative units (LAU) to reduce the gradient propagation length inside the recurrent unit.Different from conventional approaches (LSTM unit and GRU), LAUs utilizes linear associative connections between input and output of the recurrent unit, which allows unimpeded information flow through both space and time direction.The model is quite simple, but it is surprisingly effective.Our empirical study on Chinese-English translation shows that our model with proper configuration can improve by 11.7 BLEU upon Groundhog and the best reported results in the same setting.On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhengdong Lu, Jie Zhou 0016, Qun Liu 0001 |
ACL (1) | 1 |
| 2017 | Incorporating Word Reordering Knowledge into Attention-based Neural Machine TranslationabstractThis paper proposes three distortion models to explicitly incorporate the word reordering knowledge into attention-based Neural Machine Translation (NMT) for further improving translation performance.Our proposed models enable attention mechanism to attend to source words regarding both the semantic requirement and the word reordering penalty.Experiments on Chinese-English translation show that the approaches can improve word alignment quality and achieve significant translation improvements over a basic attention-based N-MT by large margins.Compared with previous works on identical corpora, our system achieves the state-of-the-art performance on translation quality. Jinchao Zhang 0001, Mingxuan Wang, Qun Liu 0001, Jie Zhou 0016 |
ACL (1) | 2 |
| 2016 | Memory-enhanced Decoder for Neural Machine TranslationabstractWe propose to enhance the RNN decoder in a neural machine translator (NMT) with external memory, as a natural but powerful extension to the state in the decoding RNN.This memory-enhanced RNN decoder is called MEMDEC.At each time during decoding, MEMDEC will read from this memory and write to this memory once, both with content-based addressing.Unlike the unbounded memory in previous work (Bahdanau et al., 2014) to store the representation of source sentence, the memory in MEMDEC is a matrix with predetermined size designed to better capture the information important for the decoding process at each time step.Our empirical study on Chinese-English translation shows that it can improve by 4.8 BLEU upon Groundhog and 5.3 BLEU upon on Moses, yielding the best performance achieved with the same training set. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
EMNLP | 1 |
| 2015 | Encoding Source Language with Convolutional Neural Network for Machine TranslationabstractFandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Fandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 3 |
| 2015 | genCNN: A Convolutional Architecture for Word Sequence PredictionabstractMingxuan Wang, Zhengdong Lu, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 1 |
| 2015 | Syntax-Based Deep Matching of Short Texts
Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
IJCAI | 1 |
| 2011 | A class of fast quaternion valued variable stepsize stochastic gradient learning algorithms for vector sensor processesabstractWe introduce a class of gradient adaptive stepsize algorithms for quaternion valued adaptive filtering based on three- and four-dimensional vector sensors. This equips the recently introduced quaternion least mean square (QLMS) algorithm with enhanced tracking ability and enables it to be more responsive to dynamically changing environments, while maintaining its desired characteristics of catering for large dynamical differences and coupling between signal components. For generality, the analysis is performed for the widely linear signal model, which by virtue of accounting for signal noncircularity, is optimal in the mean squared error (MSE) sense for both second order circular (proper) and noncircular (improper) processes. The widely linear QLMS (WL-QLMS) employing the proposed adaptive stepsize modifications is shown to provide enhanced performance for both synthetic and real world quaternion valued signals. Simulations include signals with drastically different component dynamics, such as four dimensional quaternion comprising three dimensional turbulent wind and air temperature for renewable energy applications. Mingxuan Wang, Clive Cheong Took, Danilo P. Mandic |
IJCNN | 1 |