EDBT 2026 Demo / reviewers in the wild / expert
Hua Wu 0003
dblp:27/6045-3
· DBLP profile ↗
151ranked-venue papers
10as first author
74since 2021 · last 2026
0000-0001-8254-1561ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 136 · 10 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented GenerationabstractWith the rapid advancement of large language models (LLMs), retrieval-augmented generation (RAG) has emerged as a critical approach to supplement the inherent knowledge limitations of LLMs. However, due to the typically large volume of retrieved information, RAG tends to operate with long context lengths. From the perspective of entropy engineering, we identify unconstrained entropy growth and attention dilution due to long retrieval context as significant factors affecting RAG performance. In this paper, we propose the balanced entropy-engineered RAG (BEE-RAG) framework, which improves the adaptability of RAG systems to varying context lengths through the principle of entropy invariance. By leveraging balanced context entropy to reformulate attention dynamics, BEE-RAG separates attention sensitivity from context length, ensuring a stable entropy level. Building upon this, we introduce a zero-shot inference strategy for multi-importance estimation and a parameter-efficient adaptive fine-tuning mechanism to obtain the optimal balancing factor for different settings. Extensive experiments across multiple RAG tasks demonstrate the effectiveness of BEE-RAG. Yuhao Wang 0007, Ruiyang Ren, Yucheng Wang 0006, Jing Liu 0022, Wayne Xin Zhao, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 6 |
| 2026 | Uncertainty-Aware Routing for Principled Alignment with MoE DynamicsabstractYilong Chen, Junyuan Shang, Yuchen Feng, Zhenyu Zhang, Naibin Gu, Ziqi Wang, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junyuan Shang, Zhenyu Zhang 0006, Naibin Gu, Tingwen Liu, Shuohuan Wang, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 10 |
| 2026 | AttnPO: Attention-Guided Process Supervision for Efficient ReasoningabstractShuaiyi Nie, Dingsiyu, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Weichong Yin, Yu Sun, Hua Wu, Tingwen Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shuaiyi Nie, Siyu Ding, Wenyuan Zhang 0002, Linhao Yu, Tianmeng Yang, Yao Chen 0009, Weichong Yin, Hua Wu 0003, Tingwen Liu |
ACL (1) | 9 |
| 2026 | Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language ModelsabstractShaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang, Hua Wu, Haifeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang 0001, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 7 |
| 2026 | Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented GenerationabstractLong-form question answering (LFQA) requires open-ended long-form responses that synthesize coherent, factually grounded content from multi-source evidence.This makes reinforcement learning (RL) reward design critical.The reward must be verifiable for faithful grounding and stable optimization.However, many standard rewards assume a unique target with an exact-match notion of correctness, which fits short-form QA and math but breaks in LFQA.As a result, current RAG systems still lack verifiable reward mechanisms, yielding unstable feedback signals and suboptimal optimization outcomes.We propose RioRAG, a framework for reinforced verifiable informativeness optimization.First, it defines informativeness as a measurable and externally verifiable objective for RL.Second, RioRAG uses nugget-centric verification with cross-source checks to enable self-evolution of smaller LLMs and to provide denser, actiondiscriminative rewards that mitigate reward sparsity and stabilize optimization.This formulation avoids handcrafted supervision for the policy model and strong teacher-model distillation, relying instead on externally verifiable feedback.Experiments on LongFact and RAGChecker show that RioRAG achieves higher factual recall and faithfulness, establishing verifiable reward modeling as a foundation for trustworthy long-form RAG.Our codes are available at https://github.com/RUCAIBox/ RioRAG. Yuhao Wang 0007, Ruiyang Ren, Yucheng Wang 0006, Wayne Xin Zhao, Jing Liu 0022, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 6 |
| 2025 | Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal ThinkingabstractYilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junyuan Shang, Zhenyu Zhang 0006, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 9 |
| 2025 | BeamLoRA: Beam-Constraint Low-Rank AdaptationabstractDue to the demand for efficient fine-tuning of large language models, Low-Rank Adaptation (LoRA) has been widely adopted as one of the most effective parameter-efficient fine-tuning methods. Nevertheless, while LoRA improves efficiency, there remains room for improvement in accuracy. Herein, we adopt a novel perspective to assess the characteristics of LoRA ranks. The results reveal that different ranks within the LoRA modules not only exhibit varying levels of importance but also evolve dynamically throughout the fine-tuning process, which may limit the performance of LoRA. Based on these findings, we propose BeamLoRA, which conceptualizes each LoRA module as a beam where each rank naturally corresponds to a potential sub-solution, and the fine-tuning process becomes a search for the optimal sub-solution combination. BeamLoRA dynamically eliminates underperforming sub-solutions while expanding the parameter space for promising ones, enhancing performance with a fixed rank. Extensive experiments across three base models and 12 datasets spanning math reasoning, code generation, and commonsense reasoning demonstrate that BeamLoRA consistently enhances the performance of LoRA, surpassing the other baseline methods. Naibin Gu, Zhenyu Zhang 0006, Xiyu Liu 0003, Peng Fu 0008, Zheng Lin 0001, Shuohuan Wang, Hua Wu 0003, Weiping Wang 0005, Haifeng Wang 0001 |
ACL (1) | 8 |
| 2025 | HFT: Half Fine-Tuning for Large Language ModelsabstractLarge language models (LLMs) with one or more fine-tuning phases have become necessary to unlock various capabilities, enabling LLMs to follow natural language instructions and align with human preferences. However, it carries the risk of catastrophic forgetting during sequential training, the parametric knowledge or the ability learned in previous stages may be overwhelmed by incoming training data. This paper finds that LLMs can restore some original knowledge by regularly resetting partial parameters. Inspired by this, we introduce Half Fine-Tuning (HFT) for LLMs, as a substitute for full fine-tuning (FFT), to mitigate the forgetting issues, where half of the parameters are selected to learn new tasks. In contrast, the other half are frozen to retain previous knowledge. We provide a feasibility analysis from the optimization perspective and interpret the parameter selection operation as a regularization term. HFT could be seamlessly integrated into existing fine-tuning frameworks without changing the model architecture. Extensive experiments and analysis on supervised fine-tuning, direct preference optimization, and continual learning consistently demonstrate the effectiveness, robustness, and efficiency of HFT. Compared with FFT, HFT not only significantly alleviates the forgetting problem, but also achieves the best performance in a series of downstream benchmarks, with an approximately 30% reduction in training time. Tingfeng Hui, Zhenyu Zhang 0006, Shuohuan Wang, Weiran Xu, Yu Sun 0029, Hua Wu 0003 |
ACL (1) | 6 |
| 2025 | Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter MergingabstractMixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks.However, existing methods transforming LLMs from dense to MoE face significant data requirements and typically rely on large-scale post-training.In this paper, we propose Upcycling Instruction Tuning (UpIT), a data-efficient approach for tuning a dense pre-trained model into a MoE instruction model.Specifically, we first point out that intermediate checkpoints during instruction tuning of the dense model are naturally suitable for specialized experts, and then propose an expert expansion stage to flexibly achieve models with flexible numbers of experts, where genetic algorithm and parameter merging are introduced to ensure sufficient diversity of new extended experts.To ensure that each specialized expert in the MoE model works as expected, we select a small amount of seed data that each expert excels to preoptimize the router.Extensive experiments with various data scales and upcycling settings demonstrate the outstanding performance and data efficiency of UpIT, as well as stable improvement in expert or data scaling.Further analysis reveals the importance of ensuring expert diversity in upcycling. Tingfeng Hui, Zhenyu Zhang 0006, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Sen Su |
ACL (1) | 5 |
| 2025 | Curiosity-Driven Reinforcement Learning from Human FeedbackabstractReinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity and alignment quality remains a significant challenge. Drawing inspiration from curiosity-driven exploration in reinforcement learning, we introduce curiosity-driven RLHF (CD-RLHF), a framework that incorporates intrinsic rewards for novel states, alongside traditional sparse extrinsic rewards, to optimize both output diversity and alignment quality. We demonstrate the effectiveness of CD-RLHF through extensive experiments on a range of tasks, including text summarization and instruction following. Our approach achieves significant gains in diversity on multiple diversity-oriented metrics while maintaining alignment with human preferences comparable to standard RLHF. We will make our code publicly available. Yekun Chai, Shuohuan Wang, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 5 |
| 2025 | Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval AugmentationabstractLarge language models (LLMs) have shown impressive prowess in solving a wide range of tasks with world knowledge. However, it remains unclear how well LLMs are able to perceive their factual knowledge boundaries, particularly under retrieval augmentation settings. In this study, we present the first analysis on the factual knowledge boundaries of LLMs and how retrieval augmentation affects LLMs on open-domain question answering (QA), with a bunch of important findings. Specifically, we focus on three research questions and analyze them by examining QA, priori judgement and posteriori judgement capabilities of LLMs. We show evidence that LLMs possess unwavering confidence in their knowledge and cannot handle the conflict between internal and external knowledge well. Furthermore, retrieval augmentation proves to be an effective approach in enhancing LLMs’ awareness of knowledge boundaries. We further conduct thorough experiments to examine how different factors affect LLMs and propose a simple method to dynamically utilize supporting documents with our judgement strategy. Additionally, we find that the relevance between the supporting documents and the questions significantly impacts LLMs’ QA and judgemental capabilities. Ruiyang Ren, Yuhao Wang 0007, Yingqi Qu, Wayne Xin Zhao, Jing Liu 0022, Hua Wu 0003, Ji-Rong Wen, Haifeng Wang 0001 |
COLING | 6 |
| 2025 | AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentabstractMultilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities.However, their performance and cross-lingual alignment often lag for non-dominant languages.A common solution is to fine-tune LLMs on largescale and more balanced multilingual corpora, but such approaches often lead to imprecise alignment and suboptimal knowledge transfer, struggling with limited improvements across languages.In this paper, we propose AlignX to bridge the multilingual performance gap, which is a two-stage representationlevel framework for enhancing multilingual performance of pre-trained LLMs.In the first stage, we align multilingual representations with multilingual semantic alignment and language feature integration.In the second stage, we stimulate the multilingual capability of LLMs via multilingual instruction fine-tuning.Experimental results on several pre-trained LLMs demonstrate that our approach enhances LLMs' multilingual general and cross-lingual generation capability.Further analysis indicates that AlignX brings the multilingual representations closer and improves the cross-lingual alignment.1 Mengyu Bu, Shaolei Zhang 0001, Zhongjun He, Hua Wu 0003, Yang Feng 0004 |
EMNLP | 4 |
| 2025 | Weights-Rotated Preference Optimization for Large Language ModelsabstractChenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin, Siyuan Chen, Weichong Yin, Hua Wu, Weiping Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin 0001, Weichong Yin, Hua Wu 0003, Weiping Wang 0005 |
EMNLP | 8 |
| 2025 | MA-RLHF: Reinforcement Learning from Human Feedback with Macro ActionsabstractReinforcement learning from human feedback (RLHF) has demonstrated effectiveness in aligning large language models (LLMs) with human preferences. However, token-level RLHF suffers from the credit assignment problem over long sequences, where delayed rewards make it challenging for the model to discern which actions contributed to preferred outcomes. This hinders learning efficiency and slows convergence.In this paper, we propose MA-RLHF, a simple yet effective RLHF framework that incorporates macro actions --- sequences of tokens or higher-level language constructs --- into the learning process. By operating at higher level of abstraction, our approach reduces the temporal distance between actions and rewards, facilitating faster and more accurate credit assignment. This results in more stable policy gradient estimates and enhances learning efficiency within each episode, all without increasing computational complexity during training or inference. We validate our approach through extensive experiments across various model sizes and tasks, including text summarization, dialogue generation, question answering, and program synthesis. Our method achieves substantial performance improvements over standard RLHF, with performance gains of up to 30\% in text summarization and code generation, 18\% in dialogue, and 8\% in question answering tasks. Notably, our approach reaches parity with vanilla RLHF $1.7 \sim 2$ times faster in terms of training time and continues to outperform it with further training. We make our code and data publicly available at \url{https://github.com/ernie-research/MA-RLHF}. Yekun Chai, Huang Fang, Shuohuan Wang, Hua Wu 0003 |
ICLR | 6 |
| 2025 | Mixture of Hidden-Dimensions: Not All Hidden-States' Dimensions are Needed in TransformerabstractTransformer models encounter inefficiency when scaling hidden dimensions due to the uniform expansion of parameters. When delving into the sparsity of hidden dimensions, we observe that only a small subset of dimensions are highly activated, where some dimensions are commonly activated across tokens, and some others uniquely activated for individual tokens. To leverage this, we propose MoHD (Mixture of Hidden Dimensions), a sparse architecture that combines shared sub-dimensions for common features and dynamically routes specialized sub-dimensions per token. To address the potential information loss from sparsity, we introduce activation scaling and group fusion mechanisms. MoHD efficiently expands hidden dimensions with minimal computational increases, outperforming vanilla Transformers in both parameter efficiency and task performance across 10 NLP tasks. MoHD achieves 1.7% higher performance with 50% fewer activatied parameters and 3.7% higher performance with 3$\times$ total parameters expansion at constant activated parameters cost. MoHD offers a new perspective for scaling the model, showcasing the potential of hidden dimension sparsity. Junyuan Shang, Zhenyu Zhang 0006, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001 |
ICML | 8 |
| 2025 | Residual Stream Analysis of Overfitting And Structural DisruptionsabstractEnsuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets—where unsafe prompts are paired with standard refusal templates—often leads to \emph{false refusals}, in which benign queries are declined. We first quantify this effect, showing that safety data exhibits substantially lower token entropy ($H_{1}\approx9.18$) and 2-gram diversity ($\approx$ 0.048) compared to general instruction data ($H_{1}\approx12.05$, 2-gram$\approx$0.205). To uncover the root cause, we introduce \emph{FlowLens}, a stable PCA-based tool for residual-stream geometry analysis, and reveal that higher proportions of safety examples concentrate variance along a few components, reducing representational smoothness and driving false refusals (false refusal rate rises from 63\% to 84\% as safety data increases from 0\% to 40\%). Guided by these insights, we propose \emph{Variance Concentration Loss} (VCL), an auxiliary regularizer that penalizes excessive variance concentration in mid-layer residuals. Empirical results demonstrate that VCL reduces false refusals by over 35 percentage points while maintaining or improving performance on general benchmarks such as MMLU and GSM8K. Wenquan Wu, Hua Wu 0003, Sen Su |
NeurIPS | 4 |
| 2025 | Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented GenerationabstractConsidering the inherent limitations of parametric knowledge in large language models (LLMs), retrieval-augmented generation (RAG) is widely employed to expand their knowledge scope. Since RAG has shown promise in knowledge-intensive tasks like open-domain question answering, its broader application to complex tasks and intelligent assistants has further advanced its utility. Despite this progress, the underlying knowledge utilization mechanisms of LLM-based RAG remain underexplored. In this paper, we present a systematic investigation of the intrinsic mechanisms by which LLMs integrate internal (parametric) and external (retrieved) knowledge in RAG scenarios. Specially, we employ knowledge stream analysis at the macroscopic level, and investigate the function of individual modules at the microscopic level. Drawing on knowledge streaming analyses, we decompose the knowledge utilization process into four distinct stages within LLM layers: knowledge refinement, knowledge elicitation, knowledge expression, and knowledge contestation. We further demonstrate that the relevance of passages guides the streaming of knowledge through these stages. At the module level, we introduce a new method, knowledge activation probability entropy (KAPE) for neuron identification associated with either internal or external knowledge. By selectively deactivating these neurons, we achieve targeted shifts in the LLM's reliance on one knowledge source over the other. Moreover, we discern complementary roles for multi-head attention and multi-layer perceptron layers during knowledge formation. These insights offer a foundation for improving interpretability and reliability in retrieval-augmented LLMs, paving the way for more robust and transparent generative solutions in knowledge-intensive domains. Yuhao Wang 0007, Ruiyang Ren, Yucheng Wang 0006, Wayne Xin Zhao, Jing Liu 0022, Hua Wu 0003, Haifeng Wang 0001 |
SIGIR | 6 |
| 2025 | A simple yet effective self-debiasing framework for transformer models
Suhang Wu, Jinsong Su, Hua Wu 0003 |
Artif. Intell. | 6 |
| 2025 | Towards few-shot mixed-type dialogue generation
Zeming Liu, Haifeng Wang 0001, Zeyang Lei, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
Sci. China Inf. Sci. | 5 |
| 2024 | LEMON: Reviving Stronger and Smaller LMs from Larger LMs with Linear Parameter FusionabstractYilong Chen, Junyuan Shang, Zhenyu Zhang, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junyuan Shang, Zhenyu Zhang 0006, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003 |
ACL (1) | 8 |
| 2024 | NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference TimeabstractYilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang, Tingwen Liu, Shuohuan Wang, Yu Sun, Dianhai Yu, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 10 |
| 2024 | QDMR-based Planning-and-Solving Prompting for Complex Reasoning TasksabstractChain-of-Thought prompting has improved reasoning capability of large language models (LLM). However, it still is challenging to guarantee the effectiveness and stability for questions requiring complicated reasoning. Recently, Plan-and-Solve prompting enhances the reasoning capability for complex questions by planning the solution steps firstly and then solving them step by step, but it suffers the difficulty to represent and execute the problem-solving logic of complex questions. To deal with these challenges, in this work, we propose a novel Plan-and-Solve prompting method based on Question Decomposition Meaning Representation (QDMR). Specifically, this method first allows the LLM to generate a QDMR graph to represent the problem-solving logic, which is a directed acyclic graph composed of sub-questions. Then, the LLM generates a specific solving process based on the QDMR graph. When solving each sub-question, it can locate the preceding sub-questions and their answers according to the QDMR graph, and then utilize this information for solution. Compared with existing Plan-and-Solve prompting techniques, our method can not only represent the problem-solving logic of complicated questions more accurately with the aid of QDMR graph, but also deliver the dependence information accurately for different solution steps according to the QDMR graph. In addition, with the supervised fine-tuning on the Allen Institute dataset, the decomposing capability of LLM for complicated questions can be considerably enhanced. Extensive experiments show that our method has achieve a great significance in arithmetic reasoning and commonsense reasoning task by comparing the classical Chain-of-Thought prompting and Plan-and-Solve prompting techniques, and the improvements achieved are even greater for problems with more reasoning steps. Qiaoqiao She, Wenbin Jiang 0002, Hua Wu 0003, Tong Xu 0001, Feng Wu 0001 |
LREC/COLING | 4 |
| 2024 | On Training Data Influence of GPT ModelsabstractAmidst the rapid advancements in generative language models, the investigation of how training data shapes the performance of GPT models is still emerging.This paper presents GPTfluence, a novel approach that leverages a featurized simulation to assess the impact of training examples on the training dynamics of GPT models.Our approach not only traces the influence of individual training instances on performance trajectories, such as loss and other key metrics, on targeted test points but also enables a comprehensive comparison with existing methods across various training scenarios in GPT models, ranging from 14 million to 2.8 billion parameters, across a range of downstream tasks.Contrary to earlier methods that struggle with generalization to new data, GPTfluence introduces a parameterized simulation of training dynamics, demonstrating robust generalization capabilities to unseen training data.This adaptability is evident across both fine-tuning and instruction-tuning scenarios, spanning tasks in natural language understanding and generation.We make our Yekun Chai, Qingyi Liu, Shuohuan Wang, Yu Sun 0004, Qiwei Peng 0002, Hua Wu 0003 |
EMNLP | 6 |
| 2024 | Autoregressive Pre-Training on Pixels and TextsabstractThe integration of visual and textual information represents a promising direction in the advancement of language models.In this paper, we explore the dual modality of language-both visual and textual-within an autoregressive framework, pre-trained on both document images and texts.Our method employs a multimodal training strategy, utilizing visual data through next patch prediction with a regression head and/or textual data through next token prediction with a classification head.We focus on understanding the interaction between these two modalities and their combined impact on model performance.Our extensive evaluation across a wide range of benchmarks shows that incorporating both visual and textual data significantly improves the performance of pixel-based language models.Remarkably, we find that a unidirectional pixelbased model trained solely on visual data can achieve comparable results to state-of-the-art bidirectional models on several language understanding tasks.This work uncovers the untapped potential of integrating visual and textual modalities for more effective language modeling.We release our code, data, and model checkpoints at Yekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang, Hua Wu 0003 |
EMNLP | 6 |
| 2024 | Tool-Augmented Reward ModelingabstractReward modeling (*a.k.a.*, preference modeling) is instrumental for aligning large language models with human preferences, particularly within the context of reinforcement learning from human feedback (RLHF). While conventional reward models (RMs) have exhibited remarkable scalability, they oft struggle with fundamental functionality such as arithmetic computation, code execution, and factual lookup. In this paper, we propose a tool-augmented preference modeling approach, named Themis, to address these limitations by empowering RMs with access to external environments, including calculators and search engines. This approach not only fosters synergy between tool utilization and reward grading but also enhances interpretive capacity and scoring reliability. Our study delves into the integration of external tools into RMs, enabling them to interact with diverse external sources and construct task-specific tool engagement and reasoning traces in an autoregressive manner. We validate our approach across a wide range of domains, incorporating seven distinct external tools. Our experimental results demonstrate a noteworthy overall improvement of 17.7% across eight tasks in preference ranking. Furthermore, our approach outperforms Gopher 280B by 7.3% on TruthfulQA task in zero-shot evaluation. In human evaluations, RLHF trained with Themis attains an average win rate of 32% when compared to baselines across four distinct tasks. Additionally, we provide a comprehensive collection of tool-related RM datasets, incorporating data from seven distinct tool APIs, totaling 15,000 instances. We have made the code, data, and model checkpoints publicly available to facilitate and inspire further research advancements (https://github.com/ernie-research/Tool-Augmented-Reward-Model). Lei Li 0040, Yekun Chai, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Ningyu Zhang 0001, Hua Wu 0003 |
ICLR | 7 |
| 2024 | An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text TranslationabstractPengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Pengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
NAACL-HLT | 4 |
| 2024 | Learning to Select External Knowledge With Multi-Scale Negative SamplingabstractThe Track-1 of DSTC9 aims to effectively answer user requests or questions during task-oriented dialogues, which are out of the scope of APIs/DB. By leveraging external knowledge resources, relevant information can be retrieved and encoded into the response generation for these out-of-API-coverage queries. In this work, we have explored several advanced techniques to enhance the utilization of external knowledge and boost the quality of response generation, includingschema guided knowledge decision,negatives enhanced knowledge selection, andknowledge grounded response generation. To evaluate the performance of our proposed method, comprehensive experiments have been carried out on the publicly available dataset. Our approach was ranked as the best in human evaluation of DSTC9 Track-1. Huang He, Hua Lu 0014, Siqi Bao, Fan Wang 0021, Hua Wu 0003, Zhengyu Niu, Haifeng Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Universal Information Extraction as Unified Semantic MatchingabstractThe challenge of information extraction (IE) lies in the diversity of label schemas and the heterogeneity of structures. Traditional methods require task-specific model design and rely heavily on expensive supervision, making them difficult to generalize to new schemas. In this paper, we decouple IE into two basic abilities, structuring and conceptualizing, which are shared by different tasks and schemas. Based on this paradigm, we propose to universally model various IE tasks with Unified Semantic Matching (USM) framework, which introduces three unified token linking operations to model the abilities of structuring and conceptualizing. In this way, USM can jointly encode schema and input text, uniformly extract substructures in parallel, and controllably decode target structures on demand. Empirical evaluation on 4 IE tasks shows that the proposed method achieves state-of-the-art performance under the supervised experiments and shows strong generalization ability in zero/few-shot transfer settings. Jie Lou, Yaojie Lu 0001, Dai Dai, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
AAAI | 8 |
| 2023 | Towards Boosting the Open-Domain Chatbot with Human FeedbackabstractMany open-domain dialogue models pretrained with social media comments can generate coherent replies but have difficulties producing engaging responses.This phenomenon might mainly result from the deficiency of annotated human-human conversations and the misalignment with human preference.In this paper, we propose a novel and efficient framework Diamante to boost the open-domain chatbot, where two kinds of human feedback (including explicit demonstration and implicit preference) are collected and leveraged.By asking annotators to select or amend the modelgenerated candidate responses, Diamante efficiently collects the human demonstrated responses and constructs a Chinese chit-chat dataset.To enhance the alignment with human preference, Diamante leverages the implicit preference in the data collection process and introduces the generation-evaluation joint training.Comprehensive experiments indicate that the Diamante dataset and joint training paradigm can significantly boost the performance of pre-trained dialogue models.The overall engagingness of the previous state-ofthe-art model has been improved remarkably by 50% in Chinese open-domain conversations. Hua Lu 0014, Siqi Bao, Huang He, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 5 |
| 2023 | Query Enhanced Knowledge-Intensive Conversation via Unsupervised Joint ModelingabstractIn this paper, we propose an unsupervised query enhanced approach for knowledgeintensive conversations, namely QKConv.There are three modules in QKConv: a query generator, an off-the-shelf knowledge selector, and a response generator.QKConv is optimized through joint training, which produces the response by exploring multiple candidate queries and leveraging corresponding selected knowledge.The joint training solely relies on the dialogue context and target response, getting exempt from extra query annotations or knowledge provenances.To evaluate the effectiveness of the proposed QKConv, we conduct experiments on three representative knowledgeintensive conversation datasets: conversational question-answering, task-oriented dialogue, and knowledge-grounded conversation.Experimental results reveal that QKConv performs better than all unsupervised methods across three datasets and achieves competitive performance compared to supervised methods. Mingzhu Cai, Siqi Bao, Xin Tian 0011, Huang He, Fan Wang 0021, Hua Wu 0003 |
ACL (1) | 6 |
| 2023 | Learning In-context Learning for Named Entity RecognitionabstractJiawei Chen, Yaojie Lu, Hongyu Lin, Jie Lou, Wei Jia, Dai Dai, Hua Wu, Boxi Cao, Xianpei Han, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiawei Chen 0011, Yaojie Lu 0001, Jie Lou, Dai Dai, Hua Wu 0003, Boxi Cao, Xianpei Han, Le Sun 0001 |
ACL (1) | 7 |
| 2023 | TOME: A Two-stage Approach for Model-based RetrievalabstractRecently, model-based retrieval has emerged as a new paradigm in text retrieval that discards the index in the traditional retrieval model and instead memorizes the candidate corpora using model parameters.This design employs a sequence-to-sequence paradigm to generate document identifiers, which enables the complete capture of the relevance between queries and documents and simplifies the classic indexretrieval-rerank pipeline.Despite its attractive qualities, there remain several major challenges in model-based retrieval, including the discrepancy between pre-training and fine-tuning, and the discrepancy between training and inference.To deal with the above challenges, we propose a novel two-stage model-based retrieval approach called TOME, which makes two major technical contributions, including the utilization of tokenized URLs as identifiers and the design of a two-stage generation architecture.We also propose a number of training strategies to deal with the training difficulty as the corpus size increases.Extensive experiments and analysis on MS MARCO and Natural Questions demonstrate the effectiveness of our proposed approach, and we investigate the scaling laws of TOME by examining various influencing factors. Ruiyang Ren, Wayne Xin Zhao, Jing Liu 0022, Hua Wu 0003, Ji-Rong Wen, Haifeng Wang 0001 |
ACL (1) | 4 |
| 2023 | ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-ExpertsabstractRecent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of image fidelity and text relevancy. In this paper, we propose ERNIE-ViLG 2.0, a large-scale Chinese text-to-image diffusion model, to progressively upgrade the quality of generated images by: (1) incorporating fine-grained textual and visual knowledge of key elements in the scene, and (2) utilizing different denoising experts at different denoising stages. With the proposed mechanisms, ERNIE-ViLG 2.01not only achieves a new state-of-the-art on MS-COCO with zero-shot FID-30k score of 6.75, but also significantly outperforms recent models in terms of image fidelity and image-text alignment, with side-by-side human evaluation on the bilingual prompt set ViLG-300. Zhida Feng, Zhenyu Zhang 0006, Yewei Fang, Lanxin Li, Xuyi Chen, Jiaxiang Liu 0004, Weichong Yin, Shikun Feng, Yu Sun 0004, Li Chen 0011, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
CVPR | 14 |
| 2023 | IBADR: an Iterative Bias-Aware Dataset Refinement Framework for Debiasing NLU modelsabstractAs commonly-used methods for debiasing natural language understanding (NLU) models, dataset refinement approaches heavily rely on manual data analysis, and thus maybe unable to cover all the potential biased features.In this paper, we propose IBADR, an Iterative Bias-Aware Dataset Refinement framework, which debiases NLU models without predefining biased features.We maintain an iteratively expanded sample pool.Specifically, at each iteration, we first train a shallow model to quantify the bias degree of samples in the pool.Then, we pair each sample with a bias indicator representing its bias degree, and use these extended samples to train a sample generator.In this way, this generator can effectively learn the correspondence relationship between bias indicators and samples.Furthermore, we employ the generator to produce pseudo samples with fewer biased features by feeding specific bias indicators.Finally, we incorporate the generated pseudo samples into the pool.Experimental results and in-depth analyses on two NLU tasks show that IBADR not only significantly outperforms existing dataset refinement approaches, achieving SOTA, but also is compatible with model-centric methods. 1 Yaoxiang Wang, Jinsong Su, Hua Wu 0003 |
EMNLP | 6 |
| 2023 | Less Learn Shortcut: Analyzing and Mitigating Learning of Spurious Feature-Label CorrelationabstractRecent research has revealed that deep neural networks often take dataset biases as a shortcut to make decisions rather than understand tasks, leading to failures in real-world applications. In this study, we focus on the spurious correlation between word features and labels that models learn from the biased data distribution of training data. In particular, we define the word highly co-occurring with a specific label as biased word, and the example containing biased word as biased example. Our analysis shows that biased examples are easier for models to learn, while at the time of prediction, biased words make a significantly higher contribution to the models' predictions, and models tend to assign predicted labels over-relying on the spurious correlation between words and labels. To mitigate models' over-reliance on the shortcut (i.e. spurious correlation), we propose a training strategy Less-Learn-Shortcut (LLS): our strategy quantifies the biased degree of the biased examples and down-weights them accordingly. Experimental results on Question Matching, Natural Language Inference and Sentiment Analysis tasks show that LLS is a task-agnostic strategy and can improve the model performance on adversarial data while maintaining good performance on in-domain data. Yanrui Du, Jing Yan 0004, Jing Liu 0022, Sendong Zhao, Qiaoqiao She, Hua Wu 0003, Haifeng Wang 0001, Bing Qin 0001 |
IJCAI | 7 |
| 2023 | SeSQL: A High-Quality Large-Scale Session-Level Chinese Text-to-SQL Dataset
Saihao Huang, Zhenghua Li, Chenhui Dou, Fukang Yan, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005 |
NLPCC (1) | 8 |
| 2023 | Controllable Dialogue Generation With Disentangled Multi-Grained Style Specification and Attribute Consistency RewardabstractControllable text generation is an appealing but challenging task, which allows users to specify particular attributes of the generated outputs. In this paper, we propose a controllable dialogue generation model to steer response generation under multi-attribute constraints. Specifically, we define and categorize the commonly-used control attributes into global and local ones, which possess different granularities of effects on response generation. Then, we significantly extend the conventional seq2seq framework by introducing a novel two-stage decoder, which first uses amulti-grainedstyle specification layerto impose the stylistic constraints and determine word-level control states of responses based on the attributes, and then employs aresponse generation layerto generate final responses maintaining both semantic relevancy to the contexts and fidelity to the attributes. Furthermore, we train our model with an attribute consistency reward to promote response control with explicit supervision signals. Extensive experiments and in-depth analyses on two datasets indicate that our model can significantly outperform competitive baselines in terms of response quality, content diversity and controllability. Hou Pong Chan, Xinyan Xiao, Jinsong Su, Hua Wu 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | Graph-Grounded Goal Planning for Conversational RecommendationabstractConversational recommendation casts the recommendation problem as a dialog-based interactive task, which could acquire user interest more efficiently and effectively by allowing users to express what they like. In this work, we move a step towards a new conversational recommendation task that is more suitable for real-world applications. In this task, the recommender proactively and naturally lead a dialog from non-recommendation content to approach an item being of interest to users, and allow users to ask questions for better support of user decisions. The challenge of this task lies in how to effectively control the dialog flow to complete the recommendation while appropriately responding to user utterances. To address this challenge, we first construct a Chinese recommendation dialog dataset DuRecDial. We then propose a two-stage Multi-Goal driven Conversation Generation framework, MGCG. In particular, the goal planning module leverages the global graph structure information and local goal-sequence information to effectively control the dialog flow step by step. The goal-guided responding module can produce an in-depth dialog about each goal by fully exploiting hierarchical goal information for response retrieval or generation. Results on DuRecDial demonstrate that MGCG can lead the dialog more proactively and naturally, and complete the recommendation task more effectively. Zeming Liu, Hao Liu 0026, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | RLCharge: Imitative Multi-Agent Spatiotemporal Reinforcement Learning for Electric Vehicle Charging Station RecommendationabstractElectric Vehicle (EV) has become preferable choices in modern transportation system due to its environmental and energy sustainability. However, in many large cities, EV drivers often fail to find proper spots for charging because of the limited charging infrastructures and spatiotemporally unbalanced charging demands. Indeed, the recent emergence of deep reinforcement learning provides great potential to improve charging experience over long-term horizons. In this paper, we propose RLCharge for intelligent EV charging station recommendation by jointly considering various long-term spatiotemporal factors. Specifically, by regarding each charging station as an agent, we formulate the problem as a multi-objective multi-agent reinforcement learning task. We first develop a multi-agent actor-critic framework with centralized training decentralized execution. Particularly, we propose a tailor designed centralized attentive critic with the delayed access strategy to coordinate the recommendation between geo-distributed agents during centralized training. Besides, we propose the spatio-temporal heterogeneous graph convolution module to handle the partial observability problem during decentralized execution. After that, to effectively optimize multiple divergent objectives, we develop a dynamic gradient re-weighting strategy to adaptively guide the optimization direction, and propose an adaptive imitation learning scheme to further accelerate and stabilize the policy convergence. Finally, extensive experiments on two real-world datasets demonstrate that RLCHARGE achieves the best comprehensive performance compared with ten baseline approaches. Weijia Zhang 0003, Hao Liu 0026, Hui Xiong 0001, Tong Xu 0001, Fan Wang 0021, Haoran Xin 0001, Hua Wu 0003 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Unified Structure Generation for Universal Information ExtractionabstractYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, Hua Wu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yaojie Lu 0001, Dai Dai, Xinyan Xiao, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
ACL (1) | 8 |
| 2022 | PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationabstractDespite recent progress of pre-trained language models on generating fluent text, existing methods still suffer from incoherence problems in long-form text generation tasks that require proper content control and planning to form a coherent high-level logical flow.In this work, we propose PLANET, a novel generation framework leveraging autoregressive self-attention mechanism to conduct content planning and surface realization dynamically.To guide the generation of output sentences, our framework enriches the Transformer decoder with latent representations to maintain sentence-level semantic plans grounded by bag-of-words.Moreover, we introduce a new coherence-based contrastive learning objective to further improve the coherence of output.Extensive experiments are conducted on two challenging longform text generation tasks including counterargument generation and opinion article generation.Both automatic and human evaluations show that our method significantly outperforms strong baselines and generates more coherent texts with richer contents. Hou Pong Chan, Xinyan Xiao, Hua Wu 0003, Lifu Huang |
ACL (1) | 5 |
| 2022 | Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User GoalsabstractMost dialog systems posit that users have figured out clear and specific goals before starting an interaction.For example, users have determined the departure, the destination, and the travel time for booking a flight.However, in many scenarios, limited by experience and knowledge, users may know what they need, but still struggle to figure out clear and specific goals by determining all the necessary slots. Zeming Liu, Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003 |
ACL (1) | 6 |
| 2022 | Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationabstractEnd-to-end simultaneous speech-to-text translation aims to directly perform translation from streaming source speech to target text with high translation quality and low latency.A typical simultaneous translation (ST) system consists of a speech translation model and a policy module, which determines when to wait and when to translate.Thus the policy is crucial to balance translation quality and latency.Conventional methods usually adopt fixed policies, e.g.segmenting the source speech with a fixed length and generating translation.However, this method ignores contextual information and suffers from low translation quality.This paper proposes an adaptive segmentation policy for end-toend ST.Inspired by human interpreters, the policy learns to segment the source streaming speech into meaningful units by considering both acoustic features and translation history, maintaining consistency between the segmentation and translation.Experimental results on English-German and Chinese-English show that our method achieves a good accuracylatency trade-off over recently proposed stateof-the-art methods.* Corresponding author. 1 In German, each singular noun is assigned a gender, either masculine, feminine, or neuter, which determines whether the definite article (like "The" in English) preceding the noun is "Der", "Die" or "Das".Therefore, translating "The" hastily without receiving the following noun may cause mistranslation.(b) Word-based policy ist Hund Ruiqing Zhang, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 3 |
| 2022 | HelixMO: Sample-Efficient Molecular Optimization in Scene-Sensitive Latent SpaceabstractEfficient exploration of the chemical space to search the candidate drugs that satisfy various constraints is a fundamental task of drug discovery. Advanced deep generative methods attempt to optimize the molecules in the compact latent space instead of the discrete original space, but the mapping between the original and latent spaces is always kept unchanged during the entire optimization process. The unchanged mapping makes those methods challenging to fast adapt to various optimization scenes and leads to the great demand for assessed molecules (samples) to provide optimization direction, which is a considerable expense for drug discovery. To this end, we design a sample-efficient molecular generative method, HelixMO, which explores the scene-sensitive latent space to promote sample efficiency. The scene-sensitive latent space focuses more on modeling the promising molecules by dynamically adjusting the space mapping by leveraging the correlations between the general and scene-specific characteristics during the optimization process. Extensive experiments demonstrate that HelixMO can achieve competitive performance with only a few assessed samples on four molecular optimization scenes. Ablation studies verify the positive impact of the scene-specific latent space, which is capable of identifying the critical characteristics of the promising molecules. We also deployed HelixMO on the website PaddleHelix (https://paddlehelix.baidu.com/app/drug/drugdesign/forecast) to provide drug design service. Xiaomin Fang, Zixu Hua, Yueyang Huang, Fan Wang 0021, Hua Wu 0003 |
BIBM | 6 |
| 2022 | A Fine-grained Interpretability Evaluation Benchmark for Neural NLPabstractLijie Wang, Yaozong Shen, Shuyuan Peng, Shuai Zhang, Xinyan Xiao, Hao Liu, Hongxuan Tang, Ying Chen, Hua Wu, Haifeng Wang. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022. Yaozong Shen, Shuyuan Peng, Xinyan Xiao, Hao Liu 0026, Hongxuan Tang, Ying Chen 0011, Hua Wu 0003, Haifeng Wang 0001 |
CoNLL | 9 |
| 2022 | DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineabstractIn this paper, we present DuReader retrieval , a large-scale Chinese dataset for passage retrieval.DuReader retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine.To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, we (1) reduce the false negatives in development and test sets by manually annotating results pooled from multiple retrievers, and (2) remove the training queries that are semantically similar to the development and testing queries.Additionally, we provide two outof-domain testing sets for cross-domain evaluation, as well as a set of human translated queries for for cross-lingual retrieval evaluation.The experiments demonstrate that DuReader retrieval is challenging and a number of problems remain unsolved, such as the salient phrase mismatch and the syntactic mismatch between queries and paragraphs.These experiments also show that dense retrievers do not generalize well across domains, and cross-lingual retrieval is essentially challenging.DuReader Yifu Qiu, Yingqi Qu, Ying Chen 0011, Qiaoqiao She, Jing Liu 0022, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP | 7 |
| 2022 | Q-TOD: A Query-driven Task-oriented Dialogue SystemabstractExisting pipelined task-oriented dialogue systems usually have difficulties adapting to unseen domains, whereas end-to-end systems are plagued by large-scale knowledge bases in practice.In this paper, we introduce a novel querydriven task-oriented dialogue system, namely Q-TOD.The essential information from the dialogue context is extracted into a query, which is further employed to retrieve relevant knowledge records for response generation.Firstly, as the query is in the form of natural language and not confined to the schema of the knowledge base, the issue of domain adaption is alleviated remarkably in Q-TOD.Secondly, as the query enables the decoupling of knowledge retrieval from the generation, Q-TOD gets rid of the issue of knowledge base scalability.To evaluate the effectiveness of the proposed Q-TOD, we collect query annotations for three publicly available task-oriented dialogue datasets.Comprehensive experiments verify that Q-TOD outperforms strong baselines and establishes a new state-of-the-art performance on these datasets. Xin Tian 0011, Yingzhan Lin, Mengfei Song, Siqi Bao, Fan Wang 0021, Huang He, Shu-Qi Sun, Hua Wu 0003 |
EMNLP | 8 |
| 2022 | CDConv: A Benchmark for Contradiction Detection in Chinese ConversationsabstractChujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Zhen Guo, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Minlie Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Wenquan Wu, Zhengyu Niu, Hua Wu 0003, Minlie Huang |
EMNLP | 8 |
| 2022 | DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching ModelsabstractIn this paper, we focus on the robustness evaluation of Chinese Question Matching (QM) models.Most of the previous work on analyzing robustness issues focus on just one or a few types of artificial adversarial examples.Instead, we argue that a comprehensive evaluation should be conducted on natural texts, which takes into account the fine-grained linguistic capabilities of QM models.For this purpose, we create a Chinese dataset namely DuQM which contains natural questions with linguistic perturbations to evaluate the robustness of QM models.DuQM contains 3 categories and 13 subcategories with 32 linguistic perturbations.The extensive experiments demonstrate that DuQM has a better ability to distinguish different models.Importantly, the detailed breakdown of evaluation by the linguistic phenomenon in DuQM helps us easily diagnose the strength and weakness of different models.Additionally, our experiment results show that the effect of artificial adversarial examples does not work on natural texts.Our baseline codes and a leaderboard are now publicly available.1 Hongyu Zhu 0002, Jing Yan 0004, Jing Liu 0022, Yu Hong 0001, Ying Chen 0011, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP | 7 |
| 2022 | CLOP: Video-and-Language Pre-Training with Knowledge RegularizationsabstractVideo-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner, without considering explicit structural representations of the multi-modal content. We denote such form of representations as structural knowledge, which express rich semantics of multiple granularities. There are related works that propose object-aware approaches to inject similar knowledge as inputs. However, the existing methods usually fail to effectively utilize such knowledge as regularizations to shape a superior cross-modal representation space. To this end, we propose a Cross-modaL knOwledge-enhanced Pre-training (CLOP) method with Knowledge Regularizations. There are two key designs of ours: 1) a simple yet effective Structural Knowledge Prediction (SKP) task to pull together the latent representations of similar videos; and 2) a novel Knowledge-guided sampling approach for Contrastive Learning (KCL) to push apart cross-modal hard negative samples. We evaluate our method on four text-video retrieval tasks and one multi-choice QA task. The experiments show clear improvements, outperforming prior works by a substantial margin. Besides, we provide ablations and insights of how our methods affect the latent representation space, demonstrating the value of incorporating knowledge regularizations into video-and-language pre-training. Guohao Li 0002, Zhifan Feng, Yajuan Lyu, Hua Wu 0003, Haifeng Wang 0001 |
ACM Multimedia | 6 |
| 2022 | Non-Autoregressive Chinese ASR Error Correction with Phonological TrainingabstractZheng Fang, Ruiqing Zhang, Zhongjun He, Hua Wu, Yanan Cao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zheng Fang 0002, Ruiqing Zhang, Zhongjun He, Hua Wu 0003, Yanan Cao 0001 |
NAACL-HLT | 4 |
| 2022 | Bi-SimCut: A Simple Strategy for Boosting Neural Machine TranslationabstractWe introduce Bi-SimCut: a simple but effective training strategy to boost neural machine translation (NMT) performance.It consists of two procedures: bidirectional pretraining and unidirectional finetuning.Both procedures utilize SimCut, a simple regularization method that forces the consistency between the output distributions of the original and the cutoff sentence pairs.Without leveraging extra dataset via back-translation or integrating large-scale pretrained model, Bi-SimCut achieves strong translation performance across five translation benchmarks (data sizes range from 160K to 20.2M): BLEU scores of 31.16 for en → de and 38.37 for de → en on the IWSLT14 dataset, 30.78 for en → de and 35.15 for de → en on the WMT14 dataset, and 27.17 for zh → en on the WMT17 dataset.Sim-Cut is not a new method, but a version of Cutoff (Shen et al., 2020) simplified and adapted for NMT, and it could be considered as a perturbation-based method.Given the universality and simplicity of SimCut and Bi-SimCut, we believe they can serve as strong baselines for future NMT research. Pengzhi Gao, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
NAACL-HLT | 3 |
| 2022 | DTSyn: a dual-transformer-based neural network to predict synergistic drug combinationsabstractDrug combination therapies are superior to monotherapy for cancer treatment in many ways. Identifying novel drug combinations by screening is challenging for the wet-lab experiments due to the time-consuming process of the enormous search space of possible drug pairs. Thus, computational methods have been developed to predict drug pairs with potential synergistic functions. Notwithstanding the success of current models, understanding the mechanism of drug synergy from a chemical-gene-tissue interaction perspective lacks study, hindering current algorithms from drug mechanism study. Here, we proposed a deep neural network model termed DTSyn (Dual Transformer encoder model for drug pair Synergy prediction) based on a multi-head attention mechanism to identify novel drug combinations. We designed a fine-granularity transformer encoder to capture chemical substructure-gene and gene-gene associations and a coarse-granularity transformer encoder to extract chemical-chemical and chemical-cell line interactions. DTSyn achieved the highest receiver operating characteristic area under the curve of 0.73, 0.78. 0.82 and 0.81 on four different cross-validation tasks, outperforming all competing methods. Further, DTSyn achieved the best True Positive Rate (TPR) over five independent data sets. The ablation study showed that both transformer encoder blocks contributed to the performance of DTSyn. In addition, DTSyn can extract interactions among chemicals and cell lines, representing the potential mechanisms of drug action. By leveraging the attention mechanism and pretrained gene embeddings, DTSyn shows improved interpretability ability. Thus, we envision our model as a valuable tool to prioritize synergistic drug pairs with chemical and cell line gene expression profile. Xiaomin Fang, Zijing Liu, Fan Wang 0021, Weili Huang, Hua Wu 0003 |
Briefings Bioinform. | 7 |
| 2022 | BatchDTA: implicit batch alignment enhances deep learning-based drug-target affinity estimationabstractCandidate compounds with high binding affinities toward a target protein are likely to be developed as drugs. Deep neural networks (DNNs) have attracted increasing attention for drug-target affinity (DTA) estimation owning to their efficiency. However, the negative impact of batch effects caused by measure metrics, system technologies and other assay information is seldom discussed when training a DNN model for DTA. Suffering from the data deviation caused by batch effects, the DNN models can only be trained on a small amount of 'clean' data. Thus, it is challenging for them to provide precise and consistent estimations. We design a batch-sensitive training framework, namely BatchDTA, to train the DNN models. BatchDTA implicitly aligns multiple batches toward the same protein through learning the orders of candidate compounds with respect to the batches, alleviating the impact of the batch effects on the DNN models. Extensive experiments demonstrate that BatchDTA facilitates four mainstream DNN models to enhance the ability and robustness on multiple DTA datasets (BindingDB, Davis and KIBA). The average concordance index of the DNN models achieves a relative improvement of 4.0%. The case study reveals that BatchDTA can successfully learn the ranking orders of the compounds from multiple batches. In addition, BatchDTA can also be applied to the fused data collected from multiple sources to achieve further improvement. Hongyu Luo, Yingfei Xiang, Xiaomin Fang, Wei Li 0176, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001 |
Briefings Bioinform. | 6 |
| 2022 | HelixADMET: a robust and endpoint extensible ADMET system incorporating self-supervised knowledge transferabstractMOTIVATION: Accurate ADMET (an abbreviation for 'absorption, distribution, metabolism, excretion and toxicity') predictions can efficiently screen out undesirable drug candidates in the early stage of drug discovery. In recent years, multiple comprehensive ADMET systems that adopt advanced machine learning models have been developed, providing services to estimate multiple endpoints. However, those ADMET systems usually suffer from weak extrapolation ability. First, due to the lack of labelled data for each endpoint, typical machine learning models perform frail for the molecules with unobserved scaffolds. Second, most systems only provide fixed built-in endpoints and cannot be customized to satisfy various research requirements. To this end, we develop a robust and endpoint extensible ADMET system, HelixADMET (H-ADMET). H-ADMET incorporates the concept of self-supervised learning to produce a robust pre-trained model. The model is then fine-tuned with a multi-task and multi-stage framework to transfer knowledge between ADMET endpoints, auxiliary tasks and self-supervised tasks. RESULTS: Our results demonstrate that H-ADMET achieves an overall improvement of 4%, compared with existing ADMET systems on comparable endpoints. Additionally, the pre-trained model provided by H-ADMET can be fine-tuned to generate new and customized ADMET endpoints, meeting various demands of drug research and development requirements. AVAILABILITY AND IMPLEMENTATION: H-ADMET is freely accessible at https://paddlehelix.baidu.com/app/drug/admet/train. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shanzhuo Zhang, Zhiyuan Yan 0002, Yueyang Huang, Lihang Liu, Donglong He, Xiaomin Fang, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001 |
Bioinform. | 10 |
| 2022 | Towards Knowledge-Aware Video Captioning via Transitive Visual Relationship DetectionabstractVideo captioning can be enhanced by incorporating the knowledge, which is usually represented as relationships of objects. However, the previous methods construct only superficial or static object relationships, and often introduce noise into the task through irrelevant common sense or fixed syntax templates. These problems mitigate the model interpretability and lead to the undesirable consequence. To overcome these limitations, we propose to enhance video captioning with deep-level object relationships that are adaptively explored during training. Specifically, we present a Transitive Visual Relationship Detection (TVRD) module in which we estimate the actions of the visual objects, and construct an Object-Action Graph (OAG) to describe the shallow relationship between the objects and actions. Then we bridge the gap between the objects via the actions to transitively infer an Object-Object Graph (OOG) which reflects the deep-level relationship. We further feed the OOG to a graph convolutional network to refine the object representation by deep-level relationships. With the refined representation, we capitalize on an LSTM-based decoder for caption generation. Experimental results on two benchmark datasets: MSVD, MSR-VTT demonstrate that the proposed method achieves state-of-the-art performance. Lastly, we present comprehensive ablation studies as well as visualization of visual relationships to demonstrate the effectiveness and interpretability of our model. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsabstractWe propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) across vision and language, which are essential to vision-language cross-modal tasks. Utilizing scene graphs of visual scenes, ERNIE-ViL constructs Scene Graph Prediction tasks, i.e., Object Prediction, Attribute Prediction and Relationship Prediction tasks in the pre-training phase. Specifically, these prediction tasks are implemented by predicting nodes of different types in the scene graph parsed from the sentence. Thus, ERNIE-ViL can learn the joint representations characterizing the alignments of the detailed semantics across vision and language. After pre-training on large scale image-text aligned datasets, we validate the effectiveness of ERNIE-ViL on 5 cross-modal downstream tasks. ERNIE-ViL achieves state-of-the-art performances on all these tasks and ranks the first place on the VCR leaderboard with an absolute improvement of 3.7%. Fei Yu 0010, Jiji Tang, Weichong Yin, Yu Sun 0029, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 6 |
| 2021 | Discovering Dialog Structure Graph for Coherent Dialog GenerationabstractJun Xu, Zeyang Lei, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
ACL/IJCNLP (1) | 5 |
| 2021 | ERNIE-Doc: A Retrospective Long-Document Modeling TransformerabstractSiYu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningabstractWei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei Li 0176, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu 0026, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | BASS: Boosting Abstractive Summarization with Unified Semantic GraphabstractWenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu, Ziqiang Cao, Sujian Li, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei Li 0176, Xinyan Xiao, Ziqiang Cao, Sujian Li, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | Docking-based Virtual Screening with Multi-Task LearningabstractMachine learning shows great potential in virtual screening for drug discovery. Current efforts on accelerating docking-based virtual screening do not consider using existing data of other previously developed targets. To make use of the knowledge of the other targets and take advantage of the existing data, in this work, we apply multi-task learning to the problem of docking-based virtual screening. With two large docking datasets, the results of extensive experiments show that multi-task learning can achieve better performances on docking score prediction. By learning knowledge across multiple targets, the model trained by multi-task learning shows a better ability to adapt to a new target. Additional empirical study shows that other problems in drug discovery, such as the experimental drug-target affinity prediction, may also benefit from multi-task learning. Our results demonstrate that multi-task learning is a promising machine learning approach for docking-based virtual screening and accelerating the process of drug discovery. Zijing Liu, Xianbin Ye, Xiaoming Fang, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001 |
BIBM | 5 |
| 2021 | Familia: A Configurable Topic Modeling Framework for Industrial Text Engineering
Di Jiang 0004, Yuanfeng Song, Rongzhong Lian, Siqi Bao, Jinhua Peng, Huang He, Hua Wu 0003, Chen Zhang 0013, Lei Chen 0002 |
DASFAA (3) | 7 |
| 2021 | SgSum: Transforming Multi-document Summarization into Sub-graph SelectionabstractMost of existing extractive multi-document summarization (MDS) methods score each sentence individually and extract salient sentences one by one to compose a summary, which have two main drawbacks: (1) neglecting both the intra and cross-document relations between sentences; (2) neglecting the coherence and conciseness of the whole summary.In this paper, we propose a novel MDS framework (SgSum) to formulate the MDS task as a sub-graph selection problem, in which source documents are regarded as a relation graph of sentences (e.g., similarity graph or discourse graph) and the candidate summaries are its subgraphs.Instead of selecting salient sentences, SgSum selects a salient sub-graph from the relation graph as the summary.Comparing with traditional methods, our method has two main advantages: (1) the relations between sentences are captured by modeling both the graph structure of the whole document set and the candidate sub-graphs; (2) directly outputs an integrate summary in the form of subgraph which is more informative and coherent.Extensive experiments on MultiNews and DUC datasets show that our proposed method brings substantial improvements over several strong baselines.Human evaluation results also demonstrate that our model can produce significantly more coherent and informative summaries compared with traditional MDS methods.Moreover, the proposed architecture has strong transfer ability from single to multi-document input, which can reduce the resource bottleneck in MDS tasks. 1 Moye Chen, Wei Li 0176, Xinyan Xiao, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP (1) | 5 |
| 2021 | DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationabstractIn this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendation.The difference between DuRecDial 2.0 and existing conversational recommendation datasets is that the data item (Profile, Goal, Knowledge, Context, Response) in DuRecDial 2.0 is annotated in two languages, both English and Chinese, while other datasets are built with the setting of a single language.We collect 8.2k dialogs aligned across English and Chinese languages (16.5k dialogs and 255k utterances in total) that are annotated by crowdsourced workers with strict quality control procedure.We then build monolingual, multilingual, and cross-lingual conversational recommendation baselines on DuRecDial 2.0.Experiment results show that the use of additional English data can bring performance improvement for Chinese conversational recommendation, indicating the benefits of DuRecDial 2.0.Finally, this dataset provides a challenging testbed for future studies of monolingual, multilingual, and cross-lingual conversational recommendation. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
EMNLP (1) | 4 |
| 2021 | Fine-grained Entity Typing via Label ReasoningabstractConventional entity typing approaches are based on independent classification paradigms, which make them difficult to recognize interdependent, long-tailed and fine-grained entity types.In this paper, we argue that the implicitly entailed extrinsic and intrinsic dependencies between labels can provide critical knowledge to tackle the above challenges.To this end, we propose Label Reasoning Network(LRN), which sequentially reasons finegrained entity labels by discovering and exploiting label dependencies knowledge entailed in the data.Specifically, LRN utilizes an auto-regressive network to conduct deductive reasoning and a bipartite attribute graph to conduct inductive reasoning between labels, which can effectively model, learn and reason complex label dependencies in a sequence-toset, end-to-end manner.Experiments show that LRN achieves the state-of-the-art performance on standard ultra fine-grained entity typing benchmarks, and can also resolve the long tail label problem effectively. Xinyan Xiao, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
EMNLP (1) | 6 |
| 2021 | ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual CorporaabstractRecent studies have demonstrated that pre-trained cross-lingual models achieve impressive performance in downstream cross-lingual tasks. This improvement benefits from learning a large amount of monolingual and parallel corpora. Although it is generally acknowledged that parallel corpora are critical for improving the model performance, existing methods are often constrained by the size of parallel corpora, especially for low-resource languages. In this paper, we propose Ernie-M, a new training method that encourages the model to align the representation of multiple languages with monolingual corpora, to overcome the constraint that the parallel corpus size places on the model performance. Our key insight is to integrate back-translation into the pre-training process. We generate pseudo-parallel sentence pairs on a monolingual corpus to enable the learning of semantic alignments between different languages, thereby enhancing the semantic modeling of cross-lingual models. Experimental results show that Ernie-M outperforms existing cross-lingual models and delivers new state-of-the-art results in various cross-lingual downstream tasks. The codes and pre-trained models will be made publicly available. Xuan Ouyang, Shuohuan Wang, Yu Sun 0029, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP (1) | 6 |
| 2021 | RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingabstractIn various natural language processing tasks, passage retrieval and passage re-ranking are two key procedures in finding and ranking relevant information.Since both the two procedures contribute to the final performance, it is important to jointly optimize them in order to achieve mutual improvement.In this paper, we propose a novel joint training approach for dense passage retrieval and passage reranking.A major contribution is that we introduce the dynamic listwise distillation, where we design a unified listwise training approach for both the retriever and the re-ranker.During the dynamic distillation, the retriever and the re-ranker can be adaptively improved according to each other's relevance information.We also propose a hybrid data augmentation strategy to construct diverse training instances for listwise training approach.Extensive experiments show the effectiveness of our approach on both MSMARCO and Natural Questions datasets.Our code is available at https:// github.com/PaddlePaddle/RocketQA. Ruiyang Ren, Yingqi Qu, Jing Liu 0022, Wayne Xin Zhao, Qiaoqiao She, Hua Wu 0003, Haifeng Wang 0001, Ji-Rong Wen |
EMNLP (1) | 6 |
| 2021 | Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingabstractData augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness.The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing.Previous works either require human intervention to guarantee the quality of generated data, or fail to handle complex SQL queries.This paper presents a simple yet effective data augmentation framework.First, given a database, we automatically produce a large number of SQL queries based on an abstract syntax tree grammar.For better distribution matching, we require that at least 80% of SQL patterns in the training data are covered by generated queries.Second, we propose a hierarchical SQL-to-question generation model to obtain high-quality natural language questions, which is the major contribution of this work.Finally, we design a simple sampling strategy that can greatly improve training efficiency given large amounts of generated data.Experiments on three cross-domain datasets, i.e., WikiSQL and Spider in English, and DuSQL in Chinese, show that our proposed data augmentation framework can consistently improve performance over strong baselines, and the hierarchical generation component is the key for the improvement. Kun Wu 0009, Zhenghua Li, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005, Haifeng Wang 0001 |
EMNLP (1) | 6 |
| 2021 | Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal MatchingabstractThis paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved tasks to generate high-quality event proposals. Then we incorporate contrastive loss and cycle-consistency loss typically applied to cross-modal retrieval tasks to build semantic matching between the proposals and sentences, which are eventually used to train the caption generation module. In addition, the parameters of matching module are initialized via pre-training based on annotated images to improve the matching performance. Extensive experiments on ActivityNet-Caption dataset reveal the significance of distillation-based event proposal generation and cross-modal retrieval-based semantic matching to weakly supervised DVC, and demonstrate the superiority of our method to existing state-of-the-art methods. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IJCAI | 6 |
| 2021 | RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question AnsweringabstractYingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, Haifeng Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yingqi Qu, Yuchen Ding, Jing Liu 0022, Kai Liu 0023, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu 0003, Haifeng Wang 0001 |
NAACL-HLT | 8 |
| 2021 | ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language UnderstandingabstractDongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dongling Xiao, Yu-Kun Li, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
NAACL-HLT | 6 |
| 2021 | Learning with Noisy Correspondence for Cross-modal MatchingabstractCross-modal matching, which aims to establish the correspondence between two different modalities, is fundamental to a variety of tasks such as cross-modal retrieval and vision-and-language understanding. Although a huge number of cross-modal matching methods have been proposed and achieved remarkable progress in recent years, almost all of these methods implicitly assume that the multimodal training data are correctly aligned. In practice, however, such an assumption is extremely expensive even impossible to satisfy. Based on this observation, we reveal and study a latent and challenging direction in cross-modal matching, named noisy correspondence, which could be regarded as a new paradigm of noisy labels. Different from the traditional noisy labels which mainly refer to the errors in category labels, our noisy correspondence refers to the mismatch paired samples. To solve this new problem, we propose a novel method for learning with noisy correspondence, named Noisy Correspondence Rectifier (NCR). In brief, NCR divides the data into clean and noisy partitions based on the memorization effect of neural networks and then rectifies the correspondence via an adaptive prediction model in a co-teaching manner. To verify the effectiveness of our method, we conduct experiments by using the image-text matching as a showcase. Extensive experiments on Flickr30K, MS-COCO, and Conceptual Captions verify the effectiveness of our method. The code could be accessed from www.pengxi.me . Zhenyu Huang 0005, Guocheng Niu, Xiao Liu 0040, Wenbiao Ding, Xinyan Xiao, Hua Wu 0003, Xi Peng 0001 |
NeurIPS | 6 |
| 2021 | Coherent Dialog Generation with Query GraphabstractLearning to generate coherent and informative dialogs is an enduring challenge for open-domain conversation generation. Previous work leverage knowledge graph or documents to facilitate informative dialog generation, with little attention on dialog coherence. In this article, to enhance multi-turn open-domain dialog coherence, we propose to leverage a new knowledge source, web search session data, to facilitate hierarchical knowledge sequence planning, which determines a sketch of a multi-turn dialog. Specifically, we formulate knowledge sequence planning or dialog policy learning as a graph grounded Reinforcement Learning (RL) problem. To this end, we first build a two-level query graph with queries as utterance-level vertices and their topics (entities in queries) as topic-level vertices. We then present a two-level dialog policy model that plans a high-level topic sequence and a low-level query sequence over the query graph to guide a knowledge aware response generator. In particular, to foster forward-looking knowledge planning decisions for better dialog coherence, we devise a heterogeneous graph neural network to incorporate neighbouring vertex information, or possible future RL action information, into each vertex (as an RL action) representation. Experiment results on two benchmark dialog datasets demonstrate that our framework can outperform strong baselines in terms of dialog coherence, informativeness, and engagingness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Jizhou Huang, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2020 | Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingabstractSpeech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential benefits of lower latency, smaller model size, and less error propagation. However, it is notoriously difficult to implement such a model without transcriptions as intermediate. Existing works generally apply multi-task learning to improve translation quality by jointly training end-to-end ST along with automatic speech recognition (ASR). However, different tasks in this method cannot utilize information from each other, which limits the improvement. Other works propose a two-stage model where the second model can use the hidden state from the first one, but its cascade manner greatly affects the efficiency of training and inference process. In this paper, we propose a novel interactive attention mechanism which enables ASR and ST to perform synchronously and interactively in a single model. Specifically, the generation of transcriptions and translations not only relies on its previous outputs but also the outputs predicted in the other task. Experiments on TED speech translation corpora have shown that our proposed model can outperform strong baselines on the quality of speech translation and achieve better speech recognition performances as well. Yuchen Liu 0007, Jiajun Zhang 0001, Hao Xiong 0005, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
AAAI | 6 |
| 2020 | ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingabstractRecently pre-trained models have achieved state-of-the-art results in various language understanding tasks. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides co-occurring information, there exists other valuable lexical, syntactic and semantic information in training corpora, such as named entities, semantic closeness and discourse relations. In order to extract the lexical, syntactic and semantic information from training corpora, we propose a continual pre-training framework named ERNIE 2.0 which incrementally builds pre-training tasks and then learn pre-trained models on these constructed tasks via continual multi-task learning. Based on this framework, we construct several tasks and train the ERNIE 2.0 model to capture lexical, syntactic and semantic aspects of information in the training data. Experimental results demonstrate that ERNIE 2.0 model outperforms BERT and XLNet on 16 tasks including English tasks on GLUE benchmarks and several similar tasks in Chinese. The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE. Yu Sun 0004, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 6 |
| 2020 | Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationabstractPrevious neural models on open-domain conversation generation have no effective mechanisms to manage chatting topics, and tend to produce less coherent dialogs. Inspired by the strategies in human-human dialogs, we divide the task of multi-turn open-domain conversation generation into two sub-tasks: explicit goal (chatting about a topic) sequence planning and goal completion by topic elaboration. To this end, we propose a three-layer Knowledge aware Hierarchical Reinforcement Learning based Model (KnowHRL). Specifically, for the first sub-task, the upper-layer policy learns to traverse a knowledge graph (KG) in order to plan a high-level goal sequence towards a good balance between dialog coherence and topic consistency with user interests. For the second sub-task, the middle-layer policy and the lower-layer one work together to produce an in-depth multi-turn conversation about a single topic with a goal-driven generation mechanism. The capability of goal-sequence planning enables chatbots to conduct proactive open-domain conversations towards recommended topics, which has many practical applications. Experiments demonstrate that our model outperforms state of the art baselines in terms of user-interest consistency, dialog coherence, and knowledge accuracy. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
AAAI | 4 |
| 2020 | PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableabstractPre-training models have been proved effective for a wide range of natural language processing tasks.Inspired by this, we propose a novel dialogue generation pre-training framework to support various kinds of conversations, including chit-chat, knowledge grounded dialogues, and conversational question answering.In this framework, we adopt flexible attention mechanisms to fully leverage the bi-directional context and the uni-directional characteristic of language generation.We also introduce discrete latent variables to tackle the inherent one-to-many mapping problem in response generation.Two reciprocal tasks of response generation and latent act recognition are designed and carried out simultaneously within a shared network.Comprehensive experiments on three publicly available datasets verify the effectiveness and superiority of the proposed framework. Siqi Bao, Huang He, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001 |
ACL | 4 |
| 2020 | Leveraging Graph to Improve Abstractive Multi-Document SummarizationabstractGraphs that capture relations between textual units have great benefits for detecting salient information from multiple documents and generating overall coherent summaries.In this paper, we develop a neural abstractive multidocument summarization (MDS) model which can leverage well-known graph representations of documents such as similarity graph and discourse graph, to more effectively process multiple input documents and produce abstractive summaries.Our model utilizes graphs to encode documents in order to capture cross-document relations, which is crucial to summarizing long documents.Our model can also take advantage of graphs to guide the summary generation process, which is beneficial for generating coherent and concise summaries.Furthermore, pre-trained language models can be easily combined with our model, which further improve the summarization performance significantly.Empirical results on the WikiSum and MultiNews dataset show that the proposed architecture brings substantial improvements over several strong baselines. Wei Li 0176, Xinyan Xiao, Hua Wu 0003, Haifeng Wang 0001, Junping Du 0001 |
ACL | 4 |
| 2020 | Towards Conversational Recommendation over Multi-Type DialogsabstractWe focus on the study of conversational recommendation in the context of multi-type dialogs, where the bots can proactively and naturally lead a conversation from a nonrecommendation dialog (e.g., QA) to a recommendation dialog, taking into account user's interests and feedback.To facilitate the study of this task, we create a human-to-human Chinese dialog dataset DuRecDial (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).In each dialog, the recommender proactively leads a multi-type dialog to approach recommendation targets and then makes multiple recommendations with rich interaction behavior.This dataset allows us to systematically investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how to interact with users for recommendation.Finally we establish baseline results on DuRecDial for future studies. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 4 |
| 2020 | SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment AnalysisabstractRecently, sentiment analysis has seen remarkable advance with the help of pre-training approaches. However, sentiment knowledge, such as sentiment words and aspect-sentiment pairs, is ignored in the process of pre-training, despite the fact that they are widely used in traditional sentiment analysis approaches. In this paper, we introduce Sentiment Knowledge Enhanced Pre-training (SKEP) in order to learn a unified sentiment representation for multiple sentiment analysis tasks. With the help of automatically-mined knowledge, SKEP conducts sentiment masking and constructs three sentiment knowledge prediction objectives, so as to embed sentiment information at the word, polarity and aspect level into pre-trained sentiment representation. In particular, the prediction of aspect-sentiment pairs is converted into multi-label classification, aiming to capture the dependency between words in a pair. Experiments on three kinds of sentiment tasks show that SKEP significantly outperforms strong pre-training baseline, and achieves new state-of-the-art results on most of the test datasets. We release our code at https://github.com/baidu/Senta. Hao Tian 0005, Can Gao, Xinyan Xiao, Hao Liu 0026, Bolei He, Hua Wu 0003, Haifeng Wang 0001, Feng Wu 0001 |
ACL | 6 |
| 2020 | Conversational Graph Grounded Policy Learning for Open-Domain Conversation GenerationabstractTo address the challenge of policy learning in open-domain multi-turn conversation, we propose to represent prior information about dialog transitions as a graph and learn a graph grounded dialog policy, aimed at fostering a more coherent and controllable dialog.To this end, we first construct a conversational graph (CG) from dialog corpora, in which there are vertices to represent "what to say" and "how to say", and edges to represent natural transition between a message (the last utterance in a dialog context) and its response.We then present a novel CG grounded policy learning framework that conducts dialog flow planning by graph traversal, which learns to identify a what-vertex and a how-vertex from the CG at each turn to guide response generation.In this way, we effectively leverage the CG to facilitate policy learning as follows: (1) it enables more effective long-term reward design, (2) it provides high-quality candidate actions, and (3) it gives us more control over the policy.Results on two benchmark corpora demonstrate the effectiveness of this framework. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 4 |
| 2020 | Exploring Contextual Word-level Style Relevance for Unsupervised Style TransferabstractUnsupervised style transfer aims to change the style of an input sentence while preserving its original content without using parallel training data. In current dominant approaches, owing to the lack of fine-grained control on the influence from the target style, they are unable to yield desirable output sentences. In this paper, we propose a novel attentional sequence-to-sequence (Seq2seq) model that dynamically exploits the relevance of each output word to the target style for unsupervised style transfer. Specifically, we first pretrain a style classifier, where the relevance of each input word to the original style can be quantified via layer-wise relevance propagation. In a denoising auto-encoding manner, we train an attentional Seq2seq model to reconstruct input sentences and repredict word-level previously-quantified style relevance simultaneously. In this way, this model is endowed with the ability to automatically predict the style relevance of each output word. Then, we equip the decoder of this model with a neural style component to exploit the predicted wordlevel style relevance for better style transfer. Particularly, we fine-tune this model using a carefully-designed objective function involving style transfer, style relevance consistency, content preservation and fluency modeling loss terms. Experimental results show that our proposed model achieves state-of-the-art performance in terms of both transfer accuracy and content preservation. Chulun Zhou, Xinyan Xiao, Jinsong Su, Hua Wu 0003 |
ACL | 7 |
| 2020 | Diversified Multiple Instance Learning for Document-Level Multi-Aspect Sentiment ClassificationabstractNeural Document-level Multi-aspect Sentiment Classification (DMSC) usually requires a lot of manual aspect-level sentiment annotations, which is time-consuming and laborious.As document-level sentiment labeled data are widely available from online service, it is valuable to perform DMSC with such free document-level annotations.To this end, we propose a novel Diversified Multiple Instance Learning Network (D-MILN), which is able to achieve aspect-level sentiment classification with only document-level weak supervision.Specifically, we connect aspect-level and document-level sentiment by formulating this problem as multiple instance learning, providing a way to learn aspect-level classifier from the back propagation of document-level supervision.Two diversified regularizations are further introduced in order to avoid the overfitting on document-level signals during training.Diversified textual regularization encourages the classifier to select aspect-relevant snippets, and diversified sentimental regularization prevents the aspect-level sentiments from being overly consistent with document-level sentiment.Experimental results on TripAdvisor and BeerAdvocate datasets show that D-MILN remarkably outperforms recent weaklysupervised baselines, and is also comparable to the supervised method. Yunjie Ji, Hao Liu 0026, Bolei He, Xinyan Xiao, Hua Wu 0003, Yanhua Yu |
EMNLP (1) | 5 |
| 2020 | DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL DatasetabstractDue to the lack of labeled data, previous research on text-to-SQL parsing mainly focuses on English.Representative English datasets include ATIS, WikiSQL, Spider, etc.This paper presents DuSQL, a larges-scale and pragmatic Chinese dataset for the cross-domain text-to-SQL task, containing 200 databases, 813 tables, and 23,797 question/SQL pairs.Our new dataset has three major characteristics.First, by manually analyzing questions from several representative applications, we try to figure out the true distribution of SQL queries in real-life needs.Second, DuSQL contains a considerable proportion of SQL queries involving row or column calculations, motivated by our analysis on the SQL query distributions.Finally, we adopt an effective data construction framework via human-computer collaboration.The basic idea is automatically generating SQL queries based on the SQL grammar and constrained by the given database.This paper describes in detail the construction process and data statistics of DuSQL.Moreover, we present and compare performance of several open-source textto-SQL parsers with minor modification to accommodate Chinese, including a simple yet effective extension to IRNet for handling calculation SQL queries. Kun Wu 0009, Ke Sun 0005, Zhenghua Li, Hua Wu 0003, Min Zhang 0005, Haifeng Wang 0001 |
EMNLP (1) | 6 |
| 2020 | Learning Adaptive Segmentation Policy for Simultaneous TranslationabstractBalancing accuracy and latency is a great challenge for simultaneous translation.To achieve high accuracy, the model usually needs to wait for more streaming text before translation, which results in increased latency.However, keeping low latency would probably hurt accuracy.Therefore, it is essential to segment the ASR output into appropriate units for translation.Inspired by human interpreters, we propose a novel adaptive segmentation policy for simultaneous translation.The policy learns to segment the source text by considering possible translations produced by the translation model, maintaining consistency between the segmentation and translation.Experimental results on Chinese-English and German-English translation show that our method achieves a better accuracy-latency trade-off over recently proposed state-of-the-art methods. Ruiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP (1) | 4 |
| 2020 | TopicOcean: An Ever-Increasing Topic Model With Meta-learningabstractTopic modeling has been intensively studied and widely applied in both academia and industry in the last decade. In the literature, topic models usually need to be trained from scratch for each individual corpus. Hence, the wisdom of the crowd (i.e., topic models previously trained based upon other corpora) is abandoned. Since a massive amount of in-domain data, considerable computational cost, and human labour are involved in obtaining a high-quality topic model, training from scratch for each new corpus is a huge waste of resources. In this paper, we propose the novel TopicOcean framework, which aims to integrate well-trained topic models and transfer the knowledge of accumulated topics to new corpora in order to improve the quality of their topic models. We first propose a method of constructing the ever-increasing TopicOcean, and then propose a meta-learning mechanism that transfers the meta-level knowledge (i.e., topics) in TopicOcean to the scenario of topic modeling on new corpora. Comprehensive experiments validate that the TopicOcean framework can significantly outperform the state-of-the-art (53.77% perplexity improvement on a temporal-shift corpus and 29.24% improvement on a domain-shift corpus). The well-trained high-quality topic models used to construct TopicOcean have been opensourced to promote further research.11The well-trained topic models can be accessed at Github (https://github.com/baidu/Familia/blob/master/model/download_model.sh). Yuanfeng Song, Yongxin Tong, Siqi Bao, Di Jiang 0004, Hua Wu 0003, Raymond Chi-Wing Wong |
ICDM | 5 |
| 2020 | ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language GenerationabstractCurrent pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks. To address this issue, we propose an enhanced multi-flow sequence to sequence pre-training and fine-tuning framework named ERNIE-GEN, which bridges the discrepancy between training and inference with an infilling generation mechanism and a noise-aware generation method. To make generation closer to human writing patterns, this framework introduces a span-by-span generation flow that trains the model to predict semantically-complete spans consecutively rather than predicting word by word. Unlike existing pre-training methods, ERNIE-GEN incorporates multi-granularity target sampling to construct pre-training data, which enhances the correlation between encoder and decoder. Experimental results demonstrate that ERNIE-GEN achieves state-of-the-art results with a much smaller amount of pre-training data and parameters on a range of language generation tasks, including abstractive summarization (Gigaword and CNN/DailyMail), question generation (SQuAD), dialogue generation (Persona-Chat) and generative question answering (CoQA). The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE/ernie-gen. Dongling Xiao, Yu-Kun Li, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
IJCAI | 6 |
| 2020 | Enhancing Dialog Coherence with Event Graph Grounded Content PlanningabstractHow to generate informative, coherent and sustainable open-domain conversations is a non-trivial task. Previous work on knowledge grounded conversation generation focus on improving dialog informativeness with little attention on dialog coherence. In this paper, to enhance multi-turn dialog coherence, we propose to leverage event chains to help determine a sketch of a multi-turn dialog. We first extract event chains from narrative texts and connect them as a graph. We then present a novel event graph grounded Reinforcement Learning (RL) framework. It conducts high-level response content (simply an event) planning by learning to walk over the graph, and then produces a response conditioned on the planned content. In particular, we devise a novel multi-policy decision making mechanism to foster a coherent dialog with both appropriate content ordering and high contextual relevance. Experimental results indicate the effectiveness of this framework in terms of dialog coherence and informativeness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
IJCAI | 5 |
| 2019 | Modeling Coherence for Discourse Neural Machine TranslationabstractDiscourse coherence plays an important role in the translation of one text. However, the previous reported models most focus on improving performance over individual sentence while ignoring cross-sentence links and dependencies, which affects the coherence of the text. In this paper, we propose to use discourse context and reward to refine the translation quality from the discourse perspective. In particular, we generate the translation of individual sentences at first. Next, we deliberate the preliminary produced translations, and train the model to learn the policy that produces discourse coherent text by a reward teacher. Practical results on multiple discourse test datasets indicate that our model significantly improves the translation quality over the state-of-the-art baseline system by +1.23 BLEU score. Moreover, our model generates more discourse coherent text and obtains +2.2 BLEU improvements when evaluated by discourse metrics. Hao Xiong 0005, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 3 |
| 2019 | Addressing the Under-Translation Problem from the Entropy PerspectiveabstractNeural Machine Translation (NMT) has drawn much attention due to its promising translation performance in recent years. However, the under-translation problem still remains a big challenge. In this paper, we focus on the under-translation problem and attempt to find out what kinds of source words are more likely to be ignored. Through analysis, we observe that a source word with a large translation entropy is more inclined to be dropped. To address this problem, we propose a coarse-to-fine framework. In coarse-grained phase, we introduce a simple strategy to reduce the entropy of highentropy words through constructing the pseudo target sentences. In fine-grained phase, we propose three methods, including pre-training method, multitask method and two-pass method, to encourage the neural model to correctly translate these high-entropy words. Experimental results on various translation tasks show that our method can significantly improve the translation quality and substantially reduce the under-translation cases of high-entropy words. Yang Zhao 0007, Jiajun Zhang 0001, Chengqing Zong, Zhongjun He, Hua Wu 0003 |
AAAI | 5 |
| 2019 | Know More about Each Other: Evolving Dialogue Strategy via Compound AssessmentabstractIn this paper, a novel Generation-Evaluation framework is developed for multi-turn conversations with the objective of letting both participants know more about each other.For the sake of rational knowledge utilization and coherent conversation flow, a dialogue strategy which controls knowledge selection is instantiated and continuously adapted via reinforcement learning.Under the deployed strategy, knowledge grounded conversations are conducted with two dialogue agents.The generated dialogues are comprehensively evaluated on aspects like informativeness and coherence, which are aligned with our objective and human instinct.These assessments are integrated as a compound reward to guide the evolution of dialogue strategy via policy gradient.Comprehensive experiments have been carried out on the publicly available dataset, demonstrating that the proposed method outperforms the other state-of-the-art approaches significantly. Siqi Bao, Huang He, Fan Wang 0021, Rongzhong Lian, Hua Wu 0003 |
ACL (1) | 5 |
| 2019 | ARNOR: Attention Regularization based Noise Reduction for Distant Supervision Relation ClassificationabstractDistant supervision is widely used in relation classification in order to create large-scale training data by aligning a knowledge base with an unlabeled corpus. However, it also introduces amounts of noisy labels where a contextual sentence actually does not express the labeled relation. In this paper, we propose ARNOR, a novel Attention Regularization based NOise Reduction framework for distant supervision relation classification. ARNOR assumes that a trustable relation label should be explained by the neural attention model. Specifically, our ARNOR framework iteratively learns an interpretable model and utilizes it to select trustable instances. We first introduce attention regularization to force the model to pay attention to the patterns which explain the relation labels, so as to make the model more interpretable. Then, if the learned model can clearly locate the relation patterns of a candidate instance in the training set, we will select it as a trustable instance for further training step. According to the experiments on NYT data, our ARNOR framework achieves significant improvements over state-of-the-art methods in both relation classification performance and noise reduction effect. Dai Dai, Xinyan Xiao, Hua Wu 0003 |
ACL (1) | 4 |
| 2019 | STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix FrameworkabstractMingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, Haifeng Wang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Mingbo Ma, Liang Huang 0001, Hao Xiong 0005, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 11 |
| 2019 | Proactive Human-Machine Conversation with Explicit Conversation GoalabstractThough great progress has been made for human-machine conversation, current dialogue system is still in its infancy: it usually converses passively and utters words more as a matter of response, rather than on its own initiatives. In this paper, we take a radical step towards building a human-like conversational agent: endowing it with the ability of proactively leading the conversation (introducing a new topic or maintaining the current topic). To facilitate the development of such conversation systems, we create a new dataset named Konv where one acts as a conversation leader and the other acts as the follower. The leader is provided with a knowledge graph and asked to sequentially change the discussion topics, following the given conversation goal, and meanwhile keep the dialogue as natural and engaging as possible. Konv enables a very challenging task as the model needs to both understand dialogue and plan over the given knowledge graph. We establish baseline results on this dataset (about 270K utterances and 30k dialogues) using several state-of-the-art models. Experimental results show that dialogue models that plan over the knowledge graph can make full use of related knowledge to generate more diverse multi-turn conversations. The baseline systems along with the dataset are publicly available. Wenquan Wu, Xiangyang Zhou, Hua Wu 0003, Xiyuan Zhang 0002, Rongzhong Lian, Haifeng Wang 0001 |
ACL (1) | 4 |
| 2019 | Enhancing Pre-Trained Language Representations with Rich Knowledge for Machine Reading ComprehensionabstractMachine reading comprehension (MRC) is a crucial and challenging task in NLP.Recently, pre-trained language models (LMs), especially BERT, have achieved remarkable success, presenting new state-of-the-art results in MRC.In this work, we investigate the potential of leveraging external knowledge bases (KBs) to further improve BERT for MRC.We introduce KT-NET, which employs an attention mechanism to adaptively select desired knowledge from KBs, and then fuses selected knowledge with BERT to enable context-and knowledgeaware predictions.We believe this would combine the merits of both deep LMs and curated KBs towards better MRC.Experimental results indicate that KT-NET offers significant and consistent improvements over BERT, outperforming competitive baselines on ReCoRD and SQuAD1.1 benchmarks.Notably, it ranks the 1st place on the ReCoRD leaderboard, and is also the best single model on the SQuAD1.1 leaderboard at the time of submission (March 4th, 2019). 1 An Yang, Quan Wang 0002, Jing Liu 0022, Kai Liu 0023, Yajuan Lyu, Hua Wu 0003, Qiaoqiao She, Sujian Li |
ACL (1) | 6 |
| 2019 | Multi-agent Learning for Neural Machine TranslationabstractTianchi Bi, Hao Xiong, Zhongjun He, Hua Wu, Haifeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tianchi Bi, Hao Xiong 0005, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Knowledge Aware Conversation Generation with Explainable Reasoning over Augmented GraphsabstractZhibin Liu, Zheng-Yu Niu, Hua Wu, Haifeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhengyu Niu, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Enhancing Local Feature Extraction with Global Representation for Neural Text ClassificationabstractGuocheng Niu, Hengru Xu, Bolei He, Xinyan Xiao, Hua Wu, Sheng Gao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Guocheng Niu, Hengru Xu, Bolei He, Xinyan Xiao, Hua Wu 0003 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Generating Multiple Diverse Responses with Multi-Mapping and Posterior Mapping SelectionabstractIn human conversation an input post is open to multiple potential responses, which is typically regarded as a one-to-many problem. Promising approaches mainly incorporate multiple latent mechanisms to build the one-to-many relationship. However, without accurate selection of the latent mechanism corresponding to the target response during training, these methods suffer from a rough optimization of latent mechanisms. In this paper, we propose a multi-mapping mechanism to better capture the one-to-many relationship, where multiple mapping modules are employed as latent mechanisms to model the semantic mappings from an input post to its diverse responses. For accurate optimization of latent mechanisms, a posterior mapping selection module is designed to select the corresponding mapping module according to the target response for further optimization. We also introduce an auxiliary matching loss to facilitate the optimization of posterior mapping selection. Empirical results demonstrate the superiority of our model in generating multiple diverse and informative responses over the state-of-the-art methods. Chaotao Chen, Jinhua Peng, Fan Wang 0021, Jun Xu 0027, Hua Wu 0003 |
IJCAI | 5 |
| 2019 | Learning to Select Knowledge for Response Generation in Dialog SystemsabstractEnd-to-end neural models for intelligent dialogue systems suffer from the problem of generating uninformative responses. Various methods were proposed to generate more informative responses by leveraging external knowledge. However, few previous work has focused on selecting appropriate knowledge in the learning process. The inappropriate selection of knowledge could prohibit the model from learning to make full use of the knowledge. Motivated by this, we propose an end-to-end neural model which employs a novel knowledge selection mechanism where both prior and posterior distributions over knowledge are used to facilitate knowledge selection. Specifically, a posterior distribution over knowledge is inferred from both utterances and responses, and it ensures the appropriate selection of knowledge during the training process. Meanwhile, a prior distribution, which is inferred from utterances only, is used to approximate the posterior distribution so that appropriate knowledge can be selected even without responses during the inference process. Compared with the previous work, our model can better incorporate appropriate knowledge in response generation. Experiments on both automatic and human evaluation verify the superiority of our model over previous baselines. Rongzhong Lian, Fan Wang 0021, Jinhua Peng, Hua Wu 0003 |
IJCAI | 5 |
| 2019 | End-to-End Speech Translation with Knowledge DistillationabstractEnd-to-end speech translation (ST), which directly translates from source language speech into target language text, has attracted intensive attentions in recent years.Compared to conventional pipepine systems, end-to-end ST models have advantages of lower latency, smaller model size and less error propagation.However, the combination of speech recognition and text translation in one model is more difficult than each of these two tasks.In this paper, we propose a knowledge distillation approach to improve ST model by transferring the knowledge from text translation model.Specifically, we first train a text translation model, regarded as a teacher model, and then ST model is trained to learn output probabilities from teacher model through knowledge distillation.Experiments on English-French Augmented LibriSpeech and English-Chinese TED corpus show that end-to-end ST is possible to implement on both similar and dissimilar language pairs.In addition, with the instruction of teacher model, end-to-end ST model can gain significant improvements by over 3.5 BLEU points. Yuchen Liu 0007, Hao Xiong 0005, Jiajun Zhang 0001, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
INTERSPEECH | 5 |
| 2019 | An Overview of the 2019 Language and Intelligence Challenge
Quan Wang 0002, Wenquan Wu, Yabing Shi, Wei He 0014, Ying Chen 0011, Yajuan Lyu, Hua Wu 0003 |
NLPCC (2) | 10 |
| 2019 | A Key-Phrase Aware End2end Neural Response Generation Model
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
NLPCC (2) | 4 |
| 2018 | Multi-Channel Encoder for Neural Machine TranslationabstractAttention-based Encoder-Decoder has the effective architecture for neural machine translation (NMT), which typically relies on recurrent neural networks (RNN) to build the blocks that will be lately called by attentive reader during the decoding process. This design of encoder yields relatively uniform composition on source sentence, despite the gating mechanism employed in encoding RNN. On the other hand, we often hope the decoder to take pieces of source sentence at varying levels suiting its own linguistic structure: for example, we may want to take the entity name in its raw form while taking an idiom as a perfectly composed unit. Motivated by this demand, we propose Multi-channel Encoder (MCE), which enhances encoding components with different levels of composition. More specifically, in addition to the hidden state of encoding RNN, MCE takes 1) the original word embedding for raw encoding with no composition, and 2) a particular design of external memory in Neural Turing Machine NTM) for more complex composition, while all three encoding strategies are properly blended during decoding. Empirical study on Chinese-English translation shows that our model can improve by 6.52 BLEU points upon a strong open source NMT system: DL4MT1. On the WMT14 English-French task, our single shallow system achieves BLEU=38.8, comparable with the state-of-the-art deep models. Hao Xiong 0005, Zhongjun He, Xiaoguang Hu, Hua Wu 0003 |
AAAI | 4 |
| 2018 | Multi-Turn Response Selection for Chatbots with Deep Attention Matching NetworkabstractXiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, Hua Wu. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Xiangyang Zhou, Daxiang Dong, Ying Chen 0011, Wayne Xin Zhao, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 8 |
| 2018 | Multi-Passage Machine Reading Comprehension with Cross-Passage Answer VerificationabstractYizhong Wang, Kai Liu, Jing Liu, Wei He, Yajuan Lyu, Hua Wu, Sujian Li, Haifeng Wang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Yizhong Wang, Kai Liu 0023, Jing Liu 0022, Wei He 0014, Yajuan Lyu, Hua Wu 0003, Sujian Li, Haifeng Wang 0001 |
ACL (1) | 6 |
| 2018 | Addressing Troublesome Words in Neural Machine TranslationabstractOne of the weaknesses of Neural Machine Translation (NMT) is in handling lowfrequency and ambiguous words, which we refer as troublesome words.To address this problem, we propose a novel memoryenhanced NMT method.First, we investigate different strategies to define and detect the troublesome words.Then, a contextual memory is constructed to memorize which target words should be produced in what situations.Finally, we design a hybrid model to dynamically access the contextual memory so as to correctly translate the troublesome words.The extensive experiments on Chineseto-English and English-to-German translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling troublesome words. Yang Zhao 0007, Jiajun Zhang 0001, Zhongjun He, Chengqing Zong, Hua Wu 0003 |
EMNLP | 5 |
| 2018 | A New Method of Region Embedding for Text Classification
Chao Qiao, Guocheng Niu, Daren Li, Daxiang Dong, Wei He 0014, Dianhai Yu, Hua Wu 0003 |
ICLR (Poster) | 8 |
| 2017 | An End-to-End Model for Question Answering over Knowledge Base with Cross-Attention Combining Global KnowledgeabstractWith the rapid growth of knowledge bases (KBs) on the web, how to take full advantage of them becomes increasingly important.Question answering over knowledge base (KB-QA) is one of the promising approaches to access the substantial knowledge.Meanwhile, as the neural networkbased (NN-based) methods develop, NNbased KB-QA has already achieved impressive results.However, previous work did not put more emphasis on question representation, and the question is converted into a fixed vector regardless of its candidate answers.This simple representation strategy is not easy to express the proper information in the question.Hence, we present an end-to-end neural network model to represent the questions and their corresponding scores dynamically according to the various candidate answer aspects via cross-attention mechanism.In addition, we leverage the global knowledge inside the underlying KB, aiming at integrating the rich KB information into the representation of the answers.As a result, it could alleviates the out-of-vocabulary (OOV) problem, which helps the crossattention model to represent the question more precisely.The experimental results on WebQuestions demonstrate the effectiveness of the proposed approach. Yanchao Hao, Yuanzhe Zhang, Kang Liu 0001, Shizhu He, Zhanyi Liu, Hua Wu 0003, Jun Zhao 0001 |
ACL (1) | 6 |
| 2016 | Improved Neural Machine Translation with SMT FeaturesabstractNeural machine translation (NMT) conducts end-to-end translation with a source language encoder and a target language decoder, making promising translation performance. However, as a newly emerged approach, the method has some limitations. An NMT system usually has to apply a vocabulary of certain size to avoid the time-consuming training and decoding, thus it causes a serious out-of-vocabulary problem. Furthermore, the decoder lacks a mechanism to guarantee all the source words to be translated and usually favors short translations, resulting in fluent but inadequate translations. In order to solve the above problems, we incorporate statistical machine translation (SMT) features, such as a translation model and an n-gram language model, with the NMT model under the log-linear framework. Our experiments show that the proposed method significantly improves the translation quality of the state-ofthe-art NMT system on Chinese-to-English translation tasks. Our method produces a gain of up to 2.33 BLEU score on NIST open test sets. Wei He 0014, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 3 |
| 2016 | Semi-Supervised Learning for Neural Machine TranslationabstractWhile end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource languages, it is appealing to exploit monolingual corpora to improve NMT. We propose a semi-supervised approach for training NMT models on the concatenation of labeled (parallel corpora) and unlabeled (monolingual corpora) data. The central idea is to reconstruct the monolingual corpora using an autoencoder, in which the source-to-target and target-to-source translation models serve as the encoder and decoder, respectively. Our approach can not only exploit the monolingual corpora of the target language, but also of the source language. Experiments on the Chinese-English dataset show that our approach achieves significant improvements over state-of-the-art SMT and NMT systems. Yong Cheng 0003, Wei Xu 0005, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 5 |
| 2016 | Active Learning for Dependency Parsing with Partial AnnotationabstractZhenghua Li, Min Zhang, Yue Zhang, Zhanyi Liu, Wenliang Chen, Hua Wu, Haifeng Wang. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Zhenghua Li, Min Zhang 0005, Yue Zhang 0004, Zhanyi Liu, Wenliang Chen, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 6 |
| 2016 | Minimum Risk Training for Neural Machine TranslationabstractWe propose minimum risk training for end-to-end neural machine translation.Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable.Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs.Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks. Shiqi Shen, Yong Cheng 0003, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 5 |
| 2016 | "Shall I Be Your Chat Companion?": Towards an Online Human-Computer Conversation SystemabstractTo establish an automatic conversation system between human and computer is regarded as one of the most hardcore problems in computer science. It requires interdisciplinary techniques in information retrieval, natural language processing, and data management, etc. The challenges lie in how to respond like a human, and to maintain a relevant, meaningful, and continuous conversation. The arrival of big data era reveals the feasibility to create such a system empowered by data-driven approaches. We can now organize the conversational data as a chat companion. In this paper, we introduce a chat companion system, which is a practical conversation system between human and computer as a real application. Given the human utterances as queries, our proposed system will respond with corresponding replies retrieved and highly ranked from a massive conversational data repository. Note that 'practical' here indicates effectiveness and efficiency: both issues are important for a real-time system based on a massive data repository. We have two scenarios of single-turn and multi-turn conversations. In our system, we have a base ranking without conversational context information (for single-turn) and a context-aware ranking (for multi-turn). Both rankings can be conducted either by a shallow learning or deep learning paradigm. We combine these two rankings together in optimization. In the experimental setups, we investigate the performance between effectiveness and efficiency for the proposed methods, and we also compare against a series of baselines to demonstrate the advantage of the proposed framework in terms of [email protected], MAP, and nDCG. We present a new angle to launch a practical online conversation system between human and computer. Rui Yan 0001, Yiping Song, Xiangyang Zhou, Hua Wu 0003 |
CIKM | 4 |
| 2016 | Latent Topic EmbeddingabstractTopic modeling and word embedding are two important techniques for deriving latent semantics from data. General-purpose topic models typically work in coarse granularity by capturing word co-occurrence at the document/sentence level. In contrast, word embedding models usually work in much finer granularity by modeling word co-occurrence within small sliding windows. With the aim of deriving latent semantics by considering word co-occurrence at different levels of granularity, we propose a novel model named Latent Topic Embedding (LTE), which seamlessly integrates topic generation and embedding learning in one unified framework. We further propose an efficient Monte Carlo EM algorithm to estimate the parameters of interest. By retaining the individual advantages of topic modeling and word embedding, LTE results in better latent topics and word embedding. Extensive experiments verify the superiority of LTE over the state-of-the-arts. Di Jiang 0004, Lei Shi 0016, Rongzhong Lian, Hua Wu 0003 |
COLING | 4 |
| 2016 | Chinese Poetry Generation with Planning based Neural NetworkabstractChinese poetry generation is a very challenging task in natural language processing. In this paper, we propose a novel two-stage poetry generating method which first plans the sub-topics of the poem according to the user’s writing intent, and then generates each line of the poem sequentially, using a modified recurrent neural network encoder-decoder framework. The proposed planning-based method can ensure that the generated poem is coherent and semantically consistent with the user’s intent. A comprehensive evaluation with human judgments demonstrates that our proposed approach outperforms the state-of-the-art poetry generating methods and the poem quality is somehow comparable to human poets. Wei He 0014, Hua Wu 0003, Wei Li 0176, Haifeng Wang 0001, Enhong Chen |
COLING | 3 |
| 2016 | Multi-view Response Selection for Human-Computer Conversation
Xiangyang Zhou, Daxiang Dong, Hua Wu 0003, Dianhai Yu, Hao Tian 0005, Rui Yan 0001 |
EMNLP | 3 |
| 2016 | Agreement-Based Joint Training for Bidirectional Attention-Based Neural Machine Translation
Yong Cheng 0003, Shiqi Shen, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005 |
IJCAI | 5 |
| 2016 | Learning to Respond with Deep Neural Networks for Retrieval-Based Human-Computer Conversation SystemabstractTo establish an automatic conversation system between humans and computers is regarded as one of the most hardcore problems in computer science, which involves interdisciplinary techniques in information retrieval, natural language processing, artificial intelligence, etc. The challenges lie in how to respond so as to maintain a relevant and continuous conversation with humans. Along with the prosperity of Web 2.0, we are now able to collect extremely massive conversational data, which are publicly available. It casts a great opportunity to launch automatic conversation systems. Owing to the diversity of Web resources, a retrieval-based conversation system will be able to find at least some responses from the massive repository for any user inputs. Given a human issued message, i.e., query, our system would provide a reply after adequate training and learning of how to respond. In this paper, we propose a retrieval-based conversation system with the deep learning-to-respond schema through a deep neural network framework driven by web data. The proposed model is general and unified for different conversation scenarios in open domain. We incorporate the impact of multiple data inputs, and formulate various features and factors with optimization into the deep learning framework. In the experiments, we investigate the effectiveness of the proposed deep neural network structures with better combinations of all different evidence. We demonstrate significant performance improvement against a series of standard and state-of-art baselines in terms of [email protected], MAP, nDCG, and MRR for conversational purposes. Rui Yan 0001, Yiping Song, Hua Wu 0003 |
SIGIR | 3 |
| 2015 | Multi-Task Learning for Multiple Language TranslationabstractDaxiang Dong, Hua Wu, Wei He, Dianhai Yu, Haifeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Daxiang Dong, Hua Wu 0003, Wei He 0014, Dianhai Yu, Haifeng Wang 0001 |
ACL (1) | 2 |
| 2015 | Improved beam search with constrained softmax for NMT
Xiaoguang Hu, Wei Li 0176, Hua Wu 0003, Haifeng Wang 0001 |
MTSummit | 4 |
| 2015 | Exploiting Collective Hidden Structures in Webpage Titles for Open Domain Entity ExtractionabstractWe present a novel method for open domain named entity extraction by exploiting the collective hidden structures in webpage titles. Our method uncovers the hidden textual structures shared by sets of webpage titles based on generalized URL patterns and a multiple sequence alignment technique. The highlights of our method include: 1) The boundaries of entities can be identified automatically in a collective way without any manually designed pattern, seed or class name. 2) The connections between entities are also discovered naturally based on the hidden structures, which makes it easy to incorporate distant or weak supervision. The experiments show that our method can harvest large scale of open domain entities with high precision. A large ratio of the extracted entities are long-tailed and complex and cover diverse topics. Given the extracted entities and their connections, we further show the effectiveness of our method in a weakly supervised setting. Our method can produce better domain specific entities in both precision and recall compared with the state-of-the-art approaches. Wei Song 0010, Hua Wu 0003, Haifeng Wang 0001, Lizhen Liu, Hanshi Wang |
WWW | 4 |
| 2014 | Transformation from Discontinuous to Continuous Word Alignment Improves Translation QualityabstractWe present a novel approach to improve word alignment for statistical machine translation (SMT).Conventional word alignment methods allow discontinuous alignment, meaning that a source (or target) word links to several target (or source) words whose positions are discontinuous.However, we cannot extract phrase pairs from this kind of alignments as they break the alignment consistency constraint.In this paper, we use a weighted vote method to transform discontinuous word alignment to continuous alignment, which enables SMT systems extract more phrase pairs.We carry out experiments on large scale Chineseto-English and German-to-English translation tasks.Experimental results show statistically significant improvements of BLEU score in both cases over the baseline systems.Our method produces a gain of +1.68 BLEU on NIST OpenMT04 for the phrase-based system, and a gain of +1.28 BLEU on NIST OpenMT06 for the hierarchical phrase-based system. Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Ting Liu 0001 |
EMNLP | 2 |
| 2014 | Policy Learning for Domain Selection in an Extensible Multi-domain Spoken Dialogue SystemabstractThis paper proposes a Markov Decision Process and reinforcement learning based approach for domain selection in a multidomain Spoken Dialogue System built on a distributed architecture.In the proposed framework, the domain selection problem is treated as sequential planning instead of classification, such that confirmation and clarification interaction mechanisms are supported.In addition, it is shown that by using a model parameter tying trick, the extensibility of the system can be preserved, where dialogue components in new domains can be easily plugged in, without re-training the domain selection policy.The experimental results based on human subjects suggest that the proposed model marginally outperforms a non-trivial baseline. Guanchun Wang, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP | 5 |
| 2014 | Improve Statistical Machine Translation with Context-Sensitive Bilingual Semantic Embedding ModelabstractWe investigate how to improve bilingual embedding which has been successfully used as a feature in phrase-based statistical machine translation (SMT). Despite bilingual embedding’s success, the contextual information, which is of critical importance to translation quality, was ignored in previous work. To employ the contextual information, we propose a simple and memory-efficient model for learning bilingual embedding, taking both the source phrase and context around the phrase into account. Bilingual translation scores generated from our proposed bilingual embedding model are used as features in our SMT system. Experimental results show that the proposed method achieves significant improvements on large-scale Chinese-English translation task. Daxiang Dong, Xiaoguang Hu, Dianhai Yu, Wei He 0014, Hua Wu 0003, Haifeng Wang 0001, Ting Liu 0001 |
EMNLP | 6 |
| 2014 | Improving Pivot-Based Statistical Machine Translation by Pivoting the Co-occurrence Count of Phrase PairsabstractTo overcome the scarceness of bilingual corpora for some language pairs in machine translation, pivot-based SMT uses pivot language as a "bridge" to generate source-target translation from sourcepivot and pivot-target translation.One of the key issues is to estimate the probabilities for the generated phrase pairs.In this paper, we present a novel approach to calculate the translation probability by pivoting the co-occurrence count of source-pivot and pivot-target phrase pairs.Experimental results on Europarl data and web data show that our method leads to significant improvements over the baseline systems. Zhongjun He, Hua Wu 0003, Conghui Zhu, Haifeng Wang 0001, Tiejun Zhao |
EMNLP | 3 |
| 2013 | Improving Pivot-Based Statistical Machine Translation Using Random WalkabstractThis paper proposes a novel approach that utilizes a machine learning method to improve pivot-based statistical machine translation (SMT).For language pairs with few bilingual data, a possible solution in pivot-based SMT using another language as a "bridge" to generate source-target translation.However, one of the weaknesses is that some useful sourcetarget translations cannot be generated if the corresponding source phrase and target phrase connect to different pivot phrases.To alleviate the problem, we utilize Markov random walks to connect possible translation phrases between source and target language.Experimental results on European Parliament data, spoken language data and web data show that our method leads to significant improvements on all the tasks over the baseline system. Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Conghui Zhu, Tiejun Zhao |
EMNLP | 3 |
| 2012 | Improve SMT Quality with Automatically Extracted Paraphrase Rules
Wei He 0014, Hua Wu 0003, Haifeng Wang 0001, Ting Liu 0001 |
ACL (1) | 2 |
| 2012 | Translation Model Adaptation for Statistical Machine Translation with Monolingual Topic Information
Jinsong Su, Hua Wu 0003, Haifeng Wang 0001, Yidong Chen 0001, Xiaodong Shi, Huailin Dong, Qun Liu 0001 |
ACL (1) | 2 |
| 2011 | Reordering with Source Language Collocations
Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003, Ting Liu 0001, Sheng Li 0003 |
ACL | 3 |
| 2011 | Two-Word Collocation Extraction Using Monolingual Word Alignment MethodabstractStatistical bilingual word alignment has been well studied in the field of machine translation. This article adapts the bilingual word alignment algorithm into a monolingual scenario to extract collocations from monolingual corpus, based on the fact that the words in a collocation tend to co-occur in similar contexts as in bilingual word alignment. First, the monolingual corpus is replicated to generate a parallel corpus, in which each sentence pair consists of two identical sentences. Next, the monolingual word alignment algorithm is employed to align potentially collocated words. Finally, the aligned word pairs are ranked according to the alignment scores and candidates with higher scores are extracted as collocations. We conducted experiments on Chinese and English corpora respectively. Compared to previous approaches that use association measures to extract collocations from co-occurrence word pairs within a given window, our method achieves higher precision and recall. According to human evaluation, our method achieves precisions of 62% on a Chinese corpus and 64% on an English corpus. In particular, we can extract collocations with longer spans, achieving a higher precision of 83% on the long-span (> 6 words) Chinese collocations. Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003, Sheng Li 0003 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2010 | Improving Statistical Machine Translation with Monolingual Collocation
Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003, Sheng Li 0003 |
ACL | 3 |
| 2009 | Exploiting Heterogeneous Treebanks for Parsing
Zhengyu Niu, Haifeng Wang 0001, Hua Wu 0003 |
ACL/IJCNLP | 3 |
| 2009 | Revisiting Pivot Language Approach for Machine Translation
Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP | 1 |
| 2009 | Collocation Extraction Using Monolingual Word Alignment Method
Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003, Sheng Li 0003 |
EMNLP | 3 |
| 2008 | Domain Adaptation for Statistical Machine Translation with Domain Dictionary and Monolingual Corpora
Hua Wu 0003, Haifeng Wang 0001, Chengqing Zong |
COLING | 1 |
| 2007 | Pivot Language Approach for Phrase-Based Statistical Machine Translation
Hua Wu 0003, Haifeng Wang 0001 |
ACL | 1 |
| 2007 | Using RBMT Systems to Produce Bilingual Corpus for SMT
Xiaoguang Hu, Haifeng Wang 0001, Hua Wu 0003 |
EMNLP-CoNLL | 3 |
| 2007 | Comparative study of word alignment heuristics and phrase-based SMT
Hua Wu 0003, Haifeng Wang 0001 |
MTSummit | 1 |
| 2007 | Log-linear generation models for example-based machine translation
Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003 |
MTSummit | 3 |
| 2007 | Improving statistical word alignment with various clues
Dengjun Ren, Hua Wu 0003, Haifeng Wang 0001 |
MTSummit | 2 |
| 2007 | Pivot language approach for phrase-based statistical machine translation
Hua Wu 0003, Haifeng Wang 0001 |
Mach. Transl. | 1 |
| 2006 | Word Alignment for Languages with Scarce Resources Using Bilingual Corpora of Other Language Pairs
Haifeng Wang 0001, Hua Wu 0003, Zhanyi Liu |
ACL | 2 |
| 2006 | Boosting Statistical Word Alignment Using Labeled and Unlabeled Data
Hua Wu 0003, Haifeng Wang 0001, Zhanyi Liu |
ACL | 1 |
| 2006 | Example-based machine translation based on tree-string correspondence and statistical generation
Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003 |
Mach. Transl. | 3 |
| 2005 | Alignment Model Adaptation for Domain-Specific Word AlignmentabstractThis paper proposes an alignment adaptation approach to improve domain-specific (in-domain) word alignment. The basic idea of alignment adaptation is to use out-of-domain corpus to improve in-domain word alignment results. In this paper, we first train two statistical word alignment models with the large-scale out-of-domain corpus and the small-scale in-domain corpus respectively, and then interpolate these two models to improve the domain-specific word alignment. Experimental results show that our approach improves domain-specific word alignment in terms of both precision and recall, achieving a relative error rate reduction of 6.56% as compared with the state-of-the-art technologies. Hua Wu 0003, Haifeng Wang 0001, Zhanyi Liu |
ACL | 1 |
| 2005 | Improving Statistical Word Alignment with Ensemble Methods
Hua Wu 0003, Haifeng Wang 0001 |
IJCNLP | 1 |
| 2005 | Example-based Machine Translation Based on TSC and Statistical GenerationabstractThis paper proposes a novel Example-Based Machine Translation (EBMT) method based on Tree String Correspondence (TSC) and statistical generation. In this method, the translation examples are represented as TSC, which consists of three parts: a parse tree in the source language, a string in the target language, and the correspondences between the leaf nodes of the source language tree and the substrings of the target language string. During the translation, the input sentence is first parsed into a tree. Then the TSC forest is searched out if it is best matched with the parse tree. The translation is generated by using a statistical generation model to combine the target language strings in the TSCs. The generation model consists of three parts: the semantic similarity between words, the word translation probability, and the target language model. Based on the above method, we build an English-to-Chinese Machine Translation (ECMT) system. Experimental results indicate that the performance of our system is comparable with that of the state-of-the-art commercial ECMT systems. Zhanyi Liu, Haifeng Wang 0001, Hua Wu 0003 |
MTSummit | 3 |
| 2005 | Boosting Statistical Word AlignmentabstractThis paper proposes an approach to improve statistical word alignment with the boosting method. Applying boosting to word alignment must solve two problems. The first is how to build the reference set for the training data. We propose an approach to automatically build a pseudo reference set, which can avoid manual annotation of the training set. The second is how to calculate the error rate of each individual word aligner. We solve this by calculating the error rate of a manually annotated held-out data set instead of the entire training set. In addition, the final ensemble takes into account the weights of the alignment links produced by the individual word aligners. Experimental results indicate that the boosting method proposed in this paper performs much better than the original word aligner, achieving a large error rate reduction. Hua Wu 0003, Haifeng Wang 0001 |
MTSummit | 1 |
| 2004 | Improving Statistical Word Alignment with a Rule-Based Machine Translation System
Hua Wu 0003, Haifeng Wang 0001 |
COLING | 1 |