Yan Wang 0060

dblp:59/2227-60 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
abstract
Large language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms.
Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001
ACL (1)2
2024 Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
abstract
Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate.
Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu
EMNLP5
2024 Exploring Dense Retrieval for Dialogue Response Selection
abstract
Recent progress in deep learning has continuously improved the accuracy of dialogue response selection. However, in real-world scenarios, the high computation cost forces existing dialogue response selection models to rank only a small number of candidates, recalled by a coarse-grained model, precluding many high-quality candidates. To overcome this problem, we present a novel and efficient response selection model and a set of tailor-designed learning strategies to train it effectively. The proposed model consists of a dense retrieval module and an interaction layer, which could directly select the proper response from a large corpus. We conduct re-rank and full-rank evaluations on widely used benchmarks to evaluate our proposed model. Extensive experimental results demonstrate that our proposed model notably outperforms the state-of-the-art baselines on both re-rank and full-rank evaluations. Moreover, human evaluation results show that the response quality could be improved further by enlarging the candidate pool with nonparallel corpora. In addition, we also release high-quality benchmarks that are carefully annotated for more accurate dialogue response selection evaluation. All source codes, datasets, model parameters, and other related resources have been publicly available. 1
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Yixuan Su, Heyan Huang, Xianling Mao
ACM Trans. Inf. Syst.3
2023 Generalizing Math Word Problem Solvers via Solution Diversification
abstract
Current math word problem (MWP) solvers are usually Seq2Seq models trained by the (one-problem; one-solution) pairs, each of which is made of a problem description and a solution showing reasoning flow to get the correct answer. However, one MWP problem naturally has multiple solution equations. The training of an MWP solver with (one-problem; one-solution) pairs excludes other correct solutions, and thus limits the generalizability of the MWP solver. One feasible solution to this limitation is to augment multiple solutions to a given problem. However, it is difficult to collect diverse and accurate augment solutions through human efforts. In this paper, we design a new training framework for an MWP solver by introducing a solution buffer and a solution discriminator. The buffer includes solutions generated by an MWP solver to encourage the training data diversity. The discriminator controls the quality of buffered solutions to participate in training. Our framework is flexibly applicable to a wide setting of fully, semi-weakly and weakly supervised training for all Seq2Seq MWP solvers. We conduct extensive experiments on a benchmark dataset Math23k and a new dataset named Weak12k, and show that our framework improves the performance of various MWP solvers under different settings by generating correct and diverse solutions.
Zhenwen Liang, Lei Wang 0185, Yan Wang 0060, Jie Shao 0001, Xiangliang Zhang 0001
AAAI4
2023 Copy is All You Need
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Heyan Huang, Xianling Mao
ICLR3
2023 Predicting Events in MOBA Games: Prediction, Attribution, and Evaluation
abstract
The multiplayer online battle arena (MOBA) games have become increasingly popular in recent years. Consequently, many efforts have been devoted to providing pregame or in-game predictions for them. These predictions can be used in many MOBA esports-related applications, such as artificial intelligence commentator systems, in-game data analysis, and game-assistant bots. However, these works are limited in the following two aspects: the lack of sufficient in-game features and the absence of interpretability in the prediction results. These two limitations greatly restrict the practical performance and industrial application of the current works. In this work, we collect a large-scale dataset containing rich in-game features for the popular MOBA gameHonor of Kings. We then propose to predict four types of prediction tasks in an interpretable way by attributing the predictions to the input features using two gradient-based attribution methods:Integrated GradientsandSmoothGrad. To evaluate the explanatory power of different models and attribution methods, a fidelity-based evaluation metric is further proposed. Finally, we evaluate the accuracy and fidelity of several competitive methods to assess how well machines predict events in MOBA games.
Zelong Yang 0002, Yan Wang 0060, Piji Li, Shaobin Lin, Shuming Shi 0001, Shao-Lun Huang, Wei Bi
IEEE Trans. Games2
2022 MWPToolkit: An Open-Source Framework for Deep Learning-Based Math Word Problem Solvers
abstract
While Math Word Problem (MWP) solving has emerged as a popular field of study and made great progress in recent years, most existing methods are benchmarked solely on one or two datasets and implemented with different configurations. In this paper, we introduce the first open-source library for solving MWPs called MWPToolkit, which provides a unified, comprehensive, and extensible framework for the research purpose. Specifically, we deploy 17 deep learning-based MWP solvers and 6 MWP datasets in our toolkit. These MWP solvers are advanced models for MWP solving, covering the categories of Seq2seq, Seq2Tree, Graph2Tree, and Pre-trained Language Models. And these MWP datasets are popular datasets that are commonly used as benchmarks in existing work. Our toolkit is featured with highly modularized and reusable components, which can help researchers quickly get started and develop their own models. We have released the code and documentation of MWPToolkit in https://github.com/LYH-YF/MWPToolkit.
Yihuai Lan, Lei Wang 0185, Yunshi Lan, Bing Tian Dai, Yan Wang 0060, Dongxiang Zhang, Ee-Peng Lim
AAAI6
2022 Automatic Prosody Annotation with Pre-Trained Text-Speech Model
abstract
Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability.However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming.In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders.This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}.The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.Code is released 1 .
Ziqian Dai, Jianwei Yu 0001, Yan Wang 0060, Nuo Chen 0001, Yanyao Bian, Guangzhi Li, Deng Cai 0002, Dong Yu 0001
INTERSPEECH3
2022 ASR-Robust Natural Language Understanding on ASR-GLUE dataset
Lingyun Feng, Jianwei Yu 0001, Yan Wang 0060, Songxiang Liu, Deng Cai 0002, Hai-Tao Zheng 0002
INTERSPEECH3
2022 A Contrastive Framework for Neural Text Generation
abstract
Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease the probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method---contrastive search---to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach outperforms state-of-the-art text generation methods as evaluated by both human and automatic metrics.
Yixuan Su, Tian Lan 0003, Yan Wang 0060, Dani Yogatama, Lingpeng Kong, Nigel Collier
NeurIPS3
2022 Recent Advances in Retrieval-Augmented Text Generation
abstract
Recently retrieval-augmented text generation has achieved state-of-the-art performance in many NLP tasks and has attracted increasing attention of the NLP and IR community, this tutorial thereby aims to present recent advances in retrieval-augmented text generation comprehensively and comparatively. It firstly highlights the generic paradigm of retrieval-augmented text generation, then reviews notable works for different text generation tasks including dialogue generation, machine translation, and other generation tasks, and finally points out some limitations and shortcomings to facilitate future research.
Deng Cai 0002, Yan Wang 0060, Lemao Liu, Shuming Shi 0001
SIGIR2
2022 Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA Esport
abstract
With the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability.
Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu
IEEE Trans. Games3
2021 Neural Machine Translation with Monolingual Translation Memory
abstract
Deng Cai, Yan Wang, Huayang Li, Wai Lam, Lemao Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Deng Cai 0002, Yan Wang 0060, Wai Lam, Lemao Liu
ACL/IJCNLP (1)2
2021 BoB: BERT Over BERT for Training Persona-based Dialogue Models from Limited Personalized Data
abstract
Haoyu Song, Yan Wang, Kaiyan Zhang, Wei-Nan Zhang, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haoyu Song 0002, Yan Wang 0060, Weinan Zhang 0003, Ting Liu 0001
ACL/IJCNLP (1)2
2021 Dialogue Response Selection with Hierarchical Curriculum Learning
abstract
Yixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060
ACL/IJCNLP (1)9
2021 Non-Autoregressive Text Generation with Pre-trained Language Models
abstract
Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, Nigel Collier. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Yixuan Su, Deng Cai 0002, Yan Wang 0060, David Vandyke, Simon Baker, Piji Li, Nigel Collier
EACL3
2021 PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval Memory
abstract
The ability of dialogue systems to express pre-specified style during conversations has a direct, positive impact on their usability and user satisfaction. While it has attracted much research interest, existing methods often generate stylistic responses at the cost of content quality. In this work, we introduce a prototype-to-style (PS) framework to tackle the challenge of stylistic dialogue generation. The proposed framework first exploits an Information Retrieval (IR) system and extracts a response prototype from the retrieved response. A stylistic response generator then takes the response prototype and the desired style as input to produce a high-quality and stylistic response. To effectively train the proposed model and imitate the real testing environment, we introduce a new style-aware learning objective and a denoising learning strategy. Results on three benchmark datasets (gender, emotion, and sentiment) from two languages demonstrate that the proposed approach significantly outperforms existing baselines both in terms of in-domain and cross-domain evaluations.
Yixuan Su, Yan Wang 0060, Deng Cai 0002, Simon Baker, Anna Korhonen, Nigel Collier
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Generate, Delete and Rewrite: A Three-Stage Framework for Improving Persona Consistency of Dialogue Generation
abstract
Maintaining a consistent personality in conversations is quite natural for human beings, but is still a non-trivial task for machines.The persona-based dialogue generation task is thus introduced to tackle the personalityinconsistent problem by incorporating explicit persona text into dialogue generation models.Despite the success of existing personabased models on generating human-like responses, their one-stage decoding framework can hardly avoid the generation of inconsistent persona words.In this work, we introduce a three-stage framework that employs a generate-delete-rewrite mechanism to delete inconsistent words from a generated response prototype and further rewrite it to a personality-consistent one.We carry out evaluations by both human and automatic metrics.Experiments on the Persona-Chat dataset show that our approach achieves good performance.
Haoyu Song 0002, Yan Wang 0060, Weinan Zhang 0003, Xiaojiang Liu, Ting Liu 0001
ACL2
2020 Graph-to-Tree Learning for Solving Math Word Problems
abstract
While the recent tree-based neural models have demonstrated promising results in generating solution expression for the math word problem (MWP), most of these models do not capture the relationships and order information among the quantities well.This results in poor quantity representations and incorrect solution expressions.In this paper, we propose Graph2Tree, a novel deep learning architecture that combines the merits of the graph-based encoder and tree-based decoder to generate better solution expressions.Included in our Graph2Tree framework are two graphs, namely the Quantity Cell Graph and Quantity Comparison Graph, which are designed to address limitations of existing methods by effectively representing the relationships and order information among the quantities in MWPs.We conduct extensive experiments on two available datasets.Our experiment results show that Graph2Tree outperforms the state-of-the-art baselines on two benchmark datasets significantly.We also discuss case studies and empirically examine Graph2Tree's effectiveness in translating the MWP text into solution expressions 1 .
Lei Wang 0185, Roy Ka-Wei Lee, Yi Bin, Yan Wang 0060, Jie Shao 0001, Ee-Peng Lim
ACL5
2020 The World is Not Binary: Learning to Rank with Grayscale Data for Dialogue Response Selection
abstract
Response selection plays a vital role in building retrieval-based conversation systems.Despite that response selection is naturally a learning-to-rank problem, most prior works take a point-wise view and train binary classifiers for this task: each response candidate is labeled either relevant (one) or irrelevant (zero).On the one hand, this formalization can be sub-optimal due to its ignorance of the diversity of response quality.On the other hand, annotating grayscale data for learning-to-rank can be prohibitively expensive and challenging.In this work, we show that grayscale data can be automatically constructed without human effort.Our method employs off-the-shelf response retrieval models and response generation models as automatic grayscale data generators.With the constructed grayscale data, we propose multi-level ranking objectives for training, which can (1) teach a matching model to capture more fine-grained context-response relevance difference and (2) reduce the traintest discrepancy in terms of distractor strength.Our method is simple, effective, and universal.Experiments on three benchmark datasets and four state-of-the-art matching models show that the proposed approach brings significant and consistent performance improvements.
Zibo Lin, Deng Cai 0002, Yan Wang 0060, Xiaojiang Liu, Hai-Tao Zheng 0002, Shuming Shi 0001
EMNLP (1)3
2020 Profile Consistency Identification for Open-domain Dialogue Agents
abstract
Maintaining a consistent attribute profile is crucial for dialogue agents to naturally converse with humans.Existing studies on improving attribute consistency mainly explored how to incorporate attribute information in the responses, but few efforts have been made to identify the consistency relations between response and attribute profile.To facilitate the study of profile consistency identification, we create a large-scale human-annotated dataset with over 110K single-turn conversations and their key-value attribute profiles.Explicit relation between response and profile is manually labeled.We also propose a key-value structure information enriched BERT model to identify the profile consistency, and it gained improvements over strong baselines.Further evaluations on downstream tasks demonstrate that the profile consistency identification model is conducive for improving dialogue consistency.
Haoyu Song 0002, Yan Wang 0060, Weinan Zhang 0003, Zhengyu Zhao 0003, Ting Liu 0001, Xiaojiang Liu
EMNLP (1)2
2019 Better Fine-Tuning via Instance Weighting for Text Classification
abstract
Transfer learning for deep neural networks has achieved great success in many text classification applications. A simple yet effective transfer learning method is to fine-tune the pretrained model parameters. Previous fine-tuning works mainly focus on the pre-training stage and investigate how to pretrain a set of parameters that can help the target task most. In this paper, we propose an Instance Weighting based Finetuning (IW-Fit) method, which revises the fine-tuning stage to improve the final performance on the target domain. IW-Fit adjusts instance weights at each fine-tuning epoch dynamically to accomplish two goals: 1) identify and learn the specific knowledge of the target domain effectively; 2) well preserve the shared knowledge between the source and the target domains. The designed instance weighting metrics used in IW-Fit are model-agnostic, which are easy to implement for general DNN-based classifiers. Experimental results show that IW-Fit can consistently improve the classification accuracy on the target domain.
Wei Bi, Yan Wang 0060, Xiaojiang Liu
AAAI3
2019 Modeling Intra-Relation in Math Word Problems with Different Functional Multi-Head Attentions
abstract
Several deep learning models have been proposed for solving math word problems (MWPs) automatically.Although these models have the ability to capture features without manual efforts, their approaches to capturing features are not specifically designed for MWPs.To utilize the merits of deep learning models with simultaneous consideration of MWPs' specific features, we propose a group attention mechanism to extract global features, quantity-related features, quantitypair features and question-related features in MWPs respectively.The experimental results show that the proposed approach performs significantly better than previous state-of-the-art methods, and boost performance from 66.9% to 69.5% on Math23K with training-test split, from 65.8% to 66.9% on Math23K with 5-fold cross-validation and from 69.2% to 76.1% on MAWPS.
Jierui Li, Lei Wang 0185, Yan Wang 0060, Bing Tian Dai, Dongxiang Zhang
ACL (1)4
2019 Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework
abstract
Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Deng Cai 0002, Yan Wang 0060, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi 0001
EMNLP/IJCNLP (1)2
2019 Improving Open-Domain Dialogue Systems via Multi-Turn Incomplete Utterance Restoration
abstract
Zhufeng Pan, Kun Bai, Yan Wang, Lianqiang Zhou, Xiaojiang Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhufeng Pan, Yan Wang 0060, Lianqiang Zhou, Xiaojiang Liu
EMNLP/IJCNLP (1)3
2018 Translating Math Word Problem to Expression Tree
abstract
Sequence-to-sequence (SEQ2SEQ) models have been successfully applied to automatic math word problem solving.Despite its simplicity, a drawback still remains: a math word problem can be correctly solved by more than one equations.This non-deterministic transduction harms the performance of maximum likelihood estimation.In this paper, by considering the uniqueness of expression tree, we propose an equation normalization method to normalize the duplicated equations.Moreover, we analyze the performance of three popular SEQ2SEQ models on the math word problem solving.We find that each model has its own specialty in solving problems, consequently an ensemble model is then proposed to combine their advantages.Experiments on dataset Math23K show that the ensemble model with equation normalization significantly outperforms the previous state-of-the-art methods.
Lei Wang 0185, Yan Wang 0060, Deng Cai 0002, Dongxiang Zhang, Xiaojiang Liu
EMNLP2
2017 Deep Neural Solver for Math Word Problems
abstract
This paper presents a deep neural solver to automatically solve math word problems.In contrast to previous statistical learning approaches, we directly translate math word problems to equation templates using a recurrent neural network (RNN) model, without sophisticated feature engineering.We further design a hybrid model that combines the RNN model and a similarity-based retrieval model to achieve additional performance improvement.Experiments conducted on a large dataset show that the RNN model and the hybrid model significantly outperform stateof-the-art statistical learning methods for math word problem solving. 1 We plan to make the dataset publicly available when the paper is published
Yan Wang 0060, Xiaojiang Liu, Shuming Shi 0001
EMNLP1