Wei Bi

dblp:38/1163 · DBLP profile ↗
← Back
70ranked-venue papers
10as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 9 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing
abstract
Qian Cao, Yahui Liu, Wei Bi, Yi Zhao, Ruihua Song, Xiting Wang, Ruiming Tang, Guorui Zhou, Han Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qian Cao 0001, Wei Bi, Ruihua Song, Xiting Wang, Ruiming Tang, Guorui Zhou, Han Li 0005
ACL (1)3
2026 v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
abstract
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang 0004, Qinrong Cui, Wei Bi, Song-Chun Zhu, Zilong Zheng
ACL (1)6
2026 A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
abstract
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable unknown (unk) position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (e.g., ImageNet-1 K) and sentiment analysis for text (e.g., Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks.
Weixin Ye, Wei Wang 0108, Yue Song 0002, Bin Ren 0005, Wei Bi, Rita Cucchiara, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
abstract
Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether LLMs can serve as reliable alternatives to humans, we examine the fine-grained alignment between LLM evaluators and human annotators, particularly in understanding the target evaluation tasks and conducting evaluations that meet diverse criteria. This paper explores both conventional tasks (e.g., story generation) and alignment tasks (e.g., math reasoning), each with different evaluation criteria. Our analysis shows that 1) LLM evaluators can generate unnecessary criteria or omit crucial criteria, resulting in a slight deviation from the experts. 2) LLM evaluators excel in general criteria, such as fluency, but face challenges with complex criteria, such as numerical reasoning. We also find that LLM-pre-drafting before human evaluation can help reduce the impact of human subjectivity and minimize annotation outliers in pure human evaluation, leading to more objective evaluation. All resources are available at https://github.com/qtli/CoEval.
Qintong Li, Leyang Cui, Lingpeng Kong, Wei Bi
COLING4
2025 What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
abstract
Recent advances in reasoning with large language models (LLMs) have popularized Long Chain-of-Thought (LCoT), a strategy that encourages deliberate and step-by-step reasoning before producing a final answer.While LCoTs have enabled expert-level performance in complex tasks, how the internal structures of their reasoning chains drive, or even predict, the correctness of final answers remains a critical yet underexplored question.In this work, we present LCoT2Tree, an automated framework that converts sequential LCoTs into hierarchical tree structures and thus enables deeper structural analysis of LLM reasoning.Using graph neural networks (GNNs), we reveal that structural patterns extracted by LCoT2Tree, including exploration, backtracking, and verification, serve as stronger predictors of final performance across a wide range of tasks and models.Leveraging an explainability technique, we further identify critical thought patterns such as over-branching that account for failures.Beyond diagnostic insights, the structural patterns by LCoT2Tree support practical applications, including improving Best-of-N decoding effectiveness.Overall, our results underscore the critical role of internal structures of reasoning chains, positioning LCoT2Tree as a powerful tool for diagnosing, interpreting, and improving reasoning in LLMs.The code is available at our GitHub repository.
Gangwei Jiang, Wei Bi, Linqi Song, Ying Wei 0001, Defu Lian
EMNLP4
2025 Scaling Diffusion Language Models via Adaptation from Autoregressive Models
abstract
Diffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions.
Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong
ICLR9
2025 Decider: A Dual-System Rule-Controllable Decoding Framework for Language Generation
abstract
Constrained decoding approaches aim to control the meaning or style of text generated by a Pre-trained Language Model (PLM) for various task-specific objectives at inference time. However, these methods often guide plausible continuations by greedily and explicitly selecting targets, which, while fulfilling the task requirements, may overlook the natural patterns of human language generation. In this work, we propose a novel decoding framework,Decider, which enables us to program high-level rules on how we might effectively complete tasks to control a PLM. Differing from previous works, our framework transforms the encouragement of concrete target words into the encouragement of all words that satisfy the high-level rules. Specifically,Decideris a dual system in which a PLM is equipped and controlled by a First-Order Logic (FOL) reasoner to express and evaluate the rules, along with a decision function that merges the outputs from both systems to guide the generation. Experiments on CommonGen and PersonaChat demonstrate thatDecidercan effectively follow given rules to guide a PLM in achieving generation tasks in a more human-like manner.
Tian Lan 0003, Changlong Yu, Wei Wang 0138, Qunxi Dong, Kun Qian 0003, Piji Li, Wei Bi, Bin Hu 0001
IEEE Trans. Knowl. Data Eng.10
2024 GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
abstract
Large language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks.However, there are increasing debates regarding whether these models truly understand and apply mathematical knowledge or merely rely on shortcuts for mathematical reasoning.One essential and frequently occurring evidence is that when the math questions are slightly changed, LLMs can behave incorrectly.This motivates us to evaluate the robustness of LLMs' math reasoning capability by testing a wide range of question variations.We introduce the adversarial grade school math (GSM-PLUS) dataset, an extension of GSM8K augmented with various mathematical perturbations.Our experiments on 25 LLMs and 4 prompting techniques show that while LLMs exhibit different levels of math reasoning abilities, their performances are far from robust.In particular, even for problems that have been solved in GSM8K, LLMs can make mistakes when new statements are added or the question targets are altered.We also explore whether more robust performance can be achieved by composing existing prompting methods, in which we try an iterative method that generates and verifies each intermediate thought based on its reasoning goal and calculation result.
Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong, Wei Bi
ACL (1)5
2024 MAGE: Machine-generated Text Detection in the Wild
abstract
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004
ACL (1)4
2024 SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving
abstract
Large Language Models (LLMs) have driven substantial progress in artificial intelligence in recent years, exhibiting impressive capabilities across a wide range of tasks, including mathematical problem-solving.Inspired by the success of subgoal-based methods, we propose a novel framework called SEquential subGoal Optimization (SEGO) to enhance LLMs' ability to solve mathematical problems.By establishing a connection between the subgoal breakdown process and the probability of solving problems, SEGO aims to identify better subgoals with theoretical guarantees.Addressing the challenge of identifying suitable subgoals in a large solution space, our framework generates problem-specific subgoals and adjusts them according to carefully designed criteria.Incorporating these optimized subgoals into the policy model training leads to significant improvements in problem-solving performance.We validate SEGO's efficacy through experiments on two benchmarks, GSM8K and MATH, where our approach outperforms existing methods, highlighting the potential of SEGO in AI-driven mathematical problemsolving. * This work
Xueliang Zhao, Xinting Huang, Wei Bi, Lingpeng Kong
ACL (1)3
2024 A Frustratingly Simple Decoding Method for Neural Text Generation
abstract
We introduce a frustratingly simple, highly efficient, and surprisingly effective decoding method, termed Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: We construct an anti-language model (anti-LM) based on previously generated text, which is employed to penalize the future generation of repetitive content. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD incurs no additional model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite its simplicity, FSD is surprisingly effective and generalizes across different datasets, models, and languages. Extensive experiments show that FSD outperforms established strong baselines in terms of generation quality, decoding speed, and universality.
Deng Cai 0002, Wei Bi, Wai Lam, Shuming Shi 0001
LREC/COLING4
2024 Knowledge Verification to Nip Hallucination in the Bud
abstract
While large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.In this paper, we demonstrate the feasibility of mitigating hallucinations by verifying and minimizing the inconsistency between external knowledge present in the alignment data and the intrinsic knowledge embedded within foundation LLMs.Specifically, we propose a novel approach called Knowledge Consistent Alignment (KCA), which employs a well-aligned LLM to automatically formulate assessments based on external knowledge to evaluate the knowledge boundaries of foundation LLMs.To address knowledge inconsistencies in the alignment data, KCA implements several specific strategies to deal with these data instances.We demonstrate the superior efficacy of KCA in reducing hallucinations across six benchmarks, utilizing foundation LLMs of varying backbones and scales.This confirms the effectiveness of mitigating hallucinations by reducing knowledge inconsistency.Our code, model weights, and data are openly accessible at https://github.com/fanqiwan/KCA.* Part of the work was done during his internship at Tencent AI Lab.
Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, Shuming Shi 0001
EMNLP5
2024 Rethinking Targeted Adversarial Attacks for Neural Machine Translation
abstract
Targeted adversarial attacks are widely used to evaluate the robustness of neural machine translation systems. Unfortunately, this paper first identifies a critical issue in the existing settings of NMT targeted adversarial attacks, where their attacking results are largely overestimated. To this end, this paper presents a new setting for NMT targeted adversarial attacks that could lead to reliable attacking results. Under the new setting, it then proposes a Targeted Word Gradient adversarial Attack (TWGA) method to craft adversarial examples. Experimental results demonstrate that our proposed setting could provide faithful attacking results for targeted adversarial attacks on NMT systems, and the proposed TWGA method can effectively attack such victim NMT systems. In-depth analyses on a large-scale dataset further illustrate some valuable findings.1Our code and data are available at https://github.com/wujunjie1998/TWGA.
Junjie Wu 0007, Lemao Liu, Wei Bi, Dit-Yan Yeung
ICASSP3
2024 Retrieval is Accurate Generation
abstract
Standard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. One of the most significant challenges for this paradigm shift is determining the training oracles, because a string of text can be segmented in various ways and each segment can be retrieved from numerous possible documents. To address this, we propose to initialize the training oracles using linguistic heuristics and, more importantly, bootstrap the oracles through iterative self-reinforcement. Extensive experiments show that our model not only outperforms standard language models on a variety of knowledge-intensive tasks but also demonstrates improved generation quality in open-ended text generation. For instance, compared to the standard language model counterpart, our model raises the accuracy from 23.47% to 36.27% on OpenbookQA, and improves the MAUVE score from 42.61% to 81.58% in open-ended text generation. Remarkably, our model also achieves the best performance and the lowest latency among several retrieval-augmented baselines. In conclusion, we assert that retrieval is more accurate generation and hope that our work will encourage further research on this new paradigm shift.
Bowen Cao, Deng Cai 0002, Leyang Cui, Xuxin Cheng, Wei Bi, Yuexian Zou, Shuming Shi 0001
ICLR5
2024 Knowledge Fusion of Large Language Models
abstract
While training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures—Llama-2, MPT, and OpenLLaMA—across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at \url{https://github.com/fanqiwan/FuseLLM}.
Fanqi Wan, Xinting Huang, Deng Cai 0002, Xiaojun Quan, Wei Bi, Shuming Shi 0001
ICLR5
2024 Gated Slot Attention for Efficient Linear-Time Sequence Modeling
abstract
Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch. This paper introduces Gated Slot Attention (GSA), which enhances Attention with Bounded-memory-Control (ABC) by incorporating a gating mechanism inspired by Gated Linear Attention (GLA). Essentially, GSA comprises a two-layer GLA linked via $\operatorname{softmax}$, utilizing context-aware memory reading and adaptive forgetting to improve memory capacity while maintaining compact recurrent state size. This design greatly enhances both training and inference efficiency through GLA's hardware-efficient training algorithm and reduced state size. Additionally, retaining the $\operatorname{softmax}$ operation is particularly beneficial in ``finetuning pretrained Transformers to RNNs'' (T2R) settings, reducing the need for extensive training from scratch. Extensive experiments confirm GSA's superior performance in scenarios requiring in-context recall and in T2R settings.
Yu Zhang 0092, Rui-Jie Zhu 0003, Yue Zhang 0004, Leyang Cui, Yiqiao Wang 0005, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou 0017, Guohong Fu
NeurIPS10
2024 Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models
abstract
Recently, diffusion models have garnered significant interest in the field of text processing due to their many potential advantages compared to conventional autoregressive models. In this work, we propose Diffusion-of-Thought (DoT), a novel approach that integrates diffusion models with Chain-of-Thought, a well-established technique for improving the reasoning ability of autoregressive language models. In contrast to autoregressive language models that make decisions in a left-to-right, token-by-token manner, DoT allows reasoning steps to diffuse over time through a diffusion language model and offers greater flexibility in trading-off computation for reasoning performance. Our experimental results demonstrate the effectiveness of DoT in multi-digit multiplication, boolean logic, and grade school math problems. In addition to that, DoT showcases promising self-correction abilities and benefits from existing reasoning-enhancing techniques like self-consistency decoding. Our findings contribute to the understanding and development of reasoning with diffusion language models.
Jiacheng Ye, Shansan Gong, Jiahui Gao 0002, Xin Jiang 0002, Zhenguo Li, Wei Bi, Lingpeng Kong
NeurIPS10
2024 Spatial entropy as an inductive bias for vision transformers
abstract
Abstract Recent work on Vision Transformers (VTs) showed that introducing a local inductive bias in the VT architecture helps reducing the number of samples necessary for training. However, the architecture modifications lead to a loss of generality of the Transformer backbone, partially contradicting the push towards the development of uniform architectures, shared, e.g., by both the Computer Vision and the Natural Language Processing areas. In this work, we propose a different and complementary direction, in which a local bias is introduced using an auxiliary self-supervised task, performed jointly with standard supervised training. Specifically, we exploit the observation that the attention maps of VTs, when trained with self-supervision, can contain a semantic segmentation structure which does not spontaneously emerge when training is supervised. Thus, we explicitly encourage the emergence of this spatial clustering as a form of training regularization. In more detail, we exploit the assumption that, in a given image, objects usually correspond to few connected regions, and we propose a spatial formulation of the information entropy to quantify this object-based inductive bias. By minimizing the proposed spatial entropy, we include an additional self-supervised signal during training. Using extensive experiments, we show that the proposed regularization leads to equivalent or better results than other VT proposals which include a local bias by changing the basic Transformer architecture, and it can drastically boost the VT final accuracy when using small-medium training sets. The code is available at https://github.com/helia95/SAR .
Elia Peruzzo, Enver Sangineto, Marco De Nadai, Wei Bi, Bruno Lepri, Nicu Sebe
Mach. Learn.5
2023 Pre-training Multi-party Dialogue Models with Latent Discourse Inference
abstract
Multi-party dialogues are more difficult for models to understand than one-to-one twoparty dialogues, since they involve multiple interlocutors, resulting in interweaving reply-to relations and information flows.To step over these obstacles, an effective way is to pre-train a model that understands the discourse structure of multi-party dialogues, namely, to whom each utterance is replying.However, due to the lack of explicitly annotated discourse labels in multi-party dialogue corpora, previous works fail to scale up the pre-training process by putting aside the unlabeled multi-party conversational data for nothing.To fully utilize the unlabeled data, we propose to treat the discourse structures as latent variables, then jointly infer them and pre-train the discourse-aware model by unsupervised latent variable inference methods.Experiments on multiple downstream tasks show that our pre-trained model outperforms strong baselines by large margins and achieves state-of-the-art (SOTA) results, justifying the effectiveness of our method.The official implementation of this paper is available at https://github.com/EricLee8/MPD_EMVI.
Yiyang Li 0002, Xinting Huang, Wei Bi, Hai Zhao 0001
ACL (1)3
2023 Explicit Syntactic Guidance for Neural Text Generation
abstract
Most existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity.
Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004
ACL (1)5
2023 Multi-Grained Knowledge Retrieval for End-to-End Task-Oriented Dialog
abstract
Retrieving proper domain knowledge from an external database lies at the heart of end-toend task-oriented dialog systems to generate informative responses.Most existing systems blend knowledge retrieval with response generation and optimize them with direct supervision from reference responses, leading to suboptimal retrieval performance when the knowledge base becomes large-scale.To address this, we propose to decouple knowledge retrieval from response generation and introduce a multigrained knowledge retriever (MAKER) that includes an entity selector to search for relevant entities and an attribute selector to filter out irrelevant attributes.To train the retriever, we propose a novel distillation objective that derives supervision signals from the response generator.Experiments conducted on three standard benchmarks with both small and largescale knowledge bases demonstrate that our retriever performs knowledge retrieval more effectively than existing methods.Our code has been made publicly available.
Fanqi Wan, Weizhou Shen, Xiaojun Quan, Wei Bi
ACL (1)5
2023 Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers
abstract
Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the non-occluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP.
Bin Ren 0005, Yue Song 0002, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang 0108
CVPR4
2023 RobustGEC: Robust Grammatical Error Correction Against Subtle Context Perturbation
abstract
Grammatical Error Correction (GEC) systems play a vital role in assisting people with their daily writing tasks.However, users may sometimes come across a GEC system that initially performs well but fails to correct errors when the inputs are slightly modified.To ensure an ideal user experience, a reliable GEC system should have the ability to provide consistent and accurate suggestions when encountering irrelevant context perturbations, which we refer to as context robustness.In this paper, we introduce RobustGEC, a benchmark designed to evaluate the context robustness of GEC systems.RobustGEC comprises 5,000 GEC cases, each with one original error-correct sentence pair and five variants carefully devised by human annotators.Utilizing RobustGEC, we reveal that state-of-the-art GEC systems still lack sufficient robustness against context perturbations.In addition, we propose a simple yet effective method for remitting this issue.
Yue Zhang 0004, Leyang Cui, Enbo Zhao, Wei Bi, Shuming Shi 0001
EMNLP4
2023 Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue System
abstract
Developing an efficient retriever to retrieve knowledge from a large-scale knowledge base (KB) is critical for task-oriented dialogue systems to effectively handle localized and specialized tasks.However, widely used generative models such as T5 and ChatGPT often struggle to differentiate subtle differences among the retrieved KB records when generating responses, resulting in suboptimal quality of generated responses.In this paper, we propose the application of maximal marginal likelihood to train a perceptive retriever by utilizing signals from response generation for supervision.In addition, our approach goes beyond considering solely retrieved entities and incorporates various meta knowledge to guide the generator, thus improving the utilization of knowledge.We evaluate our approach on three task-oriented dialogue datasets using T5 and ChatGPT as the backbone models.The results demonstrate that when combined with meta knowledge, the response generator can effectively leverage high-quality knowledge records from the retriever and enhance the quality of generated responses.The code of this work is available at https://github.com/shenwzh3/MK-TOD.
Weizhou Shen, Yingqi Gao, Canbin Huang, Fanqi Wan, Xiaojun Quan, Wei Bi
EMNLP6
2023 Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration
abstract
Instruction-tuning can be substantially optimized through enhanced diversity, resulting in models capable of handling a broader spectrum of tasks.However, existing data employed for such tuning often exhibit an inadequate coverage of individual domains, limiting the scope for nuanced comprehension and interactions within these areas.To address this deficiency, we propose EXPLORE-INSTRUCT, a novel approach to enhance the data coverage to be used in domain-specific instruction-tuning through active exploration via Large Language Models (LLMs).Built upon representative domain use cases, EXPLORE-INSTRUCT explores a multitude of variations or possibilities by implementing a search algorithm to obtain diversified and domain-focused instruction-tuning data.Our data-centric analysis validates the effectiveness of this proposed approach in improving domain-specific instruction coverage.Moreover, our model's performance demonstrates considerable advancements over multiple baselines, including those utilizing domainspecific data enhancement.Our findings offer a promising opportunity to improve instruction coverage, especially in domain-specific contexts, thereby advancing the development of adaptable language models.Our code, model weights, and data are public at https:// github.com/fanqiwan/Explore-Instruct.
Fanqi Wan, Xinting Huang, Tao Yang 0033, Xiaojun Quan, Wei Bi, Shuming Shi 0001
EMNLP5
2023 Predicting Events in MOBA Games: Prediction, Attribution, and Evaluation
abstract
The multiplayer online battle arena (MOBA) games have become increasingly popular in recent years. Consequently, many efforts have been devoted to providing pregame or in-game predictions for them. These predictions can be used in many MOBA esports-related applications, such as artificial intelligence commentator systems, in-game data analysis, and game-assistant bots. However, these works are limited in the following two aspects: the lack of sufficient in-game features and the absence of interpretability in the prediction results. These two limitations greatly restrict the practical performance and industrial application of the current works. In this work, we collect a large-scale dataset containing rich in-game features for the popular MOBA gameHonor of Kings. We then propose to predict four types of prediction tasks in an interpretable way by attributing the predictions to the input features using two gradient-based attribution methods:Integrated GradientsandSmoothGrad. To evaluate the explanatory power of different models and attribution methods, a fidelity-based evaluation metric is further proposed. Finally, we evaluate the accuracy and fidelity of several competitive methods to assess how well machines predict events in MOBA games.
Zelong Yang 0002, Yan Wang 0060, Piji Li, Shaobin Lin, Shuming Shi 0001, Shao-Lun Huang, Wei Bi
IEEE Trans. Games7
2022 Lexical Knowledge Internalization for Neural Dialog Generation
abstract
We propose knowledge internalization (KI), which aims to complement the lexical knowledge into neural dialog models.Instead of further conditioning the knowledge-grounded dialog (KGD) models on externally retrieved knowledge, we seek to integrate knowledge about each input token internally into the model's parameters.To tackle the challenge due to the large scale of lexical knowledge, we adopt the contrastive learning approach and create an effective token-level lexical knowledge retriever that requires only weak supervision mined from Wikipedia.We demonstrate the effectiveness and general applicability of our approach on various datasets and diversified model structures.
Zhiyong Wu 0003, Wei Bi, Xiang Li 0067, Lingpeng Kong, Ben Kao
ACL (1)2
2022 A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation
abstract
Towards building intelligent dialogue agents, there has been a growing interest in introducing explicit personas in generation models.However, with limited persona-based dialogue data at hand, it may be difficult to train a dialogue generation model well.We point out that the data challenges of this generation task lie in two aspects: first, it is expensive to scale up current persona-based dialogue datasets; second, each data sample in this task is more complex to learn with than conventional dialogue data.To alleviate the above data issues, we propose a data manipulation method, which is model-agnostic to be packed with any personabased dialogue generation model to improve its performance.The original training samples will first be distilled and thus expected to be fitted more easily.Next, we show various effective ways that can diversify such easier distilled data.A given base model will then be trained via the constructed data curricula, i.e. first on augmented distilled samples and then on original ones.Experiments illustrate the superiority of our method with two strong base dialogue models (Transformer encoderdecoder and GPT2).
Yu Cao 0014, Wei Bi, Shuming Shi 0001, Dacheng Tao
ACL (1)2
2022 On Synthetic Data for Back Translation
abstract
Jiahao Xu, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi, Lihui Chen, Lemao Liu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jiahao Xu 0001, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi 0001, Lihui Chen 0001, Lemao Liu
NAACL-HLT3
2022 Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA Esport
abstract
With the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability.
Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu
IEEE Trans. Games7
2021 Learning from My Friends: Few-Shot Personalized Conversation Systems via Social Networks
abstract
Personalized conversation models (PCMs) generate responses according to speaker preferences. Existing personalized conversation tasks typically require models to extract speaker preferences from user descriptions or their conversation histories, which are scarce for newcomers and inactive users. In this paper, we propose a few-shot personalized conversation task with an auxiliary social network. The task requires models to generate personalized responses for a speaker given a few conversations from the speaker and a social network. Existing methods are mainly designed to incorporate descriptions or conversation histories. Those methods can hardly model speakers with so few conversations or connections between speakers. To better cater for newcomers with few resources, we propose a personalized conversation model (PCM) that learns to adapt to new speakers as well as enabling new speakers to learn from resource-rich speakers. Particularly, based on a meta-learning based PCM, we propose a task aggregator (TA) to collect other speakers' information from the social network. The TA provides prior knowledge of the new speaker in its meta-learning. Experimental results show our methods outperform all baselines in appropriateness, diversity, and consistency with speakers.
Zhiliang Tian, Wei Bi, Yiping Song, Nevin Lianwen Zhang
AAAI2
2021 Data Augmentation for Text Generation Without Any Augmented Data
abstract
Wei Bi, Huayang Li, Jiacheng Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Bi, Jiacheng Huang 0005
ACL/IJCNLP (1)1
2021 Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
abstract
Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, Ben Kao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhiyong Wu 0003, Lingpeng Kong, Wei Bi, Xiang Li 0067, Ben Kao
ACL/IJCNLP (1)3
2021 Uncertainty-Aware Self-Training for Semi-Supervised Event Temporal Relation Extraction
abstract
Extracting event temporal relations is an important task for natural language understanding. Many works have been proposed for supervised event temporal relation extraction, which typically requires a large amount of human-annotated data for model training. However, the data annotation for this task is very time-consuming and challenging. To this end, we study the problem of semi-supervised event temporal relation extraction. Self-training as a widely used semi-supervised learning method can be utilized for this problem. However, it suffers from the noisy pseudo-labeling problem. In this paper, we propose the use of uncertainty-aware self-training framework (UAST) to quantify the model uncertainty for coping with pseudo-labeling errors. Specifically, UAST utilizes (1) Uncertainty Estimation module to compute the model uncertainty for pseudo-labeling unlabeled data; (2) Sample Selection with Exploration module to select informative samples based on uncertainty estimates; and (3) Uncertainty-Aware Learning module to explicitly incorporate the model uncertainty into the self-training process. Experimental results indicate that our approach significantly outperforms previous state-of-the-art methods.
Xinyu Zuo, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Wei Bi
CIKM6
2021 Set Generation Networks for End-to-End Knowledge Base Population
abstract
The task of knowledge base population (KBP) aims to discover facts about entities from texts and expand a knowledge base with these facts.Previous studies shape end-to-end KBP as a machine translation task, which is required to convert unordered fact into a sequence according to a pre-specified order.However, the facts stated in a sentence are unordered in essence.In this paper, we formulate end-to-end KBP as a direct set generation problem, avoiding considering the order of multiple facts.To solve the set generation problem, we propose networks featured by transformers with nonautoregressive parallel decoding.Unlike previous approaches that use an autoregressive decoder to generate facts one by one, the proposed networks can directly output the final set of facts in one shot.Furthermore, to train the networks, we also design a set-based loss that forces unique predictions via bipartite matching.Compared with cross-entropy loss that highly penalizes small shifts in fact order, the proposed bipartite matching loss is invariant to any permutation of predictions.Benefiting from getting rid of the burden of predicting the order of multiple facts, our proposed networks achieve state-of-the-art (SoTA) performance on two benchmark datasets.
Dianbo Sui, Chenhao Wang 0004, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Wei Bi
EMNLP (1)6
2021 Efficient Training of Visual Transformers with Small Datasets
abstract
Visual Transformers (VTs) are emerging as an architectural paradigm alternative to Convolutional networks (CNNs). Differently from CNNs, VTs can capture global relations between image elements and they potentially have a larger representation capacity. However, the lack of the typical convolutional inductive bias makes these models more data hungry than common CNNs. In fact, some local properties of the visual domain which are embedded in the CNN architectural design, in VTs should be learned from samples. In this paper, we empirically analyse different VTs, comparing their robustness in a small training set regime, and we show that, despite having a comparable accuracy when trained on ImageNet, their performance on smaller datasets can be largely different. Moreover, we propose an auxiliary self-supervised task which can extract additional information from images with only a negligible computational overhead. This task encourages the VTs to learn spatial relations within an image and makes the VT training much more robust when training data is scarce. Our task is used jointly with the standard (supervised) training and it does not depend on specific architectural choices, thus it can be easily plugged in the existing VTs. Using an extensive evaluation with different VTs and datasets, we show that our method can improve (sometimes dramatically) the final accuracy of the VTs. Our code is available at: https://github.com/yhlleo/VTs-Drloc.
Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, Marco De Nadai
NeurIPS3
2021 Adaptive Fuzzy-Region-Based Control of Euler-Lagrange Systems With Kinematically Singular Configurations
abstract
Singularity issue has long been a concern of the task-space control design for Euler-Lagrange systems. In classical task-space controls, robots are often assumed to operate in the task space, where singularities do not exist. Such an assumption limits their potential applications in various workspaces. To address the potential singularity issue associated with Euler-Lagrange systems, this article proposes an adaptive fuzzy-region-based control for Euler-Lagrange systems with kinematically singular configurations. Singular regions are described by the potential energy function. The proposed controller includes a joint-space control, which is active when the system approaches singular regions, and a task-space control, which is used to track the desired trajectory. Therefore, the system can smoothly transit from singular regions to nonsingular regions or can achieve singularity avoidance during the tracking task. In order to achieve singularity avoidance while reducing control effort, the coefficients of the potential energy function are adjusted dynamically based on the designed fuzzy system. Rigorous analysis shows that singularity issues can be properly handled, and the asymptotic stability of the system is ensured. Experiments are conducted to demonstrate the effectiveness of the proposed controller.
Hongbo Gao 0001, Wei Bi, Zhijun Li 0001, Zhen Kan, Yu Kang 0001
IEEE Trans. Fuzzy Syst.2
2021 Trail-Traced Threshold Test (T4) With a Weighted Binomial Distribution for a Psychophysical Test
abstract
Clinical visual field testing is performed with commercial perimetric devices and employs psychophysical techniques to obtain thresholds of the differential light sensitivity (DLS) at multiple retinal locations. Current thresholding algorithms are relatively inefficient and tough to get satisfied test accuracy, stability concurrently. Thus, we propose a novel Bayesian perimetric threshold method called the Trail-Traced Threshold Test (T4), which can better address the dependence of the initial threshold estimation and achieve significant improvement in the test accuracy and variability while also decreasing the number of presentations compared with Zippy Estimation by Sequential Testing (ZEST) and FT. This study compares T4 with ZEST and FT regarding presentation number, mean absolute difference (MAD between the real Visual field result and the simulate result), and measurement variability. T4 uses the complete response sequence with the spatially weighted neighbor responses to achieve better accuracy and precision than ZEST, FT, SWeLZ, and with significantly fewer stimulus presentations. T4 is also more robust to inaccurate initial threshold estimation than other methods, which is an advantage in subjective methods, such as in clinical perimetry. This method also has the potential for using in other psychophysical tests.
Yuxin Gong, Haogang Zhu, Marco Miranda, David P. Crabb, Haolan Yang, Wei Bi, David F. Garway-Heath
IEEE J. Biomed. Health Informatics6
2020 Learning to Select Bi-Aspect Information for Document-Scale Text Content Manipulation
Yawei Sun, Bing Qin 0001, Heng Gong, Wei Bi, Xiaojiang Liu, Ting Liu 0001
AAAI6
2020 Relevance-Promoting Language Model for Short-Text Conversation
abstract
Despite the effectiveness of sequence-to-sequence framework on the task of Short-Text Conversation (STC), the issue of under-exploitation of training data (i.e., the supervision signals from query text is ignored) still remains unresolved. Also, the adopted maximization-based decoding strategies, inclined to generating the generic responses or responses with repetition, are unsuited to the STC task. In this paper, we propose to formulate the STC task as a language modeling problem and tailor-make a training strategy to adapt a language model for response generation. To enhance generation performance, we design a relevance-promoting transformer language model, which performs additional supervised source attention after the self-attention to increase the importance of informative query tokens in calculating the token-level representation. The model further refines the query representation with relevance clues inferred from its multiple references during training. In testing, we adopt a randomization-over-maximization strategy to reduce the generation of generic responses. Experimental results on a large Chinese STC dataset demonstrate the superiority of the proposed model on relevance metrics and diversity metrics.1
Xin Li 0056, Piji Li, Wei Bi, Xiaojiang Liu, Wai Lam
AAAI3
2020 Improving Knowledge-Aware Dialogue Generation via Knowledge Base Question Answering
abstract
Neural network models usually suffer from the challenge of incorporating commonsense knowledge into the open-domain dialogue systems. In this paper, we propose a novel knowledge-aware dialogue generation model (called TransDG), which transfers question representation and knowledge matching abilities from knowledge base question answering (KBQA) task to facilitate the utterance understanding and factual knowledge selection for dialogue generation. In addition, we propose a response guiding attention and a multi-step decoding strategy to steer our model to focus on relevant features for response generation. Experiments on two benchmark datasets demonstrate that our model has robust superiority over compared methods in generating informative and fluent dialogues. Our code is available at https://github.com/siat-nlp/TransDG.
Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007
AAAI3
2020 Learning to Customize Model Structures for Few-shot Dialogue Generation Tasks
abstract
Training the generative models with minimal corpus is one of the critical challenges for building open-domain dialogue systems.Existing methods tend to use the meta-learning framework which pre-trains the parameters on all non-target tasks then fine-tunes on the target task.However, fine-tuning distinguishes tasks from the parameter perspective but ignores the model-structure perspective, resulting in similar dialogue models for different tasks.In this paper, we propose an algorithm that can customize a unique dialogue model for each task in the few-shot setting.In our approach, each dialogue model consists of a shared module, a gating module, and a private module.The first two modules are shared among all the tasks, while the third one will differentiate into different network structures to better capture the characteristics of the corresponding task.The extensive experiments on two datasets show that our method outperforms all the baselines in terms of task consistency, response quality, and diversity.
Yiping Song, Zequn Liu, Wei Bi, Rui Yan 0001, Ming Zhang 0004
ACL3
2020 Response-Anticipated Memory for On-Demand Knowledge Integration in Response Generation
abstract
Neural conversation models are known to generate appropriate but non-informative responses in general.A scenario where informativeness can be significantly enhanced is Conversing by Reading (CbR), where conversations take place with respect to a given external document.In previous work, the external document is utilized by (1) creating a contextaware document memory that integrates information from the document and the conversational context, and then (2) generating responses referring to the memory.In this paper, we propose to create the document memory with some anticipated responses in mind.This is achieved using a teacher-student framework.The teacher is given the external document, the context, and the ground-truth response, and learns how to build a response-aware document memory from three sources of information.The student learns to construct a response-anticipated document memory from the first two sources, and the teacher's insight on memory creation.Empirical results show that our model outperforms the previous stateof-the-art for the CbR task.
Zhiliang Tian, Wei Bi, Lanqing Xue, Yiping Song, Xiaojiang Liu, Nevin Lianwen Zhang
ACL2
2020 A Batch Normalized Inference Network Keeps the KL Vanishing Away
abstract
Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks.However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse".Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint.We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive.Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters.Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently.We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE).Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE.
Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li 0001, Dapeng Oliver Wu
ACL2
2020 TableGPT: Few-shot Table-to-Text Generation with Table Structure Reconstruction and Content Matching
abstract
Although neural table-to-text models have achieved remarkable progress with the help of largescale datasets, they suffer insufficient learning problem with limited training data.Recently, pretrained language models show potential in few-shot learning with linguistic knowledge learnt from pretraining on large-scale corpus.However, benefiting table-to-text generation in few-shot setting with the powerful pretrained language model faces three challenges, including (1) the gap between the task's structured input and the natural language input for pretraining language model.(2) The lack of modeling for table structure and ( 3) improving text fidelity with less incorrect expressions that are contradicting to the table.To address aforementioned problems, we propose TableGPT for table-to-text generation.At first, we utilize table transformation module with template to rewrite structured table in natural language as input for GPT-2.In addition, we exploit multi-task learning with two auxiliary tasks that preserve table's structural information by reconstructing the structure from GPT-2's representation and improving the text's fidelity with content matching task aligning the table and information in the generated text.By experimenting on Humans, Songs and Books, three few-shot table-to-text datasets in different domains, our model outperforms existing systems on most few-shot settings.
Heng Gong, Yawei Sun, Bing Qin 0001, Wei Bi, Xiaojiang Liu, Ting Liu 0001
COLING5
2020 Dual Dynamic Memory Network for End-to-End Multi-turn Task-oriented Dialog Systems
abstract
Existing end-to-end task-oriented dialog systems struggle to dynamically model long dialog context for interactions and effectively incorporate knowledge base (KB) information into dialog generation. To conquer these limitations, we propose a Dual Dynamic Memory Network (DDMN) for multi-turn dialog generation, which maintains two core components: dialog memory manager and KB memory manager. The dialog memory manager dynamically expands the dialog memory turn by turn and keeps track of dialog history with an updating mechanism, which encourages the model to filter irrelevant dialog history and memorize important newly coming information. The KB memory manager shares the structural KB triples throughout the whole conversation, and dynamically extracts KB information with a memory pointer at each turn. Experimental results on three benchmark datasets demonstrate that DDMN significantly outperforms the strong baselines in terms of both automatic evaluation and human evaluation. Our code is available at https://github.com/siat-nlp/DDMN.
Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007
COLING3
2020 Event Extraction as Machine Reading Comprehension
abstract
Event extraction (EE) is a crucial information extraction task that aims to extract event information in texts.Previous methods for EE typically model it as a classification task, which are data-hungry and suffer from the data scarcity problem.In this paper, we propose a new learning paradigm of EE, by explicitly casting it as a machine reading comprehension problem (MRC).Our approach includes an unsupervised question generation process, which can transfer event schema into a set of natural questions, followed by a BERTbased question-answering process to retrieve answers as EE results.This learning paradigm enables us to strengthen the reasoning process of EE, by introducing sophisticated models in MRC, and relieve the data scarcity problem, by introducing the large-scale datasets in MRC.The empirical results show that: i) our approach attains state-of-the-art performance by considerable margins over previous methods.ii) Our model is excelled in the data-scarce scenario, for example, obtaining 49.8% in F1 for event argument extraction with only 1% data, compared with 2.2% of the previous method.iii) Our model also fits with zero-shot scenarios, achieving 37.0% and 16% in F1 on two datasets without using any EE training data.
Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001, Wei Bi, Xiaojiang Liu
EMNLP (1)4
2019 Generating Multiple Diverse Responses for Short-Text Conversation
abstract
Neural generative models have become popular and achieved promising performance on short-text conversation tasks. They are generally trained to build a 1-to-1 mapping from the input post to its output response. However, a given post is often associated with multiple replies simultaneously in real applications. Previous research on this task mainly focuses on improving the relevance and informativeness of the top one generated response for each post. Very few works study generating multiple accurate and diverse responses for the same post. In this paper, we propose a novel response generation model, which considers a set of responses jointly and generates multiple diverse responses simultaneously. A reinforcement learning algorithm is designed to solve our model. Experiments on two short-text conversation tasks validate that the multiple responses generated by our model obtain higher quality and larger diversity compared with various state-ofthe-art generative models.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Shuming Shi 0001
AAAI2
2019 Better Fine-Tuning via Instance Weighting for Text Classification
abstract
Transfer learning for deep neural networks has achieved great success in many text classification applications. A simple yet effective transfer learning method is to fine-tune the pretrained model parameters. Previous fine-tuning works mainly focus on the pre-training stage and investigate how to pretrain a set of parameters that can help the target task most. In this paper, we propose an Instance Weighting based Finetuning (IW-Fit) method, which revises the fine-tuning stage to improve the final performance on the target domain. IW-Fit adjusts instance weights at each fine-tuning epoch dynamically to accomplish two goals: 1) identify and learn the specific knowledge of the target domain effectively; 2) well preserve the shared knowledge between the source and the target domains. The designed instance weighting metrics used in IW-Fit are model-agnostic, which are easy to implement for general DNN-based classifiers. Experimental results show that IW-Fit can consistently improve the classification accuracy on the target domain.
Wei Bi, Yan Wang 0060, Xiaojiang Liu
AAAI2
2019 Fine-Grained Sentence Functions for Short-Text Conversation
abstract
Sentence function is an important linguistic feature referring to a user's purpose in uttering a specific sentence.The use of sentence function has shown promising results to improve the performance of conversation models.However, there is no large conversation dataset annotated with sentence functions.In this work, we collect a new Short-Text Conversation dataset with manually annotated SEntence FUNctions (STC-Sefun).Classification models are trained on this dataset to (i) recognize the sentence function of new data in a large corpus of short-text conversations; (ii) estimate a proper sentence function of the response given a test query.We later train conversation models conditioned on the sentence functions, including information retrieval-based and neural generative models.Experimental results demonstrate that the use of sentence functions can help improve the quality of the returned responses.
Wei Bi, Xiaojiang Liu, Shuming Shi 0001
ACL (1)1
2019 Are Training Samples Correlated? Learning to Generate Dialogue Responses with Multiple References
abstract
Due to its potential applications, open-domain dialogue generation has become popular and achieved remarkable progress in recent years, but sometimes suffers from generic responses.Previous models are generally trained based on 1-to-1 mapping from an input query to its response, which actually ignores the nature of 1-to-n mapping in dialogue that there may exist multiple valid responses corresponding to the same query.In this paper, we propose to utilize the multiple references by considering the correlation of different valid responses and modeling the 1-to-n mapping with a novel two-step generation architecture.The first generation phase extracts the common features of different responses which, combined with distinctive features obtained in the second phase, can generate multiple diverse and appropriate responses.Experimental results show that our proposed model can effectively improve the quality of response and outperform existing neural dialogue models on both automatic and human evaluations.
Lisong Qiu, Juntao Li 0005, Wei Bi, Dongyan Zhao 0001, Rui Yan 0001
ACL (1)3
2019 Learning to Abstract for Memory-augmented Conversational Response Generation
abstract
Neural generative models for open-domain chit-chat conversations have become an active area of research in recent years.A critical issue with most existing generative models is that the generated responses lack informativeness and diversity.A few researchers attempt to leverage the results of retrieval models to strengthen the generative models, but these models are limited by the quality of the retrieval results.In this work, we propose a memory-augmented generative model, which learns to abstract from the training corpus and saves the useful information to the memory to assist the response generation.Our model clusters query-response samples, extracts characteristics of each cluster, and learns to utilize these characteristics for response generation.Experimental results show that our model outperforms other competitive baselines.
Zhiliang Tian, Wei Bi, Nevin Lianwen Zhang
ACL (1)2
2019 Unsupervised Rewriter for Multi-Sentence Compression
abstract
Multi-sentence compression (MSC) aims to generate a grammatical but reduced compression from multiple input sentences while retaining their key information.Previous dominating approach for MSC is the extractionbased word graph approach.A few variants further leveraged lexical substitution to yield more abstractive compression.However, two limitations exist.First, the word graph approach that simply concatenates fragments from multiple sentences may yield nonfluent or ungrammatical compression.Second, lexical substitution is often inappropriate without the consideration of context information.To tackle the above-mentioned issues, we present a neural rewriter for multisentence compression that does not need any parallel corpus.Empirical studies have shown that our approach achieves comparable results upon automatic evaluation and improves the grammaticality of compression based on human evaluation.A parallel corpus with more than 140,000 (sentence group, compression) pairs is also constructed as a by-product for future research.
Xiaoyu Shen 0001, Wei Bi, Akiko Aizawa
ACL (1)3
2019 Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework
abstract
Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Deng Cai 0002, Yan Wang 0060, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi 0001
EMNLP/IJCNLP (1)3
2019 A Discrete CVAE for Response Generation on Short-Text Conversation
abstract
Jun Gao, Wei Bi, Xiaojiang Liu, Junhui Li, Guodong Zhou, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Guodong Zhou 0001, Shuming Shi 0001
EMNLP/IJCNLP (1)2
2019 Spontaneous facial expression database for academic emotion inference in online learning
abstract
Academic emotions can produce a great impact on the learning effect. Normally, emotions are expressed externally in the students' facial expressions, speech and behaviour. In this paper, the focus is on automatic academic emotion inference based on facial expressions in online learning. Considering the lack of training samples for the inference algorithm, a spontaneous facial expression database is established. It includes the facial expressions of five common academic emotions and consists of two subsets: a video clip database and an image database. A total of 1,274 video clips and 30,184 images from 82 students are included in the database. The samples are labelled by both the participants and external coders. An extensive analysis is carried out on the image database using a convolutional neural network (CNN)‐based algorithm to infer self‐annotation. Some data augmentation algorithms are applied to improve the algorithm performance. Additionally, an adaptive data augmentation algorithm based on spatial transformer network is introduced, which can remove some confounding factors in the original images. The algorithm can obviously improve the inference performance, which has been proven by comparing some evaluation indicators before and after adoption. Such a database will certainly accelerate the application of affective computing in the educational field.
Cunling Bian, Fei Yang 0003, Wei Bi, Weigang Lu 0002
IET Comput. Vis.4
2019 Threshold changeable secret image sharing scheme based on interpolation polynomial
Yan-Xiao Liu 0001, Ching-Nung Yang, Chi-Ming Wu, Qindong Sun, Wei Bi
Multim. Tools Appl.5
2018 Towards Less Generic Responses in Neural Conversation Models: A Statistical Re-weighting Method
abstract
Sequence-to-sequence neural generation models have achieved promising performance on short text conversation tasks.However, they tend to generate generic/dull responses, leading to unsatisfying dialogue experience.We observe that in conversation tasks, each query could have multiple responses, which forms a 1-to-n or m-to-n relationship in the view of the total corpus.The objective function used in standard sequence-to-sequence models will be dominated by loss terms with generic patterns.Inspired by this observation, we introduce a statistical re-weighting method that assigns different weights for the multiple responses of the same query, and trains the standard neural generation model with the weights.Experimental results on a large Chinese dialogue corpus show that our method improves the acceptance rate of generated responses compared with several baseline models and significantly reduces the number of generated generic responses.
Wei Bi, Xiaojiang Liu, Jian Yao 0002, Shuming Shi 0001
EMNLP2
2018 MoSa: A Modeling and Sentiment Analysis System for Mobile Application Big Data
Yaocheng Zhang, Wei Ren 0002, Tianqing Zhu, Wei Bi
ICA3PP (2)4
2018 Image smoothing via a scale-aware filter and L 0 norm
abstract
It is difficult to preserve diminishing weak structures and edges, and remove complex details simultaneously in the context of image smoothing. While most of existing methods only take either local or global features into consideration, the authors propose two methods taking advantage of both to achieve smoothing, both of which consist of two steps and share the same first step. In the first step, the authors use a scale‐aware approach to generate a guidance image by blurring the small‐scale components in the input image. Such approach, based on the rolling guidance framework with domain transform filter and bilateral filter, can prevent diminishing the corners of the main structures. Subsequently, the authors use the two proposed methods, with the guidance image as input, to remove blurry details. The first method introduces two data fidelity terms into L 0 gradient minimisation and removes high‐contrast details, which is a structure‐preserving method. The other method, an edge‐preserving method, uses an adaptive L 0 gradient minimisation technique, facilitating the preservation of the weak structures and edges. The smoothing factors in such technique are decide by the corresponding gradient of each pixel of the guidance image. The authors apply both methods to various image processing fields.
Weiguo Huang, Wei Bi, Guanqi Gao, Yong Ping Zhang, Zhongkui Zhu
IET Image Process.2
2015 Bayes-Optimal Hierarchical Multilabel Classification
abstract
Hierarchical multilabel classification allows a sample to belong to multiple class labels residing on a hierarchy, which can be a tree or directed acyclic graph (DAG). However, popular hierarchical loss functions, such as the H-loss, can only be defined on tree hierarchies (but not on DAGs), and may also under- or over-penalize misclassifications near the bottom of the hierarchy. Besides, it has been relatively unexplored on how to make use of the loss functions in hierarchical multilabel classification. To overcome these deficiencies, we first propose hierarchical extensions of the Hamming loss and ranking loss which take the mistake at every node of the label hierarchy into consideration. Then, we first train a general learning model, which is independent of the loss function. Next, using Bayesian decision theory, we develop Bayes-optimal predictions that minimize the corresponding risks with the trained model. Computationally, instead of requiring an exhaustive summation and search for the optimal multilabel, the resultant optimization problem can be efficiently solved by a greedy algorithm. Experimental results on a number of real-world data sets show that the proposed Bayes-optimal classifier outperforms state-of-the-art methods.
Wei Bi, James T. Kwok
IEEE Trans. Knowl. Data Eng.1
2015 Large-Scale Nyström Kernel Matrix Approximation Using Randomized SVD
abstract
The Nyström method is an efficient technique for the eigenvalue decomposition of large kernel matrices. However, to ensure an accurate approximation, a sufficient number of columns have to be sampled. On very large data sets, the singular value decomposition (SVD) step on the resultant data submatrix can quickly dominate the computations and become prohibitive. In this paper, we propose an accurate and scalable Nyström scheme that first samples a large column subset from the input matrix, but then only performs an approximate SVD on the inner submatrix using the recent randomized low-rank matrix approximation algorithms. Theoretical analysis shows that the proposed algorithm is as accurate as the standard Nyström method that directly performs a large SVD on the inner submatrix. On the other hand, its time complexity is only as low as performing a small SVD. Encouraging results are obtained on a number of large-scale data sets for low-rank approximation. Moreover, as the most computational expensive steps can be easily distributed and there is minimal data transfer among the processors, significant speedup can be further obtained with the use of multiprocessor and multi-GPU systems.
Mu Li 0001, Wei Bi, James T. Kwok, Bao-Liang Lu
IEEE Trans. Neural Networks Learn. Syst.2
2014 Multilabel Classification with Label Correlations and Missing Labels
abstract
Many real-world applications involve multilabel classification, in which the labels can have strong inter-dependencies and some of them may even be missing.Existing multilabel algorithms are unable to handle both issues simultaneously.In this paper, we propose a probabilistic model that can automatically learn and exploit multilabel correlations.By integrating out the missing information, it also provides a disciplinedapproach to the handling of missing labels. The inference procedure is simple, and the optimization subproblems are convex. Experiments on a number of real-world data sets with both complete and missing labelsdemonstrate that the proposed algorithm can consistently outperform state-of-the-art multilabel classification algorithms.
Wei Bi, James T. Kwok
AAAI1
2014 Learning to Predict from Crowdsourced Data
Wei Bi, Liwei Wang 0009, James T. Kwok, Zhuowen Tu
UAI1
2014 Mandatory Leaf Node Prediction in Hierarchical Multilabel Classification
abstract
In hierarchical classification, the output labels reside on a tree- or directed acyclic graph (DAG)-structured hierarchy. On testing, the prediction paths of a given test example may be required to end at leaf nodes of the label hierarchy. This is called mandatory leaf node prediction (MLNP) and is particularly useful, when the leaf nodes have much stronger semantic meaning than the internal nodes. However, while there have been a lot of MLNP methods in hierarchical multiclass classification, performing MLNP in hierarchical multilabel classification is difficult. In this paper, we propose novel MLNP algorithms that consider the global label hierarchy structure. We show that the joint posterior probability over all the node labels can be efficiently maximized by dynamic programming for label trees, or greedy algorithm for label DAGs. In addition, both algorithms can be further extended for the minimization of the expected symmetric loss. Experiments are performed on real-world MLNP data sets with label trees and label DAGs. The proposed method consistently outperforms other hierarchical and flat multilabel classification methods.
Wei Bi, James T. Kwok
IEEE Trans. Neural Networks Learn. Syst.1
2013 Efficient Multi-label Classification with Many Labels
abstract
Multi-label classification deals with the problem where each instance can be associated with a set of class labels. However, in many real-world applications, the number of class labels can be in the hundreds or even thousands, and existing multi-label classification methods often become computationally inefficient. In recent years, a number of remedies have been proposed. However, they are either based on simple dimension reduction techniques or involve expensive optimization problems. In this paper, we address this problem by selecting a small subset of class labels that can approximately span the original label space. This is performed by randomized sampling where the sampling probability of each class label reflects its importance among all the labels. Theoretical analysis shows that this randomized sampling approach is highly efficient. Experiments on a number of real-world multi-label datasets with many labels demonstrate the appealing performance and efficiency of the proposed algorithm.
Wei Bi, James T. Kwok
ICML (3)1
2012 Hierarchical Multilabel Classification with Minimum Bayes Risk
abstract
Hierarchical multilabel classification (HMC) allows an instance to have multiple labels residing in a hierarchy. A popular loss function used in HMC is the H-loss, which penalizes only the first classification mistake along each prediction path. However, the H-loss metric can only be used on tree-structured label hierarchies, but not on DAG hierarchies. Moreover, it may lead to misleading predictions as not all misclassifications in the hierarchy are penalized. In this paper, we overcome these deficiencies by proposing a hierarchy-aware loss function that is more appropriate for HMC. Using Bayesian decision theory, we then develop a Bayes-optimal classifier with respect to this loss function. Instead of requiring an exhaustive summation and search for the optimal multilabel, the proposed classification problem can be efficiently solved using a greedy algorithm on both tree-and DAG-structured label hierarchies. Experimental results on a large number of real-world data sets show that the proposed algorithm outperforms existing HMC methods.
Wei Bi, James T. Kwok
ICDM1
2012 Mandatory Leaf Node Prediction in Hierarchical Multilabel Classification
abstract
In hierarchical classification, the prediction paths may be required to always end at leaf nodes. This is called mandatory leaf node prediction (MLNP) and is particularly useful when the leaf nodes have much stronger semantic meaning than the internal nodes. However, while there have been a lot of MLNP methods in hierarchical multiclass classification, performing MLNP in hierarchical multilabel classification is much more difficult. In this paper, we propose a novel MLNP algorithm that (i) considers the global hierarchy structure; and (ii) can be used on hierarchies of both trees and DAGs. We show that one can efficiently maximize the joint posterior probability of all the node labels by a simple greedy algorithm. Moreover, this can be further extended to the minimization of the expected symmetric loss. Experiments are performed on a number of real-world data sets with tree- and DAG-structured label hierarchies. The proposed method consistently outperforms other hierarchical and flat multilabel classification methods.
Wei Bi, James T. Kwok
NIPS1
2011 MultiLabel Classification on Tree- and DAG-Structured Hierarchies
Wei Bi, James T. Kwok
ICML1
2009 Extending Semi-supervised Learning Methods for Inductive Transfer Learning
abstract
Inductive transfer learning and semi-supervised learning are two different branches of machine learning. The former tries to reuse knowledge in labeled out-of-domain instances while the later attempts to exploit the usefulness of unlabeled in-domain instances. In this paper, we bridge the two branches by pointing out that many semi-supervised learning methods can be extended for inductive transfer learning, if the step of labeling an unlabeled instance is replaced by re-weighting a diff-distribution instance. Based on this recognition, we develop a new transfer learning method, namely COITL, by extending the co-training method in semi-supervised learning. Experimental results reveal that COITL can achieve significantly higher generalization and robustness, compared with two state-of-the-art methods in inductive transfer learning.
Zhen-Zhong Lan, Wei Liu 0015, Wei Bi
ICDM4