VLDB 2026 Research / reviewers in the wild / expert
Deng Cai 0002
dblp:c/DCai-2
· DBLP profile ↗
44ranked-venue papers
10as first author
31since 2021 · last 2025
0000-0003-3444-8383ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Empowering Self-Learning of LLMs: Inner Knowledge Explicitation as a CatalystabstractSelf-learning of Large Language Models (LLMs) facilitates their advancement towards super-intelligence by training with self-synthesized experiences. However, a critical challenge is the amplification of hallucinations in generated data during iterative self-learning, underscoring the need for reliable data selection. To address this, we investigate the mechanism of Inner Knowledge Explicitation, which involves explicitly extracting the inner knowledge from memory of LLMs, to concurrently improves reasoning, and enables reliable self-learning data selection. This paper introduces a Self Knowledge Explicitation Learning (SKE-Learn) framework, which equips the LLMs with meta-skills to explicitly extract, verify and utilize inner knowledge for reasoning. By leveraging these meta-skills, SKE-Learn establishes a self-learning approach that ensures reliable selection of self-synthetic data. This approach enhances performance through iterative self-learning while mitigating the problem of hallucinations. Empirical results from six benchmarks demonstrate that Inner Knowledge Explicitation improves reasoning by serving as a more effective prompting method. Additionally, SKE-Learn, based on the verifiability of explicit knowledge, shows consistent performance improvements over multiple self-training iterations, with an average performance increase from 52.79% to 56.54% across all benchmarks. Furthermore, Inner Knowledge Explicitation provides explanation and intervention space during LLM's generation process. Shijue Huang, Wanjun Zhong, Deng Cai 0002, Fanqi Wan, Mingxuan Wang, Ruifeng Xu 0001 |
AAAI | 3 |
| 2025 | ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationabstractWe introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs, requiring LMMs to generate the corresponding code for chart rendering.
ChartMimic includes $4,800$ human-curated (figure, instruction, code) triplets, which represent the authentic chart use cases found in scientific papers across various domains (e.g., Physics, Computer Science, Economics, etc). These charts span $18$ regular types and $4$ advanced types, diversifying into $201$ subcategories.
Furthermore, we propose multi-level evaluation metrics to provide an automatic and thorough assessment of the output code and the rendered charts.
Unlike existing code generation benchmarks, ChartMimic places emphasis on evaluating LMMs' capacity to harmonize a blend of cognitive capabilities, encompassing visual understanding, code generation, and cross-modal reasoning. The evaluation of $3$ proprietary models and $14$ open-weight models highlights the substantial challenges posed by ChartMimic. Even the advanced GPT-4o, InternVL2-Llama3-76B only achieved an average score across Direct Mimic and Customized Mimic tasks of $82.2$ and $61.6$, respectively, indicating significant room for improvement.
We anticipate that ChartMimic will inspire the development of LMMs, advancing the pursuit of artificial general intelligence. Cheng Yang 0002, Chufan Shi, Bo Shui, Junjie Wang 0011, Mohan Jing, Linran Xu, Siheng Li, Gongye Liu, Xiaomei Nie, Deng Cai 0002, Yujiu Yang 0001 |
ICLR | 13 |
| 2024 | WatME: Towards Lossless Watermarking Through Lexical RedundancyabstractLiang Chen, Yatao Bian, Yang Deng, Deng Cai, Shuaiyi Li, Peilin Zhao, Kam-Fai Wong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Liang Chen 0001, Yatao Bian, Yang Deng 0002, Deng Cai 0002, Shuaiyi Li, Peilin Zhao, Kam-Fai Wong |
ACL (1) | 4 |
| 2024 | A Frustratingly Simple Decoding Method for Neural Text GenerationabstractWe introduce a frustratingly simple, highly efficient, and surprisingly effective decoding method, termed Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: We construct an anti-language model (anti-LM) based on previously generated text, which is employed to penalize the future generation of repetitive content. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD incurs no additional model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite its simplicity, FSD is surprisingly effective and generalizes across different datasets, models, and languages. Extensive experiments show that FSD outperforms established strong baselines in terms of generation quality, decoding speed, and universality. Deng Cai 0002, Wei Bi, Wai Lam, Shuming Shi 0001 |
LREC/COLING | 2 |
| 2024 | Consecutive Batch Model Editing with HooK LayersabstractAs the typical retraining paradigm is unacceptably time-and resource-consuming, researchers are turning to model editing to find an effective way that supports both consecutive and batch scenarios to edit the model behavior directly.Despite all these practical expectations, existing model editing methods fail to realize all of them.Furthermore, the memory demands for such sequential model editing approaches tend to be prohibitive, frequently necessitating an external memory that grows incrementally over time.To cope with these challenges, we propose CoachHooK, a model editing method that simultaneously supports sequential and batch editing.CoachHooK is memory-friendly as it only needs a small amount of it to store several hook layers whose size remains unchanged over time.Experimental results demonstrate the superiority of our method over other batch-supportive model editing methods under both single-round and consecutive batch editing scenarios.Extensive analyses of CoachHooK have been conducted to verify the stability of our method over a number of consecutive steps. Shuaiyi Li, Yang Deng 0002, Deng Cai 0002, Hongyuan Lu, Liang Chen 0001, Wai Lam |
EMNLP | 3 |
| 2024 | A Thorough Examination of Decoding Methods in the Era of LLMsabstractDecoding methods play an indispensable role in converting language models from next-token predictors into practical task solvers.Prior research on decoding methods, primarily focusing on task-specific models, may not extend to the current era of general-purpose large language models (LLMs).Moreover, the recent influx of decoding strategies has further complicated this landscape.This paper provides a comprehensive and multifaceted analysis of various decoding methods within the context of LLMs, evaluating their performance, robustness to hyperparameter changes, and decoding speeds across a wide range of tasks, models, and deployment environments.Our findings reveal that decoding method performance is notably task-dependent and influenced by factors such as alignment, model size, and quantization.Intriguingly, sensitivity analysis exposes that certain methods achieve superior performance at the cost of extensive hyperparameter tuning, highlighting the trade-off between attaining optimal results and the practicality of implementation in varying contexts. Chufan Shi, Deng Cai 0002, Zhisong Zhang, Yujiu Yang 0001, Wai Lam |
EMNLP | 3 |
| 2024 | Retrieval is Accurate GenerationabstractStandard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. One of the most significant challenges for this paradigm shift is determining the training oracles, because a string of text can be segmented in various ways and each segment can be retrieved from numerous possible documents. To address this, we propose to initialize the training oracles using linguistic heuristics and, more importantly, bootstrap the oracles through iterative self-reinforcement. Extensive experiments show that our model not only outperforms standard language models on a variety of knowledge-intensive tasks but also demonstrates improved generation quality in open-ended text generation. For instance, compared to the standard language model counterpart, our model raises the accuracy from 23.47% to 36.27% on OpenbookQA, and improves the MAUVE score from 42.61% to 81.58% in open-ended text generation. Remarkably, our model also achieves the best performance and the lowest latency among several retrieval-augmented baselines. In conclusion, we assert that retrieval is more accurate generation and hope that our work will encourage further research on this new paradigm shift. Bowen Cao, Deng Cai 0002, Leyang Cui, Xuxin Cheng, Wei Bi, Yuexian Zou, Shuming Shi 0001 |
ICLR | 2 |
| 2024 | The Reasonableness Behind Unreasonable Translation Capability of Large Language ModelabstractMultilingual large language models trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the LLM's translation abilities. However, it has also been observed that LLM's translation capabilities persist even when incidental sentence-level bilingualism are excluded from the training corpus.
In this study, we comprehensively investigate the unreasonable effectiveness and the underlying mechanism for LLM's translation abilities, specifically addressing the question why large language models learn to translate without parallel data, using the BLOOM model series as a representative example. Through extensive experiments, our findings suggest the existence of unintentional bilingualism in the pre-training corpus, especially word alignment data significantly contributes to the large language model's acquisition of translation ability. Moreover, the translation signal derived from word alignment data is comparable to that from sentence-level bilingualism. Additionally, we study the effects of monolingual data and parameter-sharing in assisting large language model to learn to translate. Together, these findings present another piece of the broader puzzle of trying to understand how large language models acquire translation capability. Tingchen Fu, Lemao Liu, Deng Cai 0002, Guoping Huang, Shuming Shi 0001, Rui Yan 0001 |
ICLR | 3 |
| 2024 | Knowledge Fusion of Large Language ModelsabstractWhile training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures—Llama-2, MPT, and OpenLLaMA—across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at \url{https://github.com/fanqiwan/FuseLLM}. Fanqi Wan, Xinting Huang, Deng Cai 0002, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
ICLR | 3 |
| 2024 | GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu |
ACM Multimedia | 7 |
| 2024 | GLBench: A Comprehensive Benchmark for Graph with Large Language ModelsabstractThe emergence of large language models (LLMs) has revolutionized the way we interact with graphs, leading to a new paradigm called GraphLLM. Despite the rapid development of GraphLLM methods in recent years, the progress and understanding of this field remain unclear due to the lack of a benchmark with consistent experimental protocols. To bridge this gap, we introduce GLBench, the first comprehensive benchmark for evaluating GraphLLM methods in both supervised and zero-shot scenarios. GLBench provides a fair and thorough evaluation of different categories of GraphLLM methods, along with traditional baselines such as graph neural networks. Through extensive experiments on a collection of real-world datasets with consistent data processing and splitting strategies, we have uncovered several key findings. Firstly, GraphLLM methods outperform traditional baselines in supervised settings, with LLM-as-enhancers showing the most robust performance. However, using LLMs as predictors is less effective and often leads to uncontrollable output issues. We also notice that no clear scaling laws exist for current GraphLLM methods. In addition, both structures and semantics are crucial for effective zero-shot transfer, and our proposed simple baseline can even outperform several models tailored for zero-shot scenarios. The data and code of the benchmark can be found at https://github.com/NineAbyss/GLBench. Yuhan Li 0001, Peisong Wang 0002, Aochuan Chen, Haiyun Jiang, Deng Cai 0002, Wai Kin Chan, Jia Li 0009 |
NeurIPS | 6 |
| 2024 | On the Worst Prompt Performance of Large Language ModelsabstractThe performance of large language models (LLMs) is acutely sensitive to the phrasing of prompts, which raises significant concerns about their reliability in real-world scenarios. Existing studies often divide prompts into task-level instructions and case-level inputs and primarily focus on evaluating and improving robustness against variations in tasks-level instructions. However, this setup fails to fully address the diversity of real-world user queries and assumes the existence of task-specific datasets. To address these limitations, we introduce RobustAlpacaEval, a new benchmark that consists of semantically equivalent case-level queries and emphasizes the importance of using the worst prompt performance to gauge the lower bound of model performance. Extensive experiments on RobustAlpacaEval with ChatGPT and six open-source LLMs from the Llama, Mistral, and Gemma families uncover substantial variability in model performance; for instance, a difference of 45.48% between the worst and best performance for the Llama-2-70B-chat model, with its worst performance dipping as low as 9.38%. We further illustrate the difficulty in identifying the worst prompt from both model-agnostic and model-dependent perspectives, emphasizing the absence of a shortcut to characterize the worst prompt. We also attempt to enhance the worst prompt performance using existing prompt engineering and prompt consistency methods, but find that their impact is limited. These findings underscore the need to create more resilient LLMs that can maintain high performance across diverse prompts. Bowen Cao, Deng Cai 0002, Zhisong Zhang, Yuexian Zou, Wai Lam |
NeurIPS | 2 |
| 2024 | StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem SolvingabstractMost existing prompting methods suffer from the issues of generalizability and consistency, as they often rely on instance-specific solutions that may not be applicable to other instances and lack task-level consistency across the selected few-shot examples. To address these limitations, we propose a comprehensive framework, StrategyLLM, allowing LLMs to perform inductive reasoning, deriving general strategies from specific task instances, and deductive reasoning, applying these general strategies to particular task examples, for constructing generalizable and consistent few-shot prompts. It employs four LLM-based agents: strategy generator, executor, optimizer, and evaluator, working together to generate, evaluate, and select promising strategies for a given task. Experimental results demonstrate that StrategyLLM outperforms the competitive baseline CoT-SC that requires human-annotated solutions on 13 datasets across 4 challenging tasks without human involvement, including math reasoning (34.2\% $\rightarrow$ 38.8\%), commonsense reasoning (70.3\% $\rightarrow$ 72.5\%), algorithmic reasoning (73.7\% $\rightarrow$ 85.0\%), and symbolic reasoning (30.0\% $\rightarrow$ 79.2\%). Further analysis reveals that StrategyLLM is applicable to various LLMs and demonstrates advantages across numerous scenarios. Haiyun Jiang, Deng Cai 0002, Shuming Shi 0001, Wai Lam |
NeurIPS | 3 |
| 2024 | Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-ContrastabstractMixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activates a different subset of experts determined by a routing mechanism. However, the unchosen experts in MoE models do not contribute to the output, potentially leading to underutilization of the model's capacity.
In this work, we first conduct exploratory studies to demonstrate that increasing the number of activated experts does not necessarily improve and can even degrade the output quality. Then, we show that output distributions from an MoE model using different routing strategies substantially differ, indicating that different experts do not always act synergistically.
Motivated by these findings, we propose **S**elf-**C**ontrast **M**ixture-**o**f-**E**xperts (SCMoE), a training-free strategy that utilizes unchosen experts in a self-contrast manner during inference.
In SCMoE, the next-token probabilities are determined by contrasting the outputs from strong and weak activation using the same MoE model.
Our method is conceptually simple and computationally lightweight, as it incurs minimal latency compared to greedy decoding.
Experiments on several benchmarks (GSM8K, StrategyQA, MBPP and HumanEval) demonstrate that SCMoE can consistently enhance Mixtral 8x7B’s reasoning capability across various domains. For example, it improves the accuracy on GSM8K from 61.79 to 66.94.
Moreover, combining SCMoE with self-consistency yields additional gains, increasing major@20 accuracy from 75.59 to 78.31. Chufan Shi, Cheng Yang 0002, Jiahao Wang 0005, Taiqiang Wu, Siheng Li, Deng Cai 0002, Yujiu Yang 0001, Yu Meng 0001 |
NeurIPS | 7 |
| 2024 | Exploring Dense Retrieval for Dialogue Response SelectionabstractRecent progress in deep learning has continuously improved the accuracy of dialogue response selection. However, in real-world scenarios, the high computation cost forces existing dialogue response selection models to rank only a small number of candidates, recalled by a coarse-grained model, precluding many high-quality candidates. To overcome this problem, we present a novel and efficient response selection model and a set of tailor-designed learning strategies to train it effectively. The proposed model consists of a dense retrieval module and an interaction layer, which could directly select the proper response from a large corpus. We conduct re-rank and full-rank evaluations on widely used benchmarks to evaluate our proposed model. Extensive experimental results demonstrate that our proposed model notably outperforms the state-of-the-art baselines on both re-rank and full-rank evaluations. Moreover, human evaluation results show that the response quality could be improved further by enlarging the candidate pool with nonparallel corpora. In addition, we also release high-quality benchmarks that are carefully annotated for more accurate dialogue response selection evaluation. All source codes, datasets, model parameters, and other related resources have been publicly available. 1 Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Yixuan Su, Heyan Huang, Xianling Mao |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Specialist or Generalist? Instruction Tuning for Specific NLP TasksabstractThe potential of large language models (LLMs) to simultaneously perform a wide range of natural language processing (NLP) tasks has been the subject of extensive research.Although instruction tuning has proven to be a data-efficient method for transforming LLMs into such generalist models, their performance still lags behind specialist models trained exclusively for specific tasks.In this paper, we investigate whether incorporating broadcoverage generalist instruction tuning can contribute to building a specialist model.We hypothesize that its efficacy depends on task specificity and skill requirements.Our experiments assess four target tasks with distinct coverage levels, revealing that integrating generalist instruction tuning consistently enhances model performance when the task coverage is broad.The effect is particularly pronounced when the amount of task-specific training data is limited.Further investigation into three target tasks focusing on different capabilities demonstrates that generalist instruction tuning improves understanding and reasoning abilities.However, for tasks requiring factual knowledge, generalist data containing hallucinatory information may negatively affect the model's performance.Overall, our work provides a systematic guide for developing specialist models with general instruction tuning.Our code and other related resources can be found at https://github.com/DavidFanzz/ Generalist_or_Specialist. Chufan Shi, Yixuan Su, Cheng Yang 0002, Yujiu Yang 0001, Deng Cai 0002 |
EMNLP | 5 |
| 2023 | Copy is All You Need
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Heyan Huang, Xianling Mao |
ICLR | 2 |
| 2023 | Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data PerspectiveabstractThere are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning. Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su |
NeurIPS | 4 |
| 2022 | Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemabstractPre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems.Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhead.In this study, we present PPTOD, a unified plug-andplay model for task-oriented dialogue.In addition, we introduce a new dialogue multi-task pre-training strategy that allows the model to learn the primary TOD task completion skills from heterogeneous dialog corpora.We extensively test our model on three benchmark TOD tasks, including end-to-end dialogue modelling, dialogue state tracking, and intent classification.Experimental results show that PPTOD achieves new state of the art on all evaluated tasks in both high-resource and lowresource scenarios.Furthermore, comparisons against previous SOTA methods show that the responses generated by PPTOD are more factually correct and semantically coherent as judged by human annotators. 1 Yixuan Su, Lei Shu 0004, Elman Mansimov, Arshit Gupta, Deng Cai 0002, Yi-An Lai |
ACL (1) | 5 |
| 2022 | Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationabstractWe introduce a new method to improve existing multilingual sentence embeddings with Abstract Meaning Representation (AMR).Compared with the original textual input, AMR is a structured semantic representation that presents the core concepts and relations in a sentence explicitly and unambiguously.It also helps reduce surface variations across different expressions and languages.Unlike most prior work that only evaluates the ability to measure semantic similarity, we present a thorough evaluation of existing multilingual sentence embeddings and our improved versions, which include a collection of five transfer tasks in different downstream applications.Experiment results show that retrofitting multilingual sentence embeddings with AMR leads to better state-of-the-art performance on both semantic textual similarity and transfer tasks.Our codebase and evaluation scripts Deng Cai 0002, Xin Li 0056, Jackie C. S. Ho, Lidong Bing, Wai Lam |
EMNLP | 1 |
| 2022 | Linearizing Transformer with Key-Value MemoryabstractEfficient transformer variants with linear time complexity have been developed to mitigate the quadratic computational overhead of the vanilla transformer.Among them are lowrank projection methods such as Linformer and kernel-based Transformers.Despite their unique merits, they usually suffer from a performance drop comparing with the vanilla transformer on many sequence generation tasks, and often fail to obtain computation gain when the generation is short.We propose Mem-Sizer, an approach towards closing the performance gap while improving the efficiency even with short generation.It projects the source sequences into lower dimension representations like Linformer, while enjoying efficient recurrent-style incremental computation similar to kernel-based transformers.This yields linear computation time and constant memory complexity at inference time.Mem-Sizer also employs a lightweight multi-head mechanism which renders the computation as light as a single-head model.We demonstrate that MemSizer provides an improved balance between efficiency and accuracy over the vanilla transformer and other efficient transformer variants in three typical sequence generation tasks, including machine translation, abstractive text summarization, and language modeling.Our code is released at https: //github.com/jcyk/memsizer Yizhe Zhang 0002, Deng Cai 0002 |
EMNLP | 2 |
| 2022 | Automatic Prosody Annotation with Pre-Trained Text-Speech ModelabstractProsodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability.However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming.In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders.This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}.The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.Code is released 1 . Ziqian Dai, Jianwei Yu 0001, Yan Wang 0060, Nuo Chen 0001, Yanyao Bian, Guangzhi Li, Deng Cai 0002, Dong Yu 0001 |
INTERSPEECH | 7 |
| 2022 | ASR-Robust Natural Language Understanding on ASR-GLUE dataset
Lingyun Feng, Jianwei Yu 0001, Yan Wang 0060, Songxiang Liu, Deng Cai 0002, Hai-Tao Zheng 0002 |
INTERSPEECH | 5 |
| 2022 | Measuring and Reducing Model Update Regression in Structured Prediction for NLPabstractRecent advance in deep learning has led to rapid adoption of machine learning based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward compatibility requires that the new model does not regress on cases that were correctly handled by its predecessor. This work studies model update regression in structured prediction tasks. We choose syntactic dependency parsing and conversational semantic parsing as representative examples of structured prediction tasks in NLP. First, we measure and analyze model update regression in different model update settings. Next, we explore and benchmark existing techniques for reducing model update regression including model ensemble and knowledge distillation. We further propose a simple and effective method, Backward-Congruent Re-ranking (BCR), by taking into account the characteristics of structured output. Experiments show that BCR can better mitigate model update regression than model ensemble and knowledge distillation approaches. Deng Cai 0002, Elman Mansimov, Yi-An Lai, Yixuan Su, Lei Shu 0004 |
NeurIPS | 1 |
| 2022 | Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationabstractWhile large-scale neural language models, such as GPT2 and BART,have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in the human corpus (e.g., 0.02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probability of repetitive tokens and their previous repetitions in context. Through our quantitative experiments, we find that 1) Models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from synthetic repetitive data. Although our method is motivated by mitigating repetitions, our experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method. Jin Xu 0010, Xiaojiang Liu, Jianhao Yan, Deng Cai 0002, Jian Li 0015 |
NeurIPS | 4 |
| 2022 | Recent Advances in Retrieval-Augmented Text GenerationabstractRecently retrieval-augmented text generation has achieved state-of-the-art performance in many NLP tasks and has attracted increasing attention of the NLP and IR community, this tutorial thereby aims to present recent advances in retrieval-augmented text generation comprehensively and comparatively. It firstly highlights the generic paradigm of retrieval-augmented text generation, then reviews notable works for different text generation tasks including dialogue generation, machine translation, and other generation tasks, and finally points out some limitations and shortcomings to facilitate future research. Deng Cai 0002, Yan Wang 0060, Lemao Liu, Shuming Shi 0001 |
SIGIR | 1 |
| 2022 | Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA EsportabstractWith the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability. Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu |
IEEE Trans. Games | 4 |
| 2021 | Neural Machine Translation with Monolingual Translation MemoryabstractDeng Cai, Yan Wang, Huayang Li, Wai Lam, Lemao Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Deng Cai 0002, Yan Wang 0060, Wai Lam, Lemao Liu |
ACL/IJCNLP (1) | 1 |
| 2021 | Dialogue Response Selection with Hierarchical Curriculum LearningabstractYixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060 |
ACL/IJCNLP (1) | 2 |
| 2021 | Non-Autoregressive Text Generation with Pre-trained Language ModelsabstractYixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, Nigel Collier. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yixuan Su, Deng Cai 0002, Yan Wang 0060, David Vandyke, Simon Baker, Piji Li, Nigel Collier |
EACL | 2 |
| 2021 | PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval MemoryabstractThe ability of dialogue systems to express pre-specified style during conversations has a direct, positive impact on their usability and user satisfaction. While it has attracted much research interest, existing methods often generate stylistic responses at the cost of content quality. In this work, we introduce a prototype-to-style (PS) framework to tackle the challenge of stylistic dialogue generation. The proposed framework first exploits an Information Retrieval (IR) system and extracts a response prototype from the retrieved response. A stylistic response generator then takes the response prototype and the desired style as input to produce a high-quality and stylistic response. To effectively train the proposed model and imitate the real testing environment, we introduce a new style-aware learning objective and a denoising learning strategy. Results on three benchmark datasets (gender, emotion, and sentiment) from two languages demonstrate that the proposed approach significantly outperforms existing baselines both in terms of in-domain and cross-domain evaluations. Yixuan Su, Yan Wang 0060, Deng Cai 0002, Simon Baker, Anna Korhonen, Nigel Collier |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Graph Transformer for Graph-to-Sequence LearningabstractThe dominant graph-to-sequence transduction models employ graph neural networks for graph representation learning, where the structural information is reflected by the receptive field of neurons. Unlike graph neural networks that restrict the information exchange between immediate neighborhood, we propose a new model, known as Graph Transformer, that uses explicit relation encoding and allows direct communication between two distant nodes. It provides a more efficient way for global graph structure modeling. Experiments on the applications of text generation from Abstract Meaning Representation (AMR) and syntax-based neural machine translation show the superiority of our proposed model. Specifically, our model achieves 27.4 BLEU on LDC2015E86 and 29.7 BLEU on LDC2017T10 for AMR-to-text generation, outperforming the state-of-the-art results by up to 2.2 points. On the syntax-based translation tasks, our model establishes new single-model state-of-the-art BLEU scores, 21.3 for English-to-German and 14.1 for English-to-Czech, improving over the existing best results, including ensembles, by over 1 BLEU. Deng Cai 0002, Wai Lam |
AAAI | 1 |
| 2020 | AMR Parsing via Graph-Sequence Iterative InferenceabstractWe propose a new end-to-end model that treats AMR parsing as a series of dual decisions on the input sequence and the incrementally constructed graph. At each time step, our model performs multiple rounds of attention, reasoning, and composition that aim to answer two critical questions: (1) which part of the input sequence to abstract; and (2) where in the output graph to construct the new concept. We show that the answers to these two questions are mutually causalities. We design a model based on iterative inference that helps achieve better answers in both perspectives, leading to greatly improved parsing accuracy. Our experimental results significantly outperform all previously reported Smatch scores by large margins. Remarkably, without the help of any large-scale pre-trained language model (e.g., BERT), our model already surpasses previous state-of-the-art using BERT. With the help of BERT, we can push the state-of-the-art results to 80.2% on LDC2017T10 (AMR 2.0) and 75.4% on LDC2014T12 (AMR 1.0). Deng Cai 0002, Wai Lam |
ACL | 1 |
| 2020 | The World is Not Binary: Learning to Rank with Grayscale Data for Dialogue Response SelectionabstractResponse selection plays a vital role in building retrieval-based conversation systems.Despite that response selection is naturally a learning-to-rank problem, most prior works take a point-wise view and train binary classifiers for this task: each response candidate is labeled either relevant (one) or irrelevant (zero).On the one hand, this formalization can be sub-optimal due to its ignorance of the diversity of response quality.On the other hand, annotating grayscale data for learning-to-rank can be prohibitively expensive and challenging.In this work, we show that grayscale data can be automatically constructed without human effort.Our method employs off-the-shelf response retrieval models and response generation models as automatic grayscale data generators.With the constructed grayscale data, we propose multi-level ranking objectives for training, which can (1) teach a matching model to capture more fine-grained context-response relevance difference and (2) reduce the traintest discrepancy in terms of distractor strength.Our method is simple, effective, and universal.Experiments on three benchmark datasets and four state-of-the-art matching models show that the proposed approach brings significant and consistent performance improvements. Zibo Lin, Deng Cai 0002, Yan Wang 0060, Xiaojiang Liu, Hai-Tao Zheng 0002, Shuming Shi 0001 |
EMNLP (1) | 2 |
| 2020 | Describe What to Change: A Text-guided Unsupervised Image-to-image Translation ApproachabstractManipulating visual attributes of images through human-written text is a very challenging task. On the one hand, models have to learn the manipulation without the ground truth of the desired output. On the other hand, models have to deal with the inherent ambiguity of natural language. Previous research usually requires either the user to describe all the characteristics of the desired image or to use richly-annotated image captioning datasets. In this work, we propose a novel unsupervised approach, based on image-to-image translation, that alters the attributes of a given image through a command-like sentence such as "change the hair color to black". Contrarily to state-of-the-art approaches, our model does not require a human-annotated dataset nor a textual description of all the attributes of the desired image, but only those that have to be modified. Our proposed model disentangles the image content from the visual attributes, and it learns to modify the latter using the textual description, before generating a new image from the content and the modified attribute representation. Because text might be inherently ambiguous (blond hair may refer to different shadows of blond, e.g. golden, icy, sandy), our method generates multiple stochastic versions of the same translation. Experiments show that the proposed model achieves promising performances on two large-scale public datasets: CelebA and CUB. We believe our approach will pave the way to new avenues of research combining textual and speech commands with visual attributes. Marco De Nadai, Deng Cai 0002, Xavier Alameda-Pineda, Nicu Sebe, Bruno Lepri |
ACM Multimedia | 3 |
| 2020 | Neural Machine Translation With Noisy Lexical ConstraintsabstractIn neural machine translation, lexically constrained decoding generates translation outputs strictly including the constraints predefined by users, and it is beneficial to improve translation quality at the cost of more decoding overheads if the constraints are perfect. Unfortunately, those constraints may contain mistakes in real-world situations and incorrect constraints will undermine lexically constrained decoding. In this article, we propose a novel framework that is capable of improving the translation quality even if the constraints are noisy. The key to our framework is to treat the lexical constraints as external memories. More concretely, it encodes the constraints by a memory encoder and then leverages the memories by a memory integrator. Experiments demonstrate that our framework can not only deliver substantial BLEU gains in handling noisy constraints, but also achieve speedup in decoding. These results motivate us to apply our models to a new scenario where the constraints are generated without the help of users. Experiments show that our models can indeed improve the translation quality with the automatically generated constraints. Guoping Huang, Deng Cai 0002, Lemao Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Unsupervised Learning Helps Supervised Neural Word SegmentationabstractBy exploiting unlabeled data for further performance improvement for Chinese word segmentation, this work makes the first attempt at exploring adding unsupervised segmentation information into neural supervised segmenter. We survey various effective strategies, including extending the character embedding, augmenting the word score and applying multi-task learning, for leveraging unsupervised information derived from abundant unlabeled data. Experiments on standard data sets show that the explored strategies indeed improve the recall rate of out-of-vocabulary words and thus boost the segmentation accuracy. Moreover, the model enhanced by the proposed methods outperforms state-of-theart models in closed test and shows promising improvement trend when adopting three different strategies with the help of a large unlabeled data set. Our thorough empirical study eventually verifies the proposed approach outperforms the widelyused pre-training approach in terms of effectively making use of freely abundant unlabeled data. Xiaobin Wang, Deng Cai 0002, Linlin Li 0001, Hai Zhao 0001, Luo Si |
AAAI | 2 |
| 2019 | Core Semantic First: A Top-down Approach for AMR ParsingabstractDeng Cai, Wai Lam. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Deng Cai 0002, Wai Lam |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Retrieval-guided Dialogue Response Generation via a Matching-to-Generation FrameworkabstractDeng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Deng Cai 0002, Yan Wang 0060, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Charge-Based Prison Term Prediction with Deep Gating NetworkabstractHuajie Chen, Deng Cai, Wei Dai, Zehui Dai, Yadong Ding. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Huajie Chen, Deng Cai 0002, Zehui Dai, Yadong Ding |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Translating Math Word Problem to Expression TreeabstractSequence-to-sequence (SEQ2SEQ) models have been successfully applied to automatic math word problem solving.Despite its simplicity, a drawback still remains: a math word problem can be correctly solved by more than one equations.This non-deterministic transduction harms the performance of maximum likelihood estimation.In this paper, by considering the uniqueness of expression tree, we propose an equation normalization method to normalize the duplicated equations.Moreover, we analyze the performance of three popular SEQ2SEQ models on the math word problem solving.We find that each model has its own specialty in solving problems, consequently an ensemble model is then proposed to combine their advantages.Experiments on dataset Math23K show that the ensemble model with equation normalization significantly outperforms the previous state-of-the-art methods. Lei Wang 0185, Yan Wang 0060, Deng Cai 0002, Dongxiang Zhang, Xiaojiang Liu |
EMNLP | 3 |
| 2017 | Pair-Aware Neural Sentence Modeling for Implicit Discourse Relation Classification
Deng Cai 0002, Hai Zhao 0001 |
IEA/AIE (2) | 1 |
| 2017 | A Hybrid Model for Chinese Spelling CheckabstractSpelling check for Chinese has more challenging difficulties than that for other languages. A hybrid model for Chinese spelling check is presented in this article. The hybrid model consists of three components: one graph-based model for generic errors and two independently trained models for specific errors. In the graph model, a directed acyclic graph is generated for each sentence, and the single-source shortest-path algorithm is performed on the graph to detect and correct general spelling errors at the same time. Prior to that, two types of errors over functional words (characters) are first solved by conditional random fields: the confusion of “在” ( at ) (pinyin is zai in Chinese), “再” ( again , more , then ) (pinyin: zai ) and “的” ( of ) (pinyin: de ), “地” (- ly , adverb-forming particle) (pinyin: de ), and “得” ( so that , have to ) (pinyin: de ). Finally, a rule-based model is exploited to distinguish pronoun usage confusion: “她” ( she ) (pinyin: ta ), “他” ( he ) (pinyin: ta ), and some other common collocation errors. The proposed model is evaluated on the standard datasets released by the SIGHAN Bake-off shared tasks, giving state-of-the-art results. Hai Zhao 0001, Deng Cai 0002, Yang Xin 0005, Zhongye Jia |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Neural Word Segmentation Learning for ChineseabstractMost previous approaches to Chinese word segmentation formalize this problem as a character-based sequence labeling task so that only contextual information within fixed sized local windows and simple interactions between adjacent tags can be captured.In this paper, we propose a novel neural framework which thoroughly eliminates context windows and can utilize complete segmentation history.Our model employs a gated combination neural network over characters to produce distributed representations of word candidates, which are then given to a long shortterm memory (LSTM) language scoring model.Experiments on the benchmark datasets show that without the help of feature engineering as most existing approaches, our models achieve competitive or better performances with previous stateof-the-art methods. Deng Cai 0002, Hai Zhao 0001 |
ACL (1) | 1 |