EDBT 2026 Demo / reviewers in the wild / expert
Juntao Li 0005
dblp:32/971-5
· DBLP profile ↗
74ranked-venue papers
9as first author
61since 2021 · last 2026
0000-0002-6286-7529ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 7 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context ReasoningabstractHigh-quality multi-hop instruction data is critical for enhancing the reasoning capabilities of large language models (LLMs) in complex long-context scenarios, e.g., long-form reasoning. Nevertheless, there is currently a notable scarcity of such datasets within the community, and existing data synthesis approaches typically fail to provide explicit modeling of intermediate reasoning steps, resulting in unverifiable and potentially erroneous samples. To mitigate above issue, we design the Concept-Graph based Multi-hop Instructions Synthesis (CGMIS) framework, which constructs long-form reasoning paths via concept graph traversal and automatically generates verifiable multi-hop data. The CGMIS framework not only guarantees the accuracy and verifiability of the synthesized data but also enables the construction of high-quality multi-hop instruction datasets from arbitrary corpora. Experiments show that fine-tuning with CGMIS-generated data achieves state-of-the-art performance across 13 long-context reasoning tasks on various models, using only 10% of the data volume required by existing methods. Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Qiaoming Zhu |
AAAI | 3 |
| 2026 | Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-ReflectionabstractLinxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi, Zhang Yijun, Ji Qi, Wenpeng Zhu, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi 0005, Zhang Yijun, Wenpeng Zhu, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 10 |
| 2026 | DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-RewardabstractXiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Zhe Zhao, Kehai Chen, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Kehai Chen, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 10 |
| 2026 | Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsabstractYi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yi Su 0006, Dian Yu 0001, Linfeng Song, Juntao Li 0005, Haitao Mi, Zhaopeng Tu, Min Zhang 0005, Dong Yu 0001 |
ACL (1) | 4 |
| 2026 | IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural ThinkingabstractZechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang 0005 |
ACL (1) | 4 |
| 2026 | BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMsabstractYue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Xin He, Qu Yang, Qingxuan Jiang, Fanghua Ye, Juntao Li, Zhaopeng Tu, Xiaolong Li, Liefeng Bo, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yue Wang 0039, Ruotian Ma, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Qu Yang, Qingxuan Jiang, Fanghua Ye 0004, Juntao Li 0005, Zhaopeng Tu, Liefeng Bo, Min Zhang 0005 |
ACL (1) | 13 |
| 2026 | When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient ReasoningabstractLarge reasoning models (LRMs) have achieved remarkable performance in complex reasoning tasks, driven by their powerful inference-time scaling capability.However, LRMs often suffer from overthinking, which results in substantial computational redundancy and significantly reduces efficiency.Early-exit methods aim to mitigate this issue by terminating reasoning once sufficient evidence has been generated, yet existing approaches mostly rely on handcrafted or empirical indicators that are unreliable and impractical.In this work, we introduce Dynamic Thought Sufficiency in Reasoning (DTSR), a novel framework for efficient reasoning that enables the model to dynamically assess the sufficiency of its chain-of-thought (CoT) and determine the optimal point for early exit.Inspired by human metacognition, DTSR operates in two stages: (1) Reflection Signal Monitoring, which identifies reflection signals as potential cues for early exit, and (2) Thought Sufficiency Check, which evaluates whether the current CoT is sufficient to derive the final answer.Experimental results on the Qwen3 models show that DTSR reduces reasoning length by 28.9%-34.9%with minimal performance loss, effectively mitigating overthinking.We further discuss overconfidence in LRMs and self-evaluation paradigms, providing valuable insights for early-exit reasoning. Yang Xiang 0003, Yixin Ji, Ruotao Xu, Zheming Yang, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 6 |
| 2026 | ALD2: Adaptive layer-wise denoising decoding for hallucinations mitigation in large vision-language models
Yuechi Zhou, Morunliu Yang, Juntao Li 0005, Siwei Feng |
Inf. Process. Manag. | 4 |
| 2026 | Data Foundations of Long-Context Language Models: A Survey
Zechen Sun, Zhaochen Su, Zecheng Tang, Juntao Li 0005, Wenliang Chen, Min Zhang 0005 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | Accurate KV Cache Quantization with Outlier Tokens TracingabstractThe impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput. Yi Su 0006, Yuechi Zhou, Quantong Qiu, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | \mathcalA³: Automatic Alignment Framework for Attributed Text GenerationabstractAttributed text generation aims to enhance the reliability of content generated from large language models by providing citations for each claim, which thereby enables users to easily verify the correctness of the responses.However, the scarcity of high-quality training samples presents a significant challenge in aligning large language models to generate texts with citations, revealing considerable room for improvement in existing attribution systems.Besides, existing approaches of aligning large language models to follow user instructions can lead to an undue emphasis on irrelevant documents, which in turn reduces the quality of responses.To address the above problems, we propose Automatic Alignment Framework for Attributed Text Generation (A 3 ), a novel framework designed to automatically generate highquality attributed query-response pairs for both supervised fine-tuning and preference optimization stages without human annotation.With the help of A 3 , Mistral-7B can achieve a citation recall of 84.4 and a precision of 87.0 precision on ASQA, which notably surpasses GPT-4's citation recall of 73.0 and precision of 76.5. 1 Yue Wang 0039, Haoke Zhang, Juntao Li 0005, Jinxiong Chang, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | Unleashing LLM Reasoning Capability via Scalable Question Synthesis from ScratchabstractImproving the mathematical reasoning capabilities of Large Language Models (LLMs) is critical for advancing artificial intelligence. However, access to extensive, diverse, and high-quality reasoning datasets remains a significant challenge, particularly for the open-source community. In this paper, we propose ScaleQuest, a novel, scalable, and cost-effective data synthesis method that enables the generation of large-scale mathematical reasoning datasets using lightweight 7B-scale models. ScaleQuest introduces a two-stage question-tuning process comprising Question Fine-Tuning (QFT) and Question Preference Optimization (QPO) to unlock the question generation capabilities of problem-solving models. By generating diverse questions from scratch – without relying on powerful proprietary models or seed data – we produce a dataset of 1 million problem-solution pairs. Our experiments demonstrate that models trained on our data outperform existing open-source datasets in both in-domain and out-of-domain evaluations. Furthermore, our approach shows continued performance improvement as the volume of training data increases, highlighting its potential for ongoing data scaling. The extensive improvements observed in code reasoning tasks demonstrate the generalization capabilities of our proposed method. Our work provides the open-source community with a practical solution to enhance the mathematical reasoning abilities of LLMs. Yuyang Ding, Xinyu Shi 0005, Xiaobo Liang, Juntao Li 0005, Zhaopeng Tu, Qiaoming Zhu, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Generative Reward Modeling via Synthetic Criteria Preference LearningabstractGenerative Reward Models (GenRMs) leverage synthesized Chains of Thought (CoT) to reduce the need for massive labeled data, but this approach introduces risks of overoptimization due to the inability to guarantee the correctness of the CoTs.Identifying and optimizing unexpected behaviors within these synthesized CoT remains a challenge, as it heavily depends on precise annotations of intermediate behavior, similar to process supervision.In this work, we introduce a criteria-based preference tree for reward modeling, where each path in the tree represents a reasoning trajectory based on synthesized criteria.Crucially, each reasoning trajectory can be independently optimized through RL algorithm.These fine-grained process reward signals are derived from the inferencetime computations and predefined rules, eliminating the need for human supervision.In experiments, SyncPL 1 showed significant improvements over baselines on multiple human preference benchmarks.We further demonstrate that synthesized data can be learned using a long CoT format, analogous to an o1-like model, further enhancing performance while keeping stability and efficiency during training. Xiaobo Liang, Haoke Zhang, Juntao Li 0005, Kehai Chen, Qiaoming Zhu, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | L-CiteEval: A Suite for Evaluating Fidelity of Long-context ModelsabstractLong-context models (LCMs) have witnessed remarkable advancements in recent years, facilitating real-world tasks like long-document QA.The success of LCMs is founded on the hypothesis that the model demonstrates strong fidelity, enabling it to respond based on the provided long context rather than relying solely on the intrinsic knowledge acquired during pre-training.Yet, in this paper, we find that open-sourced LCMs are not as faithful as expected.We introduce L-CiteEval, an out-of-the-box suite that can assess both generation quality and fidelity in long-context understanding tasks.It covers 11 tasks with context lengths ranging from 8K to 48K and a corresponding automatic evaluation pipeline.Evaluation of 11 cutting-edge closed-source and open-source LCMs indicates that, while there are minor differences in their generation, open-source models significantly lag behind closed-source counterparts in terms of fidelity.Furthermore, we analyze the benefits of citation generation for LCMs from both the perspective of explicit model output and the internal attention mechanism 1 . Zecheng Tang, Keyan Zhou, Juntao Li 0005, Baibei Ji, Jianye Hou, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | An Empirical Study of Iterative Refinements for Non-autoregressive TranslationabstractIterative non-autoregressive (NAR) models share a spirit of mixed autoregressive (AR) and fully NAR models, seeking a balance between generation quality and inference efficiency.These models have recently demonstrated impressive performance in varied generation tasks, surpassing the autoregressive Transformer.However, they also face several challenges that impede further development.In this work, we target building more efficient and competitive iterative NAR models.Firstly, we produce two simple metrics to identify the potential problems existing in current refinement processes, and look back on the various iterative NAR models to find the key factors for realizing our purpose.Subsequently, based on the analyses of the limitations of previous inference algorithms, we propose a simple yet effective strategy to conduct efficient refinements without performance declines.Experiments on five widely used datasets show that our final models set the new state-of-the-art performance compared to all previous NAR models, even with fewer decoding steps, and outperform AR Transformer by around one BLEU on average.Our codes and models are available on Github 1 . Yisheng Xiao, Pei Guo, Zechen Sun, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Unveiling the Potential of BERT-family: A New Recipe for Building Scalable, General and Competitive Large Language ModelsabstractBERT-family have been increasingly explored for adaptation to scenarios beyond language understanding tasks, with more recent efforts focused on enabling them to become good instruction followers.These explorations have endowed BERT-family with new roles and human expectations, showcasing their potential on par with current state-of-the-art (SOTA) large language models (LLMs).However, several certain shortcomings in previous BERT-family, such as the relatively sub-optimal training corpora, learning procedure, and model architecture, all impede the further advancement of these models for serving as general and competitive LLMs.Therefore, we aim to address these deficiencies in this paper.Our study not only introduces a more suitable pre-training task that helps BERT-family excel in wider applications to realize generality but also explores the integration of cutting-edge technologies into our model to further enhance their capabilities.Our final models, termed Bidirectional General Language Models (BiGLM), exhibit performance levels comparable to current SOTA LLMs across a spectrum of tasks.Moreover, we conduct detailed analyses to study the effects of scaling and training corpora for BiGLM.To the best of our knowledge, our work represents the early attempt to offer a recipe for building novel types of scalable, general, and competitive LLMs that diverge from current autoregressive modeling methodology.Our codes and models are available on Github 1 . Yisheng Xiao, Juntao Li 0005, Wenpeng Hu, Zhunchen Luo, Min Zhang 0005 |
ACL (1) | 2 |
| 2025 | A Survey of Generative Information ExtractionabstractGenerative information extraction (Generative IE) aims to generate structured text sequences from unstructured text using a generative framework. Scaling in model size yields variations in adaptation and generalization, and also drives fundamental shifts in the techniques and approaches used within this domain. In this survey, we first review generative information extraction (IE) methods based on pre-trained language models (PLMs) and large language models (LLMs), focusing on their adaptation and generalization capabilities. We also discuss the connection between these methods and these two aspects. Furthermore, to balance task performance with the substantial computational demands associated with LLMs, we emphasize the importance of model collaboration. Finally, given the advanced capabilities of LLMs, we explore methods for integrating diverse IE tasks into unified models. Zikang Zhang, Wangjie You, Tianci Wu, Juntao Li 0005, Min Zhang 0005 |
COLING | 5 |
| 2025 | Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional VerificationabstractRecent works have revealed the great potential of speculative decoding in accelerating the autoregressive generation process of large language models.The success of these methods relies on the alignment between draft candidates and the sampled outputs of the target model.Existing methods mainly achieve draft-target alignment with training-based methods, e.g., EAGLE, Medusa, involving considerable training costs.In this paper, we present a trainingfree alignment-augmented speculative decoding algorithm.We propose alignment sampling, which leverages output distribution obtained in the prefilling phase to provide more aligned draft candidates.To further benefit from highquality but non-aligned draft candidates, we also introduce a simple yet effective flexible verification strategy.Through an adaptive probability threshold, our approach can improve generation accuracy while further improving inference efficiency.Experiments on 8 datasets (including question answering, summarization and code completion tasks) show that our approach increases the average generation score by 3.3 points for the LLaMA3 model.Our method achieves a mean acceptance length up to 2.39 and speed up generation by 2.23×. Zhenxu Tian, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 3 |
| 2025 | Beware of Calibration Data for Pruning Large Language ModelsabstractAs large language models (LLMs) are widely applied across various fields, model
compression has become increasingly crucial for reducing costs and improving
inference efficiency. Post-training pruning is a promising method that does not
require resource-intensive iterative training and only needs a small amount of
calibration data to assess the importance of parameters. Recent research has enhanced post-training pruning from different aspects but few of them systematically
explore the effects of calibration data, and it is unclear if there exist better calibration data construction strategies. We fill this blank and surprisingly observe that
calibration data is also crucial to post-training pruning, especially for high sparsity. Through controlled experiments on important influence factors of calibration
data, including the pruning settings, the amount of data, and its similarity with
pre-training data, we observe that a small size of data is adequate, and more similar data to its pre-training stage can yield better performance. As pre-training data
is usually inaccessible for advanced LLMs, we further provide a self-generating
calibration data synthesis strategy to construct feasible calibration data. Experimental results on recent strong open-source LLMs (e.g., DCLM, and LLaMA-3)
show that the proposed strategy can enhance the performance of strong pruning
methods (e.g., Wanda, DSnoT, OWL) by a large margin (up to 2.68%). Yixin Ji, Yang Xiang 0003, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 3 |
| 2025 | Revealing and Mitigating Over-Attention in Knowledge EditingabstractLarge Language Models~(LLMs) have demonstrated superior performance across a wide range of tasks, but they still exhibit undesirable errors due to incorrect knowledge learned from the training data. To avoid this, knowledge editing methods emerged to precisely edit the specific model knowledge via efficiently modifying a very small percentage of parameters. However, those methods can lead to the problem of **Specificity Failure**, where the existing knowledge and capabilities are severely degraded due to editing.
Our preliminary indicates that Specificity Failure primarily stems from the model's attention heads assigning excessive attention scores to entities related to the edited knowledge, thereby unduly focusing on specific snippets within the context, which we denote as the **Attention Drift** phenomenon.
To mitigate such Attention Drift issue, we introduce a simple yet effective method **S**elective **A**ttention **D**rift **R**estriction(**SADR**), which introduces an additional regularization term during the knowledge editing process to restrict changes in the attention weight distribution, thereby preventing undue focus on the edited entity.
Experiments on five frequently-used strong LLMs demonstrate the effectiveness of our method, where SADR can significantly mitigate Specificity Failure in the predominant knowledge editing tasks. Pinzheng Wang, Zecheng Tang, Keyan Zhou, Juntao Li 0005, Qiaoming Zhu, Min Zhang 0005 |
ICLR | 4 |
| 2025 | LOGO - Long cOntext aliGnment via efficient preference OptimizationabstractLong-context models (LCMs) have shown great potential in processing long input sequences (even more than 100M tokens) conveniently and effectively. With significant progress, recent research has pointed out that LCMs can accurately locate token-level salient information within the context. Yet, the generation performance of these LCMs is far from satisfactory and might result in misaligned responses, such as hallucinations. To enhance the generation capability of LCMs, existing works have investigated the effects of data size and quality for both pre-training and instruction tuning. Though achieving meaningful improvement, previous methods fall short in either effectiveness or efficiency. In this paper, we introduce LOGO (Long cOntext aliGnment via efficient preference Optimization), a training strategy that first introduces preference optimization for long-context alignment. To overcome the GPU memory-bound issue caused by the long sequence, LOGO employs a reference-free preference optimization strategy and adopts a position synthesis method to construct the training data. By training with only 0.3B data on a single 8 x A800 GPU machine for 16 hours, LOGO allows the Llama-3-8B-Instruct-80K model to achieve comparable performance with GPT-4 in real-world long-context tasks while preserving the model's original capabilities on other tasks, e.g., language modeling and MMLU. Moreover, LOGO can extend the model's context window size while enhancing its generation performance. Zecheng Tang, Zechen Sun, Juntao Li 0005, Qiaoming Zhu, Min Zhang 0005 |
ICML | 3 |
| 2025 | Improving Rationality in the Reasoning Process of Language Models through Self-playing GameabstractLarge language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding.
However, recent studies indicate that even the best models lack true comprehension of their reasoning processes.
In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models.
We design a $\textit{\textbf{C}ritic-\textbf{D}iscernment \textbf{G}ame}~(\textbf{CDG})$ in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution.
These critiques either aim to assist or mislead the prover.
The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback.
Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process. Pinzheng Wang, Juntao Li 0005, Zecheng Tang, Haijia Gui, Min Zhang 0005 |
ICML | 2 |
| 2025 | Taming the Titans: A Survey of Efficient LLM Inference ServingabstractLarge Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by their vast number of parameters, combined with the high computational demands of the attention mechanism, poses significant challenges in achieving low latency and high throughput for LLM inference services. Recent advancements, driven by groundbreaking research, have significantly accelerated progress in this field. This paper provides a comprehensive survey of these methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, and emerging scenarios. At the instance level, we review model placement, request scheduling, decoding length prediction, storage management, and the disaggregation paradigm. At the cluster level, we explore GPU cluster deployment, multi-instance load balancing, and cloud service solutions. Additionally, we discuss specific tasks, modules, and auxiliary methods in emerging scenarios. Finally, we outline potential research directions to further advance the field of LLM inference serving. Ranran Zhen, Juntao Li 0005, Yixin Ji, Zhenlin Yang, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
INLG | 2 |
| 2025 | SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningabstractProcess reward models (PRMs) offer fine-grained, step-level evaluations that facilitate deeper reasoning processes in large language models (LLMs), proving effective in complex tasks like mathematical reasoning.
However, developing PRMs is challenging due to the high cost and limited scalability of human-annotated data.
Synthetic data from Monte Carlo (MC) estimation is a promising alternative but suffers from a high noise ratio, which can cause overfitting and hinder large-scale training.
In this work, we conduct a preliminary study on the noise distribution in synthetic data from MC estimation, identifying that annotation models tend to both underestimate and overestimate step correctness due to limitations in their annotation capabilities.
Building on these insights, we propose {\bf S}elf-Denoising Monte {\bf C}arlo {\bf An}notation (\textsc{Scan}), an efficient data synthesis and noise-tolerant learning framework.
Our key findings indicate that:
(1) Even lightweight models (e.g., 1.5B parameters) can produce high-quality annotations through self-denoising strategy, enabling PRMs to achieve superior performance with only 6\% the inference cost required by vanilla MC estimation.
(2) With our robust learning strategy, PRMs can effectively learn from this weak supervision, achieving a 39.2 F1 score improvement (from 19.9 to 59.1) in ProcessBench.
Despite using only a compact synthetic dataset, our models surpass strong baselines, including those trained on large-scale human-annotated datasets such as PRM800K.
Furthermore, performance continues to improve as we scale up the synthetic data, highlighting the potential of \textsc{Scan} for scalable, cost-efficient, and robust PRM training. Yuyang Ding, Xinyu Shi 0005, Juntao Li 0005, Xiaobo Liang, Zhaopeng Tu, Min Zhang 0005 |
NeurIPS | 3 |
| 2025 | XIFBench: Evaluating Large Language Models on Multilingual Instruction FollowingabstractLarge Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking fine-grained constraint analysis across diverse linguistic contexts. We introduce XIFBench, a comprehensive constraint-based benchmark for evaluating multilingual instruction-following abilities of LLMs, comprising 558 instructions with 0-5 additional constraints across five categories (Content, Style, Situation, Format, and Numerical) in six languages spanning different resource levels. To support reliable and consistent cross-lingual evaluation, we implement three methodological innovations: cultural accessibility annotation, constraint-level translation validation, and requirement-based evaluation using English requirements as semantic anchors across languages. Extensive experiments with various LLMs not only quantify performance disparities across resource levels but also provide detailed insights into how language resources, constraint categories, instruction complexity, and cultural specificity influence multilingual instruction-following. Our code and data are available at https://github.com/zhenyuli801/XIFBench. Kehai Chen, Xuefeng Bai 0001, Yaoyin Zhang, Xuchen Wei, Juntao Li 0005, Min Zhang 0005 |
NeurIPS | 7 |
| 2025 | Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model ReasoningabstractReasoning capability is pivotal for Large Language Models (LLMs) to solve complex tasks, yet achieving reliable and scalable reasoning remains challenging. While Chain-of-Thought (CoT) prompting has become a mainstream approach, existing methods often suffer from uncontrolled generation, insufficient quality, and limited diversity in reasoning paths.
Recent efforts leverage code to enhance CoT by grounding reasoning in executable steps, but such methods are typically constrained to predefined mathematical problems, hindering scalability and generalizability.
In this work, we propose \texttt{Caco} (Code-Assisted Chain-of-ThOught), a novel framework that automates the synthesis of high-quality, verifiable, and diverse instruction-CoT reasoning data through code-driven augmentation. Unlike prior work, \texttt{Caco} first fine-tunes a code-based CoT generator on existing math and programming solutions in a unified code format, then scales the data generation to a large amount of diverse reasoning traces. Crucially, we introduce automated validation via code execution and rule-based filtering to ensure logical correctness and structural diversity, followed by reverse-engineering filtered outputs into natural language instructions and language CoTs to enrich task adaptability. This closed-loop process enables fully automated, scalable synthesis of reasoning data with guaranteed executability.
Experiments on our created \texttt{Caco}-1.3M dataset demonstrate that \texttt{Caco}-trained models achieve strong competitive performance on mathematical reasoning benchmarks, outperforming existing strong baselines. Further analysis reveals that \texttt{Caco}’s code-anchored verification and instruction diversity contribute to superior generalization across unseen tasks. Our work establishes a paradigm for building self-sustaining, trustworthy reasoning systems without human intervention. Honglin Lin, Qizhi Pei, Zhuoshi Pan, Yu Li 0006, Xin Gao 0001, Juntao Li 0005, Conghui He, Lijun Wu 0003 |
NeurIPS | 6 |
| 2025 | Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsabstractLong reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking. Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 9 |
| 2025 | OpenBA: an open-sourced 15B bilingual asymmetric Seq2Seq model pre-trained from scratch
Juntao Li 0005, Zecheng Tang, Yuyang Ding, Pinzheng Wang, Pei Guo, Wangjie You, Wenliang Chen, Guohong Fu, Qiaoming Zhu, Guodong Zhou 0001, Min Zhang 0005 |
Sci. China Inf. Sci. | 1 |
| 2025 | OPT-Tree: Speculative Decoding with Adaptive Draft Tree StructureabstractAbstract Autoregressive language models demonstrate excellent performance in various scenarios. However, the inference efficiency is limited by its one-step-one-word generation mode, which has become a pressing problem recently as the models become increasingly larger. Speculative decoding employs a “draft and then verify” mechanism to allow multiple tokens to be generated in one step, realizing lossless acceleration. Existing methods mainly adopt fixed heuristic draft structures, which do not adapt to different situations to maximize the acceptance length during verification. To alleviate this dilemma, we propose OPT-Tree, an algorithm to construct adaptive and scalable draft trees, which can be applied to any autoregressive draft model. It searches the optimal tree structure that maximizes the mathematical expectation of the acceptance length in each decoding step. Experimental results reveal that OPT-Tree outperforms the existing draft structures and achieves a speed-up ratio of up to 3.2 compared with autoregressive decoding. If the draft model is powerful enough and the node budget is sufficient, it can generate more than ten tokens in a single step. Our code is available at https://github.com/Jikai0Wang/OPT-Tree. Yi Su 0006, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Towards DS-NER: Unveiling and Addressing Latent Noise in Distant AnnotationsabstractDistantly supervised named entity recognition (DS-NER) has emerged as a cheap and convenient alternative to traditional human annotation methods, enabling the automatic generation of training data by aligning text with external resources. Despite the many efforts in noise measurement methods, few works focus on the latent noise distribution between different distant annotation methods. In this work, we explore the effectiveness and robustness of DS-NER by two aspects: (1) distant annotation techniques, which encompasses both traditional rule-based methods and the innovative large language model supervision approach, and (2) noise assessment, for which we introduce a novel framework. This framework addresses the challenges by distinctly categorizing them into theunlabeled-entity problem (UEP)and thenoisy-entity problem (NEP), subsequently providing specialized solutions for each. Our proposed method achieves significant improvements on eight real-world distant supervision datasets originating from three different data sources and involving four distinct annotation techniques, confirming its superiority over current state-of-the-art methods. Yuyang Ding, Juntao Li 0005, Jiajie Xu 0001, Pingfu Chao, Xiaofang Zhou 0001, Min Zhang 0005 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?abstractZhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaochen Su, Juntao Li 0005, Jun Zhang 0069, Tong Zhu 0002, Xiaoye Qu, Pan Zhou 0001, Yan Bowen, Yu Cheng 0001, Min Zhang 0005 |
ACL (1) | 2 |
| 2024 | When and How to Grow? On Efficient Pre-training via Model Growth
Juntao Li 0005, Min Zhang 0005, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai |
ACML | 2 |
| 2024 | Exploring and Mitigating Shortcut Learning for Generative Large Language ModelsabstractRecent generative large language models (LLMs) have exhibited incredible instruction-following capabilities while keeping strong task completion ability, even without task-specific fine-tuning. Some works attribute this to the bonus of the new scaling law, in which the continuous improvement of model capacity yields emergent capabilities, e.g., reasoning and universal generalization. However, we point out that recent LLMs still show shortcut learning behavior, where the models tend to exploit spurious correlations between non-robust features and labels for prediction, which might lead to overestimating model capabilities. LLMs memorize more complex spurious correlations (i.e., task \leftrightarrow feature \leftrightarrow label) compared with that learned from previous pre-training and task-specific fine-tuning paradigm (i.e., feature \leftrightarrow label). Based on our findings, we propose FSLI, a framework for encouraging LLMs to Forget Spurious correlations and Learn from In-context information. Experiments on three tasks show that FSFI can effectively mitigate shortcut learning. Besides, we argue not to overestimate the capabilities of LLMs and conduct evaluations in more challenging and complete test scenarios. Zechen Sun, Yisheng Xiao, Juntao Li 0005, Yixin Ji, Wenliang Chen, Min Zhang 0005 |
LREC/COLING | 3 |
| 2024 | Towards More Realistic Chinese Spell Checking with New Benchmark and Specialized Expert ModelabstractLarge Language Models (LLMs) hold considerable promise for artificial general intelligence, given their intrinsic abilities to accomplish a wide range of open-domain tasks either independently or in tandem with specialized expert models. However, despite these capabilities, the performance of LLMs has yet to be comprehensively evaluated in realistic scenarios. To this end, in this work, we introduce a novel task, the Realistic Chinese Spell Checking (RCSC), to evaluate the effectiveness of existing methods comprehensively. In contrast to existing works that solely address Chinese character misspellings or pinyin conversions, our task aims to convert the realistic Chinese text into the corresponding correct text. The realistic Chinese text may potentially contain both Chinese misspellings and pinyin conversions. We first present the Realistic Chinese Spell Checking Benchmark (RCSCB), which consists of two subsets and contains a total of 581,657 samples. Then, we benchmark the performance of various baselines and find that all the existing methods, including instruction-based LLMs, achieve unsatisfactory results on RCSCB. To further improve the performance on RCSCB, we propose Pinyin-Enhanced Spell Checker (PESC), which is specifically designed to address pinyin-related misspellings. Experimental results demonstrate that PESC can achieve state-of-the-art performance on RCSCB. Despite the progress made, the current state-of-the-art performance is still far from satisfactory. We expect further progress on this crucial and challenging task. Yue Wang 0039, Zilong Zheng, Juntao Li 0005, Jinxiong Chang, Qishen Zhang, Zhongyi Liu 0001, Min Zhang 0005 |
LREC/COLING | 3 |
| 2024 | TabMedBERT: A Tabular Knowledge Enhanced Biomedical Pretrained Language ModelabstractMost existing biomedical language models are trained on plain text with general learning goals such as random word infilling, failing to capture the knowledge in the biomedical corpus sufficiently. Since biomedical articles usually contain many tables summarising the main entities and their relations, in the paper, we propose a Tabular knowledge enhanced bioMedical pretrained language model, called TabMedBERT. Specifically, we align entities between table cells, and article text spans with pre-defined rules. Then we add two table-related self-supervised tasks to integrate tabular knowledge into the language model: Entity Infilling (EI) and Table Cloze Test (TCT). While EI masks tokens within aligned entities in the article, TCT converts aligned entities in the table layout into a cloze text by erasing one entity and prompts the model to extract the appropriate span to fill in the blank. Experimental results demonstrate that TabMedBERT surpasses all competing language models without adding additional parameters, establishing a new state-of-the-art performance of 85.59% (+1.29%) on the BLURB biomedical NLP benchmark and 7 additional information extraction datasets. Moreover, the model architecture for TCT provides a straightforward solution to revise information extraction with paired entities. Lei Geng, Ziqiang Cao, Juntao Li 0005, Wenjie Li 0002, Sujian Li, Yang Yang 0074, Jun Zhang 0069 |
ECAI | 4 |
| 2024 | CMD: a framework for Context-aware Model self-DetoxificationabstractText detoxification aims to minimize the risk of language models producing toxic content.However, existing detoxification methods fail to balance the detoxification effectiveness and generation quality.This issue arises from neglecting the constraints imposed by the context: language models are designed to generate output that closely matches the given context, while detoxification methods strive to ensure the safety of the output, even if it deviates semantically from the context.Given this, we introduce a Context-aware Model self-Detoxification (CMD) framework that pays attention to both the context and the detoxification process, i.e., first detoxifying the context and then making the language model generate along the safe context.Specifically, CMD framework involves two phases: utilizing language models to synthesize data and applying these data for training.We also introduce a toxic contrastive loss that encourages the model generation away from the negative toxic samples.Experiments on various LLMs have verified the effectiveness of our MSD framework, which can yield the best performance compared to baselines. 1 Warning: cases in this paper may contain offensive content. Zecheng Tang, Keyan Zhou, Juntao Li 0005, Yuyang Ding, Pinzheng Wang, Yan Bowen, Renjie Hua, Min Zhang 0005 |
EMNLP | 3 |
| 2024 | Are Bert Family Good Instruction Followers? A Study on Their Potential And LimitationsabstractLanguage modeling at scale has proven very effective and brought unprecedented success to natural language models. Many typical representatives, especially decoder-only models, e.g., BLOOM and LLaMA, and encoder-decoder models, e.g., Flan-T5 and AlexaTM, have exhibited incredible instruction-following capabilities while keeping strong task completion ability. These large language models can achieve superior performance in various tasks and even yield emergent capabilities, e.g., reasoning and universal generalization. Though the above two paradigms are mainstream and well explored, the potential of the BERT family, which are encoder-only based models and have ever been one of the most representative pre-trained models, also deserves attention, at least should be discussed. In this work, we adopt XML-R to explore the effectiveness of the BERT family for instruction following and zero-shot learning. We first design a simple yet effective strategy to utilize the encoder-only models for generation tasks and then conduct multi-task instruction tuning. Experimental results demonstrate that our fine-tuned model, Instruct-XMLR, outperforms Bloomz on all evaluation tasks and achieves comparable performance with mT0 on most tasks. Surprisingly, Instruct-XMLR also possesses strong task and language generalization abilities, indicating that Instruct-XMLR can also serve as a good instruction follower and zero-shot learner. Besides, Instruct-XMLR can accelerate decoding due to its non-autoregressive generation manner, achieving around 3 times speedup compared with current autoregressive large language models. Although we also witnessed several limitations through our experiments, such as the performance decline in long-generation tasks and the shortcoming of length prediction, Instruct-XMLR can still become a good member of the family of current large language models. Yisheng Xiao, Juntao Li 0005, Zechen Sun, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 2 |
| 2024 | ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001 |
NeurIPS | 7 |
| 2024 | Towards Better Chinese Spelling Check for Search Engines: A New Dataset and Strong BaselineabstractMisspellings in search engine queries may prevent search engines from returning accurate results. For Chinese mobile search engines, due to the different input methods (e.g., hand-written and T9 input methods), more types of misspellings exist, making this problem more challenging. As an essential module of search engines, Chinese Spelling Check~(CSC) models aim to detect and correct misspelled Chinese characters from user-issued queries. Despite the great value of CSC to the search engine, there is no CSC benchmark collected from real-world search engine queries. To fill this blank, we construct and release the Alipay Search Engine Query (AlipaySEQ) spelling check dataset. To the best of our knowledge, AlipaySEQ is the first Chinese Spelling Check dataset collected from the real-world scenario of Chinese mobile search engines. It consists of 15,522 high-quality human annotated and 1,175,151 automatically generated samples. To demonstrate the unique challenges of AlipaySEQ in the era of Large Language Models~(LLMs), we conduct a thorough study to analyze the difference between AlipaySEQ and existing SIGHAN benchmarks and compare the performance of various baselines, including existing task-specific methods and LLMs. We observe that all baselines fail to perform satisfactorily due to the over-correction problem. Especially, LLMs exhibit below-par performance on AlipaySEQ, which is rather surprising. Therefore, to alleviate the over-correction problem, we introduce a model-agnostic CSC Self-Refine Framework (SRF) to construct a strong baseline. Comprehensive experiments demonstrate that our proposed SRF, though more effective against existing models on both the AlipaySEQ and SIGHAN15, is still far from achieving satisfactory performance on our real-world dataset. With the newly collected real-world dataset and strong baseline, we hope more progress can be achieved on such a challenging and valuable task. Yue Wang 0039, Zilong Zheng, Zecheng Tang, Juntao Li 0005, Kunlong Chen, Jinxiong Chang, Qishen Zhang, Zhongyi Liu 0001, Min Zhang 0005 |
WSDM | 4 |
| 2024 | Randomness Regularization With Simple Consistency Training for Neural NetworksabstractRandomness is widely introduced in neural network training to simplify model optimization or avoid the over-fitting problem. Among them, dropout and its variations in different aspects (e.g., data, model structure) are prevalent in regularizing the training of deep neural networks. Though effective and performing well, the randomness introduced by these dropout-based methods causes nonnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize such randomness, namely R-Drop, which forces two output distributions sampled by each type of randomness to be consistent. Specifically, R-Drop minimizes the bidirectional KL-divergence between two output distributions produced by dropout-based randomness for each training sample. Theoretical analysis reveals that R-Drop can reduce the above inconsistency by reducing the inconsistency among the sampled sub structures and bridging the gap between the loss calculated by the full model and sub structures. Experiments on 7 widely-used deep learning tasks ( 23 datasets in total) demonstrate that R-Drop is universally effective for different types of neural networks (i.e., feed-forward, recurrent, and graph neural networks) and different learning paradigms (supervised, parameter-efficient, and semi-supervised). In particular, it achieves state-of-the-art performances with the vanilla Transformer model on WMT14 English → German translation ( 30.91 BLEU) and WMT14 English → French translation ( 43.95 BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Juntao Li 0005, Xiaobo Liang, Lijun Wu 0003, Yue Wang 0039, Tao Qin 0001, Min Zhang 0005, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Enhancing Low-Resource NLP by Consistency Training With Data and Model PerturbationsabstractNatural language processing (NLP) has recently shown significant progress in rich-resource scenarios. However, it is much less effective for low-resource scenarios due to the model easily overfitting to limited training data and generalizing poorly on testing data. In recent years, consistency training has been widely adopted and shown great promise in deep learning, but still remains unexplored in low-resource settings. In this work, we propose DM-CT, a framework that incorporates both data-level and model-level consistency training as well as advanced data augmentation techniques for low-resource scenarios. Concretely, the input data is first augmented, and the output distributions of different sub-models generated by model variance are forced to be consistent (model-level consistency). Meanwhile, the predictions of the original input and the augmented one are also constrained to be consistent (data-level consistency). Experiments on different low-resource NLP tasks, including neural machine translation (4 IWSLT14 translation tasks, multilingual translation task, and WMT16 Romanian$\to$English translation), natural language understanding tasks (GLUE benchmark), and named entity recognition (Conll2003 and WikiGold), well demonstrate the superiority of DM-CT by obtaining significant and consistent performance improvements. Xiaobo Liang, Runze Mao, Lijun Wu 0003, Juntao Li 0005, Min Zhang 0005, Qing Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | RenewNAT: Renewing Potential Translation for Non-autoregressive TransformerabstractNon-autoregressive neural machine translation (NAT) models are proposed to accelerate the inference process while maintaining relatively high performance. However, existing NAT models are difficult to achieve the desired efficiency-quality trade-off. For one thing, fully NAT models with efficient inference perform inferior to their autoregressive counterparts. For another, iterative NAT models can, though, achieve comparable performance while diminishing the advantage of speed. In this paper, we propose RenewNAT, a flexible framework with high efficiency and effectiveness, to incorporate the merits of fully and iterative NAT models. RenewNAT first generates the potential translation results and then renews them in a single pass. It can achieve significant performance improvements at the same expense as traditional NAT models (without introducing additional model parameters and decoding latency). Experimental results on various translation benchmarks (e.g., 4 WMT) show that our framework consistently improves the performance of strong fully NAT methods (e.g., GLAT and DSLP) without additional speed overhead. Pei Guo, Yisheng Xiao, Juntao Li 0005, Min Zhang 0005 |
AAAI | 3 |
| 2023 | AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelabstractTransformer-based autoregressive (AR) methods have achieved appealing performance for varied sequence-to-sequence generation tasks, e.g., neural machine translation, summarization, and code generation, but suffer from low inference efficiency. To speed up the inference stage, many non-autoregressive (NAR) strategies have been proposed in the past few years. Among them, the conditional masked language model (CMLM) is one of the most versatile frameworks, as it can support many different sequence generation scenarios and achieve very competitive performance on these tasks. In this paper, we further introduce a simple yet effective adaptive masking over masking strategy to enhance the refinement capability of the decoder and make the encoder optimization easier. Experiments on 3 different tasks (neural machine translation, summarization, and code generation) with 15 datasets in total confirm that our proposed simple method achieves significant performance improvement over the strong CMLM model. Surprisingly, our proposed model yields state-of-the-art performance on neural machine translation (34.62 BLEU on WMT16 EN to RO, 34.82 BLEU on WMT16 RO to EN, and 34.84 BLEU on IWSLT De to En) and even better performance than the AR Transformer on 7 benchmark datasets with at least 2.2x speedup. Our code is available at GitHub. Yisheng Xiao, Lijun Wu 0003, Juntao Li 0005, Tao Qin 0001, Tie-Yan Liu, Min Zhang 0005 |
AAAI | 4 |
| 2023 | Dynamic and Efficient Inference for Text Generation via BERT FamilyabstractDespite the excellent performance of Pretrained Language Models on many text generation tasks, they suffer from inefficient inference on computation and memory due to their largescale parameters and the universal autoregressive decoding paradigm.In this work, we propose a novel fine-tuning method DEER, which can make a single pre-trained model support Dynamic and Efficient infERence and achieve an adaptive trade-off between model performance and latency.In particular, our critical insight is to jointly utilize the non-autoregressive (NAR) generation and dynamic parameter pruning techniques, which can flexibly control the decoding iteration steps and model sizes according to memory and latency limitations.Besides, we also explore the effectiveness of the pre-trained MLMs (i.e., the BERT family) for text generation tasks since their bidirectional attention nature is more suitable for the NAR training objective.Extensive experiments on both monolingual and multilingual pre-trained MLMs demonstrate the effectiveness of our proposed DEER method by consistently achieving (1) higher BLEU scores than the strong autoregressive Transformer model on three neural machine translation tasks with 3 → 12 times speedup, (2) competitive performance (but with much faster inference speed) compared with the BART model on four GLGE benchmark tasks.Our code will be publicly available at GitHub 1 . Xiaobo Liang, Juntao Li 0005, Lijun Wu 0003, Ziqiang Cao, Min Zhang 0005 |
ACL (1) | 2 |
| 2023 | Open-ended Long Text Generation via Masked Language ModelingabstractPre-trained autoregressive (AR) language models such as BART and GPTs have dominated Open-ended Long Text Generation (Open-LTG).However, the AR nature will decrease the inference efficiency along with the increase of generation length, which hinder their application in Open-LTG.To improve inference efficiency, we alternatively explore the potential of the pre-trained masked language models (MLMs) along with a representative iterative non-autoregressive (NAR) decoding strategy for Open-LTG.Our preliminary study shows that pre-trained MLMs can merely generate short text and will collapse for long text modeling.To enhance the long text generation capability of MLMs, we introduce two simple yet effective strategies for the iterative NAR model: dynamic sliding window attention (DSWA) and linear temperature decay (LTD).It can alleviate long-distance collapse problems and achieve longer text generation with a flexible trade-off between performance and inference speedup.Experiments on the storytelling and multi-paragraph opinionated article writing tasks show that pre-trained MLMs can achieve more than 3 × → 13 × speedup with better performance than strong AR models.Our code is available at GitHub * . Xiaobo Liang, Zecheng Tang, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 3 |
| 2023 | CORE: Cooperative Training of Retriever-Reranker for Effective Dialogue Response SelectionabstractChongyang Tao, Jiazhan Feng, Tao Shen, Chang Liu, Juntao Li, Xiubo Geng, Daxin Jiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chongyang Tao, Jiazhan Feng, Tao Shen 0001, Chang Liu 0076, Juntao Li 0005, Xiubo Geng, Daxin Jiang |
ACL (1) | 5 |
| 2023 | Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringabstractAlthough pre-trained language models (PLM) have achieved great success in question answering (QA), their robustness is still insufficient to support their practical applications, especially in the face of distribution shifts.Recently, testtime adaptation (TTA) has shown great potential for solving this problem, which adapts the model to fit the test samples at test time.However, TTA sometimes causes model collapse, making almost all the model outputs incorrect, which has raised concerns about its stability and reliability.In this paper, we delve into why TTA causes model collapse and find that the imbalanced label distribution inherent in QA is the reason for it.To address this problem, we propose Anti-Collapse Fast test-time adaptation (Anti-CF), which utilizes the source model's output to regularize the update of the adapted model during test time.We further design an efficient side block to reduce its inference time.Extensive experiments on various distribution shift scenarios and pre-trained language models (e.g., XLM-RoBERTa, BLOOM) demonstrate that our method can achieve comparable or better results than previous TTA methods at a speed close to vanilla forward propagation, which is 1.8× to 4.4× speedup compared to previous TTA methods.Our code is available at https://github.com/yisunlp/Anti-CF. Yi Su 0006, Yixin Ji, Juntao Li 0005, Hai Ye, Min Zhang 0005 |
EMNLP | 3 |
| 2023 | INFORM : Information eNtropy based multi-step reasoning FOR large language ModelsabstractLarge language models (LLMs) have demonstrated exceptional performance in reasoning tasks with dedicated Chain-of-Thought (CoT) prompts.Further enhancing CoT prompts with exquisite exemplars can significantly improve reasoning performance.However, the effectiveness of CoT prompts may fluctuate dramatically with different choices of in-context examples.Additionally, manual construction of rationale steps can be time-consuming, presenting challenges for the widespread adoption of CoT prompting.In this work, we propose a novel approach by introducing information entropy (IE) as a criteria on for CoT prompt selection.We extend this criterion to the CoT generation and inference stages, automatically generating CoT prompts with higher information entropy scores and adaptively determining the number of samples.These three stages together form our proposed information entropy based multi-step reasoning for large language models, named INFORM.Our experiments across seven reasoning benchmarks utilizing two language models(GPT-3.5-Turboand text-davinci-003) demonstrate the superiority of INFORM both in performance and efficiency. 1 Chuyue Zhou, Wangjie You, Juntao Li 0005, Kehai Chen, Min Zhang 0005 |
EMNLP | 3 |
| 2023 | CT4Rec: Simple yet Effective Consistency Training for Sequential RecommendationabstractSequential recommendation methods are increasingly important in cutting-edge recommender systems. Through leveraging historical records, the systems can capture user interests and perform recommendations accordingly. State-of-the-art sequential recommendation models proposed very recently combine contrastive learning techniques for obtaining high-quality user representations. Though effective and performing well, the models based on contrastive learning require careful selection of data augmentation methods and pretext tasks, efficient negative sampling strategies, and massive hyper-parameters validation. In this paper, we propose an ultra-simple alternative for obtaining better user representations and improving sequential recommendation performance. Specifically, we present a simple yet effective Consistency T braining method for sequential Recommendation (CT4Rec) in which only two extra training objectives are utilized without any structural modifications and data augmentation. Experiments on three benchmark datasets and one large newly crawled industrial corpus demonstrate that our proposed method outperforms SOTA models by a large margin and with much less training time than these based on contrastive learning. Online evaluation on real-world content recommendation system also achieves 2.717% improvement on the click-through rate and 3.679% increase on the average click number per capita. Further exploration reveals that such a simple method has great potential for CTR prediction. Our code is available at https://github.com/ct4rec/CT4Rec.git. Xiaoyang Liu 0012, Rongqin Zheng, Xiaobo Liang, Juntao Li 0005, Lijun Wu 0003, Min Zhang 0005, Leyu Lin |
KDD | 6 |
| 2023 | Beyond Hard Samples: Robust and Effective Grammatical Error Correction with Cycle Self-Augmenting
Kaiqi Feng, Zecheng Tang, Juntao Li 0005, Min Zhang 0005 |
NLPCC (2) | 3 |
| 2023 | Are the BERT family zero-shot learners? A study on their potential and limitations
Yue Wang 0039, Lijun Wu 0003, Juntao Li 0005, Xiaobo Liang, Min Zhang 0005 |
Artif. Intell. | 3 |
| 2023 | A Survey on Non-Autoregressive Generation for Neural Machine Translation and BeyondabstractNon-autoregressive (NAR) generation, which is first proposed in neural machine translation (NMT) to speed up inference, has attracted much attention in both machine learning and natural language processing communities. While NAR generation can significantly accelerate inference speed for machine translation, the speedup comes at the cost of sacrificed translation accuracy compared to its counterpart, autoregressive (AR) generation. In recent years, many new models and algorithms have been designed/proposed to bridge the accuracy gap between NAR generation and AR generation. In this paper, we conduct a systematic survey with comparisons and discussions of various non-autoregressive translation (NAT) models from different aspects. Specifically, we categorize the efforts of NAT into several groups, including data manipulation, modeling methods, training criterion, decoding algorithms, and the benefit from pre-trained models. Furthermore, we briefly review other applications of NAR models beyond machine translation, such as grammatical error correction, text summarization, text style transfer, dialogue, semantic parsing, automatic speech recognition, and so on. In addition, we also discuss potential directions for future exploration, including releasing the dependency of KD, reasonable training objectives, pre-training for NAR, and wider applications, etc. We hope this survey can help researchers capture the latest progress in NAR generation, inspire the design of advanced NAR models and algorithms, and enable industry practitioners to choose appropriate solutions for their applications. Yisheng Xiao, Lijun Wu 0003, Junliang Guo, Juntao Li 0005, Min Zhang 0005, Tao Qin 0001, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | SelfMix: Robust Learning against Textual Label Noise with Self-Mixup TrainingabstractThe conventional success of textual classification relies on annotated data, and the new paradigm of pre-trained language models (PLMs) still requires a few labeled data for downstream tasks. However, in real-world applications, label noise inevitably exists in training data, damaging the effectiveness, robustness, and generalization of the models constructed on such data. Recently, remarkable achievements have been made to mitigate this dilemma in visual data, while only a few explore textual data. To fill this gap, we present SelfMix, a simple yet effective method, to handle label noise in text classification tasks. SelfMix uses the Gaussian Mixture Model to separate samples and leverages semi-supervised learning. Unlike previous works requiring multiple models, our method utilizes the dropout mechanism on a single model to reduce the confirmation bias in self-training and introduces a textual level mixup training strategy. Experimental results on three text classification benchmarks with different types of text show that the performance of our proposed method outperforms these strong baselines designed for both textual and visual data under different noise ratios and noise types. Our anonymous code is available at https://github.com/noise-learning/SelfMix. Chenchen Dai, Yuyang Ding, Juntao Li 0005, Wenliang Chen, Min Zhang 0005 |
COLING | 4 |
| 2022 | JANUS: Joint Autoregressive and Non-autoregressive Training with Auxiliary Loss for Sequence GenerationabstractTransformer-based autoregressive and nonautoregressive models have played an essential role in sequence generation tasks.The autoregressive model can obtain excellent performance, while the non-autoregressive model brings fast decoding speed for inference.In this paper, we propose JANUS, a Joint Autoregressive and Non-autoregressive training method using aUxiliary losS to enhance the model performance in both AR and NAR manner simultaneously and effectively alleviate the problem of distribution discrepancy.Further, we pre-train BART with JANUS on a large corpus with minimal cost (16 GPU days) and make the BART-JANUS capable of nonautoregressive generation, demonstrating that our approach can transfer the AR knowledge to NAR.Empirically, we show our approach and BART-JANUS can achieve significant improvement on multiple generation tasks, including machine translation and GLGE benchmarks.Our code is available at Github 1 . Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Min Zhang 0005 |
EMNLP | 3 |
| 2022 | Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic ChangeabstractRecent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., language model pre-trained on static data from past years performs worse over time on emerging data.Existing methods mainly perform continual training to mitigate such a misalignment.While effective to some extent but is far from being addressed on both the language modeling and downstream tasks.In this paper, we empirically observe that temporal generalization is closely affiliated with lexical semantic change, which is one of the essential phenomena of natural languages.Based on this observation, we propose a simple yet effective lexical-level masking strategy to post-train a converged language model.Experiments on two pre-trained language models, two different classification tasks, and four benchmark datasets demonstrate the effectiveness of our proposed method over existing temporal adaptation methods, i.e., continual training with new data.Our code is available at https: //github.com/zhaochen0110/LMLM. Zhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu 0003, Min Zhang 0005, Juntao Li 0005 |
EMNLP | 6 |
| 2022 | Image-text Retrieval: A Survey on Recent Research and DevelopmentabstractIn the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the queries are from one modality and the retrieval galleries from another modality. This paper presents a comprehensive and up-to-date survey on the ITR approaches from four perspectives. By dissecting an ITR system into two processes: feature extraction and feature alignment, we summarize the recent advance of the ITR approaches from these two perspectives. On top of this, the efficiency-focused study on the ITR system is introduced as the third perspective. To keep pace with the times, we also provide a pioneering overview of the cross-modal pre-training ITR approaches as the fourth perspective. Finally, we outline the common benchmark datasets and evaluation metric for ITR, and conduct the accuracy comparison among the representative ITR approaches. Some critical yet less studied issues are discussed at the end of the paper. Min Cao 0005, Shiping Li, Juntao Li 0005, Liqiang Nie, Min Zhang 0005 |
IJCAI | 3 |
| 2022 | Multi-Teacher Distillation With Single Model for Neural Machine TranslationabstractKnowledge distillation (KD) is an effective strategy for neural machine translation (NMT) to improve the performance of a student model. Usually, the teacher can guide the student to be better by distilling the soft label or data knowledge from the teacher itself. However, the data diversity and teacher knowledge are limited with only one teacher model. Though a natural solution is to adopt multiple randomized teacher models, one big shortcoming is that the model parameters and training costs are largely increased with the number of teacher models. In this work, we explore to mimic multiple teacher distillation from the sub-network space and permuted variants of one single teacher model. Specifically, we train a teacher by multiple sub-network extraction paradigms: sub-layer reordering, layer-drop, and dropout variants. In doing so, one teacher model can provide multiple outputs variants and causes neither additional parameters nor much extra training cost. Experiments on $8$ IWSLT datasets: (IWSLT14 En $\leftrightarrow$ De, En $\leftrightarrow$ Es, and IWSLT17 En $\leftrightarrow$ Fr, En $\leftrightarrow$ Zh) and the large WMT14 EN $\to$ DE translation tasks show that our method even achieves nearly comparable performance with multiple teacher models with different randomized parameters, both word-level, and sequence-level knowledge distillation. Our code is available at GitHub\footnote{https://github.com/dropreg/RLD} Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Tao Qin 0001, Min Zhang 0005, Tie-Yan Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Content Learning with Structure-Aware Writing: A Graph-Infused Dual Conditional Variational Autoencoder for Automatic StorytellingabstractRecent automatic storytelling methods mainly rely on keyword planning or plot skeleton generation to model long-range dependencies and create consistent narrative texts. However, these approaches generate story plans or plots sequentially, leaving the non-sequential conception and structural design processes of human writers unexplored. To mimic human writers and exploit the fine-grained, intrinsic structural information of each story, we decompose automatic story generation into sub-problems of graph construction, graph generation, and graph-infused sequence generation. Specifically, we propose a graph-infused dual conditional variational autoencoder model to capture multi-level intra-story structures (i.e., graph) by continuous variational latent variables and generate consistent stories through dual-infusion of story structure planning and content learning. Experimental results on the ROCStories dataset and the CMU Movie Summary corpus confirm that our proposed model outperforms strong baselines in both human judges and widely-used automatic metrics. Meng-Hsuan Yu, Juntao Li 0005, Zhangming Chan, Rui Yan 0001, Dongyan Zhao 0001 |
AAAI | 2 |
| 2021 | Learning to Organize a Bag of Words into Sentences with Neural Networks: An Empirical StudyabstractChongyang Tao, Shen Gao, Juntao Li, Yansong Feng, Dongyan Zhao, Rui Yan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Chongyang Tao, Shen Gao, Juntao Li 0005, Yansong Feng 0002, Dongyan Zhao 0001, Rui Yan 0001 |
NAACL-HLT | 3 |
| 2021 | R-Drop: Regularized Dropout for Neural NetworksabstractDropout is a powerful and widely used technique to regularize the training of deep neural networks. Though effective and performing well, the randomness introduced by dropout causes unnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize dropout, namely R-Drop, which forces the output distributions of different sub models generated by dropout to be consistent with each other. Specifically, for each training sample, R-Drop minimizes the bidirectional KL-divergence between the output distributions of two sub models sampled by dropout. Theoretical analysis reveals that R-Drop reduces the above inconsistency. Experiments on $\bf{5}$ widely used deep learning tasks ($\bf{18}$ datasets in total), including neural machine translation, abstractive summarization, language understanding, language modeling, and image classification, show that R-Drop is universally effective. In particular, it yields substantial improvements when applied to fine-tune large-scale pre-trained models, e.g., ViT, RoBERTa-large, and BART, and achieves state-of-the-art (SOTA) performances with the vanilla Transformer model on WMT14 English$\to$German translation ($\bf{30.91}$ BLEU) and WMT14 English$\to$French translation ($\bf{43.95}$ BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Our code is available at GitHub\footnote{\url{https://github.com/dropreg/R-Drop}}. Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Yue Wang 0039, Tao Qin 0001, Wei Chen 0034, Min Zhang 0005, Tie-Yan Liu |
NeurIPS | 3 |
| 2021 | Dialogue History Matters! Personalized Response Selection in Multi-Turn Retrieval-Based ChatbotsabstractExisting multi-turn context-response matching methods mainly concentrate on obtaining multi-level and multi-dimension representations and better interactions between context utterances and response. However, in real-place conversation scenarios, whether a response candidate is suitable not only counts on the given dialogue context but also other backgrounds, e.g., wording habits, user-specific dialogue history content. To fill the gap between these up-to-date methods and the real-world applications, we incorporate user-specific dialogue history into the response selection and propose a personalized hybrid matching network (PHMN). Our contributions are two-fold: (1) our model extracts personalized wording behaviors from user-specific dialogue history as extra matching information; (2) we perform hybrid representation learning on context-response utterances and explicitly incorporate a customized attention mechanism to extract vital information from context-response interactions so as to improve the accuracy of matching. We evaluate our model on two large datasets with user identification, i.e., personalized Ubuntu dialogue Corpus (P-Ubuntu) and personalized Weibo dataset (P-Weibo). Experimental results confirm that our method significantly outperforms several strong models by combining personalized attention, wording behaviors, and hybrid representation learning. Juntao Li 0005, Chang Liu 0076, Chongyang Tao, Zhangming Chan, Dongyan Zhao 0001, Min Zhang 0005, Rui Yan 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2020 | Cross-Lingual Low-Resource Set-to-Description Retrieval for Global E-CommerceabstractWith the prosperous of cross-border e-commerce, there is an urgent demand for designing intelligent approaches for assisting e-commerce sellers to offer local products for consumers from all over the world. In this paper, we explore a new task of cross-lingual information retrieval, i.e., cross-lingual set-to-description retrieval in cross-border e-commerce, which involves matching product attribute sets in the source language with persuasive product descriptions in the target language. We manually collect a new and high-quality paired dataset, where each pair contains an unordered product attribute set in the source language and an informative product description in the target language. As the dataset construction process is both time-consuming and costly, the new dataset only comprises of 13.5k pairs, which is a low-resource setting and can be viewed as a challenging testbed for model development and evaluation in cross-border e-commerce. To tackle this cross-lingual set-to-description retrieval task, we propose a novel cross-lingual matching network (CLMN) with the enhancement of context-dependent cross-lingual mapping upon the pre-trained monolingual BERT representations. Experimental results indicate that our proposed CLMN yields impressive results on the challenging task and the context-dependent cross-lingual mapping on BERT yields noticeable improvement over the pre-trained multi-lingual BERT model. Juntao Li 0005, Chang Liu 0076, Lidong Bing, Hongsong Li, Xiaozhong Liu 0001, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 1 |
| 2020 | A Character-Centric Neural Model for Automated Story GenerationabstractAutomated story generation is a challenging task which aims to automatically generate convincing stories composed of successive plots correlated with consistent characters. Most recent generation models are built upon advanced neural networks, e.g., variational autoencoder, generative adversarial network, convolutional sequence to sequence model. Although these models have achieved prompting results on learning linguistic patterns, very few methods consider the attributes and prior knowledge of the story genre, especially from the perspectives of explainability and consistency. To fill this gap, we propose a character-centric neural storytelling model, where a story is created encircling the given character, i.e., each part of a story is conditioned on a given character and corresponded context environment. In this way, we explicitly capture the character information and the relations between plots and characters to improve explainability and consistency. Experimental results on open dataset indicate that our model yields meaningful improvements over several strong baselines on both human and automatic evaluations. Juntao Li 0005, Meng-Hsuan Yu, Ziming Huang, Gongshen Liu, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 2 |
| 2020 | Draft and Edit: Automatic Storytelling Through Multi-Pass Hierarchical Conditional Variational AutoencoderabstractAutomatic Storytelling has consistently been a challenging area in the field of natural language processing. Despite considerable achievements have been made, the gap between automatically generated stories and human-written stories is still significant. Moreover, the limitations of existing automatic storytelling methods are obvious, e.g., the consistency of content, wording diversity. In this paper, we proposed a multi-pass hierarchical conditional variational autoencoder model to overcome the challenges and limitations in existing automatic storytelling models. While the conditional variational autoencoder (CVAE) model has been employed to generate diversified content, the hierarchical structure and multi-pass editing scheme allow the story to create more consistent content. We conduct extensive experiments on the ROCStories Dataset. The results verified the validity and effectiveness of our proposed model and yields substantial improvement over the existing state-of-the-art approaches. Meng-Hsuan Yu, Juntao Li 0005, Bo Tang 0016, Haisong Zhang, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 2 |
| 2020 | Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-TrainingabstractAdapting pre-trained language models (PrLMs) (e.g., BERT) to new domains has gained much attention recently.Instead of fine-tuning PrLMs as done in most previous work, we investigate how to adapt the features of PrLMs to new domains without fine-tuning.We explore unsupervised domain adaptation (UDA) in this paper.With the features from PrLMs, we adapt the models trained with labeled data from the source domain to the unlabeled target domain.Self-training is widely used for UDA, and it predicts pseudo labels on the target domain data for training.However, the predicted pseudo labels inevitably include noise, which will negatively affect training a robust model.To improve the robustness of self-training, in this paper we present class-aware feature self-distillation (CFd) to learn discriminative features from PrLMs, in which PrLM features are self-distilled into a feature adaptation module and the features from the same class are more tightly clustered.We further extend CFd to a cross-language setting, in which language discrepancy is studied.Experiments on two monolingual and multilingual Amazon review datasets show that CFd can consistently improve the performance of self-training in cross-domain and cross-language settings. Hai Ye, Ruidan He, Juntao Li 0005, Hwee Tou Ng, Lidong Bing |
EMNLP (1) | 4 |
| 2020 | Unsupervised Domain Adaptation of a Pretrained Cross-Lingual Language ModelabstractRecent research indicates that pretraining cross-lingual language models on large-scale unlabeled texts yields significant performance improvements over various cross-lingual and low-resource tasks. Through training on one hundred languages and terabytes of texts, cross-lingual language models have proven to be effective in leveraging high-resource languages to enhance low-resource language processing and outperform monolingual models. In this paper, we further investigate the cross-lingual and cross-domain (CLCD) setting when a pretrained cross-lingual language model needs to adapt to new domains. Specifically, we propose a novel unsupervised feature decomposition method that can automatically extract domain-specific features and domain-invariant features from the entangled pretrained cross-lingual representations, given unlabeled raw texts in the source language. Our proposed model leverages mutual information estimation to decompose the representations computed by a cross-lingual model into domain-invariant and domain-specific parts. Experimental results show that our proposed method achieves significant performance improvements over the state-of-the-art pretrained cross-lingual language model in the CLCD setting. Juntao Li 0005, Ruidan He, Hai Ye, Hwee Tou Ng, Lidong Bing, Rui Yan 0001 |
IJCAI | 1 |
| 2019 | Learning to Write Stories with Thematic Consistency and Wording NoveltyabstractAutomatic story generation is a challenging task, which involves automatically comprising a sequence of sentences or words with a consistent topic and novel wordings. Although many attention has been paid to this task and prompting progress has been made, there still exists a noticeable gap between generated stories and those created by humans, especially in terms of thematic consistency and wording novelty. To fill this gap, we propose a cache-augmented conditional variational autoencoder for story generation, where the cache module allows to improve thematic consistency while the conditional variational autoencoder part is used for generating stories with less common words by using a continuous latent variable. For combing the cache module and the autoencoder part, we further introduce an effective gate mechanism. Experimental results on ROCStories and WritingPrompts indicate that our proposed model can generate stories with consistency and wording novelty, and outperforms existing models under both automatic metrics and human evaluations. Juntao Li 0005, Lidong Bing, Lisong Qiu, Dongmin Chen, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 1 |
| 2019 | Insufficient Data Can Also Rock! Learning to Converse Using Smaller Data with AugmentationabstractRecent successes of open-domain dialogue generation mainly rely on the advances of deep neural networks. The effectiveness of deep neural network models depends on the amount of training data. As it is laboursome and expensive to acquire a huge amount of data in most scenarios, how to effectively utilize existing data is the crux of this issue. In this paper, we use data augmentation techniques to improve the performance of neural dialogue models on the condition of insufficient data. Specifically, we propose a novel generative model to augment existing data, where the conditional variational autoencoder (CVAE) is employed as the generator to output more training data with diversified expressions. To improve the correlation of each augmented training pair, we design a discriminator with adversarial training to supervise the augmentation process. Moreover, we thoroughly investigate various data augmentation schemes for neural dialogue system with generative models, both GAN and CVAE. Experimental results on two open corpora, Weibo and Twitter, demonstrate the superiority of our proposed data augmentation model. Juntao Li 0005, Lisong Qiu, Bo Tang 0016, Dongmin Chen, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 1 |
| 2019 | Are Training Samples Correlated? Learning to Generate Dialogue Responses with Multiple ReferencesabstractDue to its potential applications, open-domain dialogue generation has become popular and achieved remarkable progress in recent years, but sometimes suffers from generic responses.Previous models are generally trained based on 1-to-1 mapping from an input query to its response, which actually ignores the nature of 1-to-n mapping in dialogue that there may exist multiple valid responses corresponding to the same query.In this paper, we propose to utilize the multiple references by considering the correlation of different valid responses and modeling the 1-to-n mapping with a novel two-step generation architecture.The first generation phase extracts the common features of different responses which, combined with distinctive features obtained in the second phase, can generate multiple diverse and appropriate responses.Experimental results show that our proposed model can effectively improve the quality of response and outperform existing neural dialogue models on both automatic and human evaluations. Lisong Qiu, Juntao Li 0005, Wei Bi, Dongyan Zhao 0001, Rui Yan 0001 |
ACL (1) | 2 |
| 2019 | Stick to the Facts: Learning towards a Fidelity-oriented E-Commerce Product Description GenerationabstractZhangming Chan, Xiuying Chen, Yongliang Wang, Juntao Li, Zhiqiang Zhang, Kun Gai, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhangming Chan, Xiuying Chen, Juntao Li 0005, Zhiqiang Zhang 0011, Kun Gai, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Modeling Personalization in Continuous Space for Response Generation via Augmented Wasserstein AutoencodersabstractZhangming Chan, Juntao Li, Xiaopeng Yang, Xiuying Chen, Wenpeng Hu, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhangming Chan, Juntao Li 0005, Xiaopeng Yang 0002, Xiuying Chen, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Boosting Variational Generative Model via Condition Enhancing and Lexical-Editing
Zhengwei Tao, Waiman Si, Juntao Li 0005, Dongyan Zhao 0001, Rui Yan 0001 |
PRICAI (1) | 3 |
| 2018 | Generating Classical Chinese Poems via Conditional Variational Autoencoder and Adversarial TrainingabstractIt is a challenging task to automatically compose poems with not only fluent expressions but also aesthetic wording.Although much attention has been paid to this task and promising progress is made, there exist notable gaps between automatically generated ones with those created by humans, especially on the aspects of term novelty and thematic consistency.Towards filling the gap, in this paper, we propose a conditional variational autoencoder with adversarial training for classical Chinese poem generation, where the autoencoder part generates poems with novel terms and a discriminator is applied to adversarially learn their thematic consistency with their titles.Experimental results on a large poetry corpus confirm the validity and effectiveness of our model, where its automatic and human evaluation scores outperform existing models. Juntao Li 0005, Yan Song 0003, Haisong Zhang, Dongmin Chen, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP | 1 |
| 2018 | Overview of the NLPCC 2018 Shared Task: Multi-turn Human-Computer Conversations
Juntao Li 0005, Rui Yan 0001 |
NLPCC (2) | 1 |