Lingpeng Kong

dblp:144/7656 · DBLP profile ↗
← Back
95ranked-venue papers
6as first author
82since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 91 · 4 first-author · 78 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI8
2026 Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
abstract
Xiachong Feng, Deyi Yin, Xiaocheng Feng, Yi Jiang, Libo Qin, Yangfan Ye, Lei Huang, Weitao Ma, Qiming Li, Yuxuan Gu, Bing Qin, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiachong Feng, Deyi Yin, Libo Qin 0001, Yangfan Ye, Lei Huang 0021, Weitao Ma, Yuxuan Gu 0004, Bing Qin 0001, Lingpeng Kong
ACL (1)12
2026 ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models
abstract
Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval.This gap is critical: effective assistants must automatically apply learned procedures or avoid failed actions without explicit reminders.We introduce IMPLIC-ITMEMBENCH, the first systematic benchmark evaluating implicit memory through three cognitively grounded constructs drawn from standard cognitive-science accounts of nondeclarative memory: Procedural Memory (oneshot skill acquisition after interference), Priming (theme-driven bias via paired experimental/control instances), and Classical Conditioning (Conditioned Stimulus-Unconditioned Stimulus (CS-US) associations shaping first decisions).Our 300-item suite employs a unified Learning/Priming-Interfere-Test protocol with first-attempt scoring.Evaluation of 17 models reveals severe limitations: no model exceeds 66% overall, with top performers DeepSeek-R1 (65.3%),Qwen3-32B (64.1%), and GPT-5 (63.0%) far below human baselines.Analysis uncovers dramatic asymmetries (inhibition 17.6% vs. preference 75.0%) and universal bottlenecks requiring architectural innovations beyond parameter scaling.IMPLICITMEM-BENCH reframes evaluation from "what agents recall" to "what they automatically enact" 1 .
Chonghan Qin, Xiachong Feng, Weitao Ma, Lingpeng Kong
ACL (1)5
2026 OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
abstract
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong
ACL (1)14
2025 Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention Learning
abstract
Lei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001
ACL (1)13
2025 Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
abstract
Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether LLMs can serve as reliable alternatives to humans, we examine the fine-grained alignment between LLM evaluators and human annotators, particularly in understanding the target evaluation tasks and conducting evaluations that meet diverse criteria. This paper explores both conventional tasks (e.g., story generation) and alignment tasks (e.g., math reasoning), each with different evaluation criteria. Our analysis shows that 1) LLM evaluators can generate unnecessary criteria or omit crucial criteria, resulting in a slight deviation from the experts. 2) LLM evaluators excel in general criteria, such as fluency, but face challenges with complex criteria, such as numerical reasoning. We also find that LLM-pre-drafting before human evaluation can help reduce the impact of human subjectivity and minimize annotation outliers in pure human evaluation, leading to more objective evaluation. All resources are available at https://github.com/qtli/CoEval.
Qintong Li, Leyang Cui, Lingpeng Kong, Wei Bi
COLING3
2025 VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io.
Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049
CVPR11
2025 Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning Transition
abstract
Chenxin An, Zhihui Xie, Xiaonan Li, Ming Zhong, Shansan Gong, Lei Li, Jun Zhang, Jingjing Xu, Lingpeng Kong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chenxin An, Zhihui Xie 0002, Ming Zhong 0005, Shansan Gong, Lei Li 0039, Jun Zhang 0003, Jingjing Xu 0001, Lingpeng Kong
EMNLP9
2025 UNComp: Can Matrix Entropy Uncover Sparsity? - A Compressor Design from an Uncertainty-Aware Perspective
abstract
Jing Xiong, Jianghan Shen, Fanghua Ye, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Xun Wu, Chuanyang Zheng, Zhijiang Guo, Min Yang, Lingpeng Kong, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jianghan Shen, Fanghua Ye 0001, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Chuanyang Zheng, Zhijiang Guo, Min Yang 0007, Lingpeng Kong, Ngai Wong 0001
EMNLP11
2025 QSpec: Speculative Decoding with Complementary Quantization Schemes
abstract
Quantization is widely adopted to accelerate inference and reduce memory consumption in large language models (LLMs).While activation-weight joint quantization enables efficient low-precision decoding, it suffers from substantial performance degradation on multistep reasoning tasks.We propose QSPEC, a novel quantization paradigm that decouples efficiency from quality by integrating two complementary schemes via speculative decoding: low-precision joint quantization for fast drafting and high-precision weight-only quantization for accurate verification.QSPEC reuses both weights and KV cache across stages, enabling near-zero-cost switching without retraining or auxiliary models.Compared to highprecision baselines, QSPEC achieves up to 1.64× speedup without quality degradation, and outperforms state-of-the-art speculative decoding methods by up to 1.55× in batched settings.Furthermore, QSPEC supports plug-andplay deployment and generalizes well across model scales, quantization methods, and workloads.These properties make QSPEC a practical and scalable solution for high-fidelity quantized LLM serving under memory-constrained scenarios.Our code is available at https: //github.com/hku-netexplo-lab/QSpec.
Wenhao Lu, Lingpeng Kong
EMNLP4
2025 ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
abstract
Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks.However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation.To tackle this issue, we first identify the drawbacks of existing solutions (i.e., limited multi-modal reasoning capacities, and insufficient and irrelevant visual descriptions).We then decompose visual reasoning process into two stages: proactive visual perception (i.e., eyesight) and textual reasoning (i.e., wisdom), and introduce a novel visual reasoning framework named PROREASON.This framework features decoupled vision-reasoning capabilities and multi-run proactive perception.Briefly, given a multi-modal question, PRORE-ASON iterates proactive information collection and reasoning until the answer can be concluded with necessary and sufficient visual descriptions.Notably, the disassociation of capabilities allows seamless integration of existing large language models (LLMs) to compensate for the reasoning deficits of LVLMs.Our extensive experiments demonstrate that PROREASON outperforms existing multi-step reasoning frameworks on various benchmarks for both open-source and closed-source models, with the average performance gain reaching 13.2%.Besides, the integration of LLMs allows PROREASON to produce high-quality visual reasoning data, which empowers PRORE-ASON-distilled models (i.e., ProReason-VL and ProReason-Q3) to achieve superior performance in downstream tasks.Our insights into existing solutions and the decoupled perspective for feasible integration of LLMs illuminate future research on visual reasoning techniques, especially LLM-assisted ones.
Jingqi Zhou, Jingwei Dong, Lei Li 0039, Jiahui Gao 0002, Jiyue Jiang, Lingpeng Kong
EMNLP8
2025 Jailbreaking as a Reward Misspecification Problem
abstract
The widespread adoption of large language models (LLMs) has raised concerns about their safety and reliability, particularly regarding their vulnerability to adversarial attacks. In this paper, we propose a new perspective that attributes this vulnerability to reward misspecification during the alignment process. This misspecification occurs when the reward function fails to accurately capture the intended behavior, leading to misaligned model outputs. We introduce a metric ReGap to quantify the extent of reward misspecification and demonstrate its effectiveness and robustness in detecting harmful backdoor prompts. Building upon these insights, we present ReMiss, a system for automated red teaming that generates adversarial prompts in a reward-misspecified space. ReMiss achieves state-of-the-art attack success rates on the AdvBench benchmark against various target aligned LLMs while preserving the human readability of the generated prompts. Furthermore, these attacks on open-source models demonstrate high transferability to closed-source models like GPT-4o and out-of-distribution tasks from HarmBench. Detailed analysis highlights the unique advantages of the proposed reward misspecification objective compared to previous methods, offering new insights for improving LLM safety and robustness.
Zhihui Xie 0002, Jiahui Gao 0002, Lei Li 0039, Zhenguo Li, Qi Liu 0049, Lingpeng Kong
ICLR6
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR8
2025 Why Does the Effective Context Length of LLMs Fall Short?
abstract
Advancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their training lengths. In this work, we attribute this limitation to the left-skewed frequency distribution of relative positions formed in LLMs pretraining and post-training stages, which impedes their ability to effectively gather distant information. To address this challenge, we introduce Shifted Rotray Position Embedding (STRING). STRING shifts well-trained positions to overwrite the original ineffective positions during inference, enhancing performance within their existing training lengths. Experimental results show that without additional training, STRING dramatically improves the performance of the latest large-scale models, such as Llama3.1 70B and Qwen2 72B, by over 10 points on popular long-context benchmarks RULER and InfiniteBench, establishing new state-of-the-art results for open-source LLMs. Compared to commercial models, Llama 3.1 70B with STRING even achieves better performance than GPT-4-128K and clearly surpasses Claude 2 and Kimi-chat.
Chenxin An, Jun Zhang 0003, Ming Zhong 0005, Lei Li 0039, Shansan Gong, Yao Luo, Jingjing Xu 0001, Lingpeng Kong
ICLR8
2025 G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
abstract
Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, current work has been largely focused on text-based mathematical problems, with limited investigation in problems involving multi-modal geometric information. Addressing this gap, we aim to enable LLMs to solve geometric problems by understanding image input. We first identify the limitations of current Multimodal Large Language Models (MLLMs) in this area: they struggle to accurately comprehend basic geometric elements and their relationships. To address these challenges, we leverage the inherent attribute of logical structure compactness in geometric figures, utilizing text-only Large Language Models (LLMs) to curate a comprehensive multimodal geometry dataset. This dataset, named Geo170k, contains more than 170K geometric image-caption and question-answer pairs. Utilizing the Geo170k dataset, we introduce G-LLaVA, a model that demonstrates exceptional performance in solving geometric problems. It significantly outperforms GPT4-V on the geometry task of MathVista benchmark with only 7B parameters.
Jiahui Gao 0002, Renjie Pi, Jiacheng Ye, Wanjun Zhong, Yufei Wang 0005, Lanqing Hong, Jianhua Han, Hang Xu 0004, Zhenguo Li, Lingpeng Kong
ICLR11
2025 Scaling Diffusion Language Models via Adaptation from Autoregressive Models
abstract
Diffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions.
Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong
ICLR12
2025 Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced Exploration
abstract
Large language models (LLMs) have significantly benefited from training on diverse, high-quality task-specific data, leading to impressive performance across a range of downstream applications. Current methods often rely on human-annotated data or predefined task templates to direct powerful LLMs in synthesizing task-relevant data for effective model training. However, this dependence on manually designed components may constrain the scope of generated data, potentially overlooking critical edge cases or novel scenarios that could challenge the model. In this paper, we present a novel approach, ReverseGen, designed to automatically generate effective training samples that expose the weaknesses of LLMs. Specifically, we introduce a dedicated proposer trained to produce queries that lead target models to generate unsatisfactory responses. These failure-inducing queries are then used to construct training data, helping to address the models' shortcomings and improve overall performance. Our approach is flexible and can be applied to models of various scales (3B, 7B, and 8B). We evaluate ReverseGen on three key applications—safety, honesty, and math—demonstrating that our generated data is both highly effective and diverse. Models fine-tuned with ReverseGen-generated data consistently outperform those trained on human-annotated or general model-generated data, offering a new perspective on data synthesis for task-specific LLM enhancement.
Qintong Li, Jiahui Gao 0002, Renjie Pi, Xueliang Zhao, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR9
2025 Non-myopic Generation of Language Models for Reasoning and Planning
abstract
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning and planning. Despite their success in various domains, such as mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to the inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By reweighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements across a wide range of tasks in math, coding, and agent-based scenarios. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines while utilizing inference compute more effectively. This study provides insights into optimizing LLM planning capabilities.
Haiteng Zhao, Junlei Zhang, Junxian He, Lingpeng Kong
ICLR5
2025 MoS: Unleashing Parameter Efficiency of Low-Rank Adaptation with Mixture of Shards
abstract
The rapid scaling of large language models necessitates more lightweight finetuning methods to reduce the explosive GPU memory overhead when numerous customized models are served simultaneously. Targeting more parameter-efficient low-rank adaptation (LoRA), parameter sharing presents a promising solution. Empirically, our research into high-level sharing principles highlights the indispensable role of differentiation in reversing the detrimental effects of pure sharing. Guided by this finding, we propose Mixture of Shards (MoS), incorporating both inter-layer and intra-layer sharing schemes, and integrating four nearly cost-free differentiation strategies, namely subset selection, pair dissociation, vector sharding, and shard privatization. Briefly, it selects a designated number of shards from global pools with a Mixture-of-Experts (MoE)-like routing mechanism before sequentially concatenating them to low-rank matrices. Hence, it retains all the advantages of LoRA while offering enhanced parameter efficiency, and effectively circumvents the drawbacks of peer parameter-sharing methods. Our empirical experiments demonstrate approximately $8\times$ parameter savings in a standard LoRA setting. The ablation study confirms the significance of each component. Our insights into parameter sharing and MoS method may illuminate future developments of more parameter-efficient finetuning methods. The code is officially available at https://github.com/Forence1999/MoS.
Pengan Chen, Jingwei Dong, Boyang Xue, Jiyue Jiang, Lingpeng Kong
ICLR7
2025 Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
abstract
Autoregressive language models, despite their impressive capabilities, struggle with complex reasoning and long-term planning tasks. We introduce discrete diffusion models as a novel solution to these challenges. Through the lens of subgoal imbalance, we demonstrate how diffusion models effectively learn difficult subgoals that elude autoregressive approaches. We propose Multi-Granularity Diffusion Modeling (MGDM), which prioritizes subgoals based on difficulty during learning. On complex tasks like Countdown, Sudoku, and Boolean Satisfiability Problems, MGDM significantly outperforms autoregressive models without using search techniques. For instance, MGDM achieves 91.5\% and 100\% accuracy on Countdown and Sudoku, respectively, compared to 45.8\% and 20.7\% for autoregressive models. Our work highlights the potential of diffusion-based approaches in advancing AI capabilities for sophisticated language understanding and problem-solving tasks. All associated codes are available at \href{https://github.com/HKUNLP/diffusion-vs-ar}{https://github.com/HKUNLP/diffusion-vs-ar}.
Jiacheng Ye, Jiahui Gao 0002, Shansan Gong, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR7
2025 Implicit Search via Discrete Diffusion: A Study on Chess
abstract
In the post-AlphaGo era, there has been a renewed interest in search techniques such as Monte Carlo Tree Search (MCTS), particularly in their application to Large Language Models (LLMs). This renewed attention is driven by the recognition that current next-token prediction models often lack the ability for long-term planning. Is it possible to instill search-like abilities within the models to enhance their planning abilities without relying on explicit search? We propose DiffuSearch , a model that does \textit{implicit search} by looking into the future world via discrete diffusion modeling. We instantiate DiffuSearch on a classical board game, Chess, where explicit search is known to be essential. Through extensive controlled experiments, we show DiffuSearch outperforms both the searchless and explicit search-enhanced policies. Specifically, DiffuSearch outperforms the one-step policy by 19.2\% and the MCTS-enhanced policy by 14\% on action accuracy. Furthermore, DiffuSearch demonstrates a notable 30\% enhancement in puzzle-solving abilities compared to explicit search-based policies, along with a significant 540 Elo increase in game-playing strength assessment. These results indicate that implicit search via discrete diffusion is a viable alternative to explicit search over a one-step policy. All codes are publicly available at \href{https://github.com/HKUNLP/DiffuSearch}{https://github.com/HKUNLP/DiffuSearch}.
Jiacheng Ye, Jiahui Gao 0002, Zhiyong Wu 0003, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR7
2025 Teaching Language Models to Critique via Reinforcement Learning
abstract
Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide *accurate judgments* and *actionable suggestions*. In this work, we study LLM critics for code generation and propose $\texttt{CTRL}$, a framework for $\texttt{C}$ritic $\texttt{T}$raining via $\texttt{R}$einforcement $\texttt{L}$earning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with $\texttt{CTRL}$ significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models. Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1\% relative improvements across challenging code generation benchmarks.
Zhihui Xie 0002, Liyu Chen, Weichao Mao, Jingjing Xu 0001, Lingpeng Kong
ICML6
2025 ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
abstract
Extrapolating ultra-long contexts (text length $>$128K) remains a major challenge for large language models (LLMs), as most training-free extrapolation methods are not only severely limited by memory bottlenecks, but also suffer from the attention sink, which restricts their scalability and effectiveness in practice. In this work, we propose ParallelComp, a parallel long-context compression method that effectively overcomes the memory bottleneck, enabling 8B-parameter LLMs to extrapolate from 8K to 128K tokens on a single A100 80GB GPU in a training-free setting. ParallelComp splits the input into chunks, dynamically evicting redundant chunks and irrelevant tokens, supported by a parallel KV cache eviction mechanism. Importantly, we present a systematic theoretical and empirical analysis of attention biases in parallel attention—including the attention sink, recency bias, and middle bias—and reveal that these biases exhibit distinctive patterns under ultra-long context settings. We further design a KV cache eviction technique to mitigate this phenomenon. Experimental results show that ParallelComp enables an 8B model (trained on 8K context) to achieve 91.17% of GPT-4’s performance under ultra-long contexts, outperforming closed-source models such as Claude-2 and Kimi-Chat. We achieve a 1.76x improvement in chunk throughput, thereby achieving a 23.50x acceleration in the prefill stage with negligible performance loss and pave the way for scalable and robust ultra-long contexts extrapolation in LLMs. We release the code at https://github.com/menik1126/ParallelComp.
Jianghan Shen, Chuanyang Zheng, Zhongwei Wan, Chiwun Yang, Fanghua Ye 0001, Hongxia Yang, Lingpeng Kong, Ngai Wong 0001
ICML9
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia11
2025 FactTrack: Time-Aware World State Tracking in Story Outlines
abstract
Zhiheng Lyu, Kevin Yang, Lingpeng Kong, Dan Klein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhiheng Lyu, Kevin Yang, Lingpeng Kong
NAACL (Long Papers)3
2025 ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
abstract
Xijia Tao, Shuai Zhong, Lei Li, Qi Liu, Lingpeng Kong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xijia Tao, Shuai Zhong, Lei Li 0039, Qi Liu 0049, Lingpeng Kong
NAACL (Long Papers)5
2025 TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning
abstract
Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases and low-variation prompts, resulting in limited diversity and biased distribution with the increase of data scales. To tackle this challenge, we introduce TreeSynth, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness, before synthesizing samples within each atomic subspace. This globally divide-and-synthesize method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the re-balancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently validates the superior data diversity, model performance, and robust scalability of TreeSynth compared to both human-crafted datasets and peer data synthesis methods, with the average performance gain reaching 10%. Besides, the consistent improvements of TreeSynth-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth.
Pengan Chen, Jingqi Zhou, Qintong Li, Jingwei Dong, Jiahui Gao 0002, Boyang Xue, Jiyue Jiang, Lingpeng Kong
NeurIPS9
2025 DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
abstract
In modern sequential decision-making systems, the construction of an optimal candidate action space is critical to efficient inference. However, existing approaches either rely on manually defined action spaces that lack scalability or utilize unstructured spaces that render exhaustive search computationally prohibitive. In this paper, we propose a novel framework named \textsc{DynaAct} for automatically constructing a compact action space to enhance sequential reasoning in complex problem-solving scenarios. Our method first estimates a proxy for the complete action space by extracting general sketches observed in a corpus covering diverse complex reasoning problems using large language models. We then formulate a submodular function that jointly evaluates candidate actions based on their utility to the current state and their diversity, and employ a greedy algorithm to select an optimal candidate set. Extensive experiments on six diverse standard benchmarks demonstrate that our approach significantly improves overall performance, while maintaining efficient inference without introducing substantial latency. The implementation is available at \url{https://github.com/zhaoxlpku/DynaAct}.
Xueliang Zhao, Wei Wu 0014, Jian Guan 0002, Qintong Li, Lingpeng Kong
NeurIPS5
2025 LOGOWheat: deep learning-based prediction of regulatory effects for noncoding variants in wheats
abstract
Identifying the regulatory effects of noncoding variants presents a significant challenge. Recently, the accumulation of epigenomic profiling data in wheat has provided an opportunity to model the functional impacts of these variants. In this study, we introduce Language of Genome for Wheat (LOGOWheat), a deep learning-based tool designed to predict the regulatory effects of noncoding variants in wheat. LOGOWheat initially employs a self-attention-based, contextualized pretrained language model to acquire bidirectional representations of the unlabeled wheat reference genome. Epigenomic profiling data are also collected and utilized to fine-tune the model, enabling it to discern the regulatory code inherent in genomic sequences. The test results suggest that LOGOWheat is highly effective in predicting multiple chromatin features, achieving an average area under the receiver operating characteristic (AUROC) of 0.8531 and an average area under the precision-recall curve (AUPRC) of 0.7633. Two case studies illustrate and demonstrate the main functions provided by LOGOWheat: assigning scores and prioritizing causal variants within a given variant set and constructing a saturated mutagenesis map in silico to discover high-impact sites or functional motifs in a given sequence. Finally, we propose the concept of extracting potential functional variations from the wheat population by integrating evolutionary conservation information. LOGOWheat is available at http://logowheat.cn/.
Lingpeng Kong, Kun Zhu 0019
Briefings Bioinform.1
2025 Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong
Int. J. Comput. Vis.9
2024 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
abstract
Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes.However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains.To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension.ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains.Drawing from ArXivCap, we introduce ArXivQA, a questionanswering dataset generated by prompting GPT-4V based on scientific figures.ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs.Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domainspecific training yields substantial performance gains.Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.
Lei Li 0039, Yuqi Wang 0003, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu 0049
ACL (1)6
2024 L-Eval: Instituting Standardized Evaluation for Long Context Language Models
abstract
Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chenxin An, Shansan Gong, Ming Zhong 0005, Xingjian Zhao, Mukai Li, Jun Zhang 0003, Lingpeng Kong, Xipeng Qiu
ACL (1)7
2024 GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
abstract
Large language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks.However, there are increasing debates regarding whether these models truly understand and apply mathematical knowledge or merely rely on shortcuts for mathematical reasoning.One essential and frequently occurring evidence is that when the math questions are slightly changed, LLMs can behave incorrectly.This motivates us to evaluate the robustness of LLMs' math reasoning capability by testing a wide range of question variations.We introduce the adversarial grade school math (GSM-PLUS) dataset, an extension of GSM8K augmented with various mathematical perturbations.Our experiments on 25 LLMs and 4 prompting techniques show that while LLMs exhibit different levels of math reasoning abilities, their performances are far from robust.In particular, even for problems that have been solved in GSM8K, LLMs can make mistakes when new statements are added or the question targets are altered.We also explore whether more robust performance can be achieved by composing existing prompting methods, in which we try an iterative method that generates and verifies each intermediate thought based on its reasoning goal and calculation result.
Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong, Wei Bi
ACL (1)4
2024 Large Language Models are not Fair Evaluators
abstract
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui
ACL (1)8
2024 PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
abstract
Sheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang, Liheng Chen, Lingpeng Kong, Chuan Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Boyang Xue, Jiacheng Ye, Jiyue Jiang, Lingpeng Kong
ACL (1)6
2024 SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving
abstract
Large Language Models (LLMs) have driven substantial progress in artificial intelligence in recent years, exhibiting impressive capabilities across a wide range of tasks, including mathematical problem-solving.Inspired by the success of subgoal-based methods, we propose a novel framework called SEquential subGoal Optimization (SEGO) to enhance LLMs' ability to solve mathematical problems.By establishing a connection between the subgoal breakdown process and the probability of solving problems, SEGO aims to identify better subgoals with theoretical guarantees.Addressing the challenge of identifying suitable subgoals in a large solution space, our framework generates problem-specific subgoals and adjusts them according to carefully designed criteria.Incorporating these optimized subgoals into the policy model training leads to significant improvements in problem-solving performance.We validate SEGO's efficacy through experiments on two benchmarks, GSM8K and MATH, where our approach outperforms existing methods, highlighting the potential of SEGO in AI-driven mathematical problemsolving. * This work
Xueliang Zhao, Xinting Huang, Wei Bi, Lingpeng Kong
ACL (1)4
2024 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
abstract
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049
EMNLP9
2024 Retrieved Sequence Augmentation for Protein Representation Learning
abstract
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Young Lu, Qi Liu, Sheng Wang, Lingpeng Kong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiteng Zhao, Jiayi Xin, Qintong Li, Zhi-Hong Deng 0001, Qi Liu 0049, Lingpeng Kong
EMNLP11
2024 Lemur: Harmonizing Natural Language and Code for Language Agents
abstract
We introduce Lemur and Lemur-Chat, openly accessible language models optimized for both natural language and coding capabilities to serve as the backbone of versatile language agents. The evolution from language chat models to functional language agents demands that models not only master human interaction, reasoning, and planning but also ensure grounding in the relevant environments. This calls for a harmonious blend of language and coding capabilities in the models. Lemur and Lemur-Chat are proposed to address this necessity, demonstrating balanced proficiencies in both domains, unlike existing open-source models that tend to specialize in either. Through meticulous pretraining using a code-intensive corpus and instruction fine-tuning on text and code data, our models achieve state-of-the-art averaged performance across diverse text and coding benchmarks. Comprehensive experiments demonstrate Lemur’s superiority over existing open-source models and its proficiency across various agent tasks involving human communication, tool usage, and interaction under fully- and partially- observable environments. The harmonization between natural and programming languages enables Lemur-Chat to significantly narrow the gap with proprietary models on agent abilities, providing key insights into developing advanced open-source agents adept at reasoning, planning, and operating seamlessly across environments. Our model and code have been open-sourced at https://github.com/OpenLemur/Lemur.
Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu 0033, Binyuan Hui, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu 0009
ICLR13
2024 Training-Free Long-Context Scaling of Large Language Models
abstract
The ability of Large Language Models (LLMs) to process and generate coherent text is markedly weakened when the number of input tokens exceeds their pretraining length. Given the expensive overhead of finetuning large-scale models with longer sequences, we propose a training-free approach named Dual Chunk Attention (DCA), which enables Llama2 70B to support context windows of up to 100k tokens. By decomposing the attention computation for long sequences into chunk-based modules, DCA manages to effectively capture the relative positional information of tokens within the same chunk (Intra-Chunk) and across distinct chunks (Inter-Chunk), as well as integrates seamlessly with Flash Attention. In addition to its impressive extrapolation capability, DCA achieves performance on practical long-context tasks that is comparable to or even better than that of models built through continual training. All code and data used in this work are released at https://github.com/HKUNLP/ChunkLlama.
Chenxin An, Fei Huang 0005, Jun Zhang 0003, Shansan Gong, Xipeng Qiu, Chang Zhou 0005, Lingpeng Kong
ICML7
2024 Subgoal-based Demonstration Learning for Formal Theorem Proving
abstract
Large language models (LLMs) present a promising pathway for advancing the domain of formal theorem proving. In this paper, we aim to improve the performance of LLMs in formal theorem proving by thoroughly examining the structure and organization of demonstrative in-context examples. We introduce a subgoal-based demonstration learning framework, specifically designed to enhance the efficiency of proof search in LLMs. First, drawing upon the insights of subgoal learning from reinforcement learning and robotics, we propose the construction of distinct subgoals for each demonstration example and refine these subgoals in accordance with the pertinent theories of subgoal learning. Second, we build upon recent advances in diffusion models to predict the optimal organization, simultaneously addressing two intricate issues that persist within the domain of demonstration organization: subset selection and order determination. Our integration of subgoal-based learning has notably increased proof accuracy from 38.9% to 44.1% on the miniF2F benchmark. Furthermore, the adoption of diffusion models for demonstration organization can lead to an additional enhancement in accuracy to 45.5%, or a $5\times$ improvement in sampling efficiency compared to previously established methods.
Xueliang Zhao, Wenda Li 0001, Lingpeng Kong
ICML3
2024 Self-Infilling Code Generation
abstract
In this work, we introduce self-infilling code generation, a general framework that incorporates infilling operations into auto-regressive decoding. Our approach capitalizes on the observation that recent infilling-capable code language models can perform self-infilling: whereas conventional infilling is designed to fill in the middle based on a predefined prefix and suffix, self-infilling sequentially generates both such surrounding context and the infilled content. We utilize self-infilling to introduce novel interruption and looping mechanisms in conventional decoding, evolving it into a non-monotonic process. Interruptions allow for postponing the generation of specific code until a definitive suffix is established, enhancing control during decoding. Meanwhile, the looping mechanism, which leverages the complementary nature of self-infilling and left-to-right decoding, can iteratively update and synchronize each piece of generation cyclically. Extensive experiments across a variety of code generation benchmarks demonstrate that decoding with self-infilling not only improves the output quality but also regularizes the overall generation, which effectively mitigates potential degeneration and scaffolds code to be more consistent with intended functionality.
Zhi Zhang 0005, Hongxia Yang, Lingpeng Kong
ICML5
2024 AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
abstract
Evaluating large language models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a significant step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
Junlei Zhang, Cheng Yang 0007, Yujiu Yang 0001, Yaohui Jin, Zhen-Zhong Lan, Lingpeng Kong, Junxian He
NeurIPS8
2024 Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models
abstract
Recently, diffusion models have garnered significant interest in the field of text processing due to their many potential advantages compared to conventional autoregressive models. In this work, we propose Diffusion-of-Thought (DoT), a novel approach that integrates diffusion models with Chain-of-Thought, a well-established technique for improving the reasoning ability of autoregressive language models. In contrast to autoregressive language models that make decisions in a left-to-right, token-by-token manner, DoT allows reasoning steps to diffuse over time through a diffusion language model and offers greater flexibility in trading-off computation for reasoning performance. Our experimental results demonstrate the effectiveness of DoT in multi-digit multiplication, boolean logic, and grade school math problems. In addition to that, DoT showcases promising self-correction abilities and benefits from existing reasoning-enhancing techniques like self-consistency decoding. Our findings contribute to the understanding and development of reasoning with diffusion language models.
Jiacheng Ye, Shansan Gong, Jiahui Gao 0002, Xin Jiang 0002, Zhenguo Li, Wei Bi, Lingpeng Kong
NeurIPS11
2023 Unsupervised Explanation Generation via Correct Instantiations
abstract
While large pre-trained language models (PLM) have shown their great skills at solving discriminative tasks, a significant gap remains when compared with humans for explanation-related tasks. Among them, explaining the reason why a statement is wrong (e.g., against commonsense) is incredibly challenging. The major difficulty is finding the conflict point, where the statement contradicts our real world. This paper proposes Neon, a two-phrase, unsupervised explanation generation framework. Neon first generates corrected instantiations of the statement (phase I), then uses them to prompt large PLMs to find the conflict point and complete the explanation (phase II). We conduct extensive experiments on two standard explanation benchmarks, i.e., ComVE and e-SNLI. According to both automatic and human evaluations, Neon outperforms baselines, even for those with human-annotated instantiations. In addition to explaining a negative prediction, we further demonstrate that Neon remains effective when generalizing to different scenarios. The resources of Neon are available at: https://github.com/Shark-NLP/Neon.
Sijie Cheng, Zhiyong Wu 0003, Jiangjie Chen, Lingpeng Kong
AAAI6
2023 Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering
abstract
Despite the impressive few-shot performance of in-context learning (ICL), it remains a common practice to randomly select examples to serve as the context.In this paper, we advocate self-adaptive in-context learning, a new principle for ICL, in which the self-adaption mechanism is introduced to help each input find an in-context example organization (i.e., selection and permutation) that can derive the correct output, thus maximizing performance.To validate the effectiveness of self-adaptive ICL, we propose a general select-then-rank framework and a set of novel selection and ranking algorithms.Upon extensive evaluation on eight different NLP datasets, our self-adaptive ICL method achieves a 40% relative improvement over the common practice setting.Further analysis reveals the great potential of selfadaptive ICL as a promising method to close the gap between ICL and finetuning.Our code will be released to facilitate future research.
Zhiyong Wu 0003, Yaoxiang Wang, Jiacheng Ye, Lingpeng Kong
ACL (1)4
2023 A Cognitive Stimulation Dialogue System with Multi-source Knowledge Fusion for Elders with Cognitive Impairment
abstract
When communicating with elders with cognitive impairment, cognitive stimulation (CS) help to maintain the cognitive health of elders.Data sparsity is the main challenge in building CS-based dialogue systems, particularly in the Chinese language.To fill this gap, we construct a Chinese CS conversation (CSConv) dataset, which contains about 2.6K groups of dialogues with therapy principles and emotional support strategy labels.Making chit chat while providing emotional support is overlooked by the majority of existing cognitive dialogue systems.In this paper, we propose a multi-source knowledge fusion method for CS dialogue (CSD), to generate open-ended responses guided by the therapy principle and emotional support strategy.We first use a progressive mask method based on external knowledge to learn encoders as effective classifiers, which is the prerequisite to predict the therapy principle and emotional support strategy of the target response.Then a decoder interacts with the perceived therapy principle and emotional support strategy to generate responses.Extensive experiments conducted on the CSConv dataset demonstrate the effectiveness of the proposed method, while there is still a large space for improvement compared to human performance 1 .
Jiyue Jiang, Qintong Li, Lingpeng Kong
ACL (1)4
2023 INK: Injecting kNN Knowledge in Nearest Neighbor Machine Translation
abstract
Neural machine translation has achieved promising results on many translation tasks.However, previous studies have shown that neural models induce a non-smooth representation space, which harms its generalization results.Recently, kNN-MT has provided an effective paradigm to smooth the prediction based on neighbor representations during inference.Despite promising results, kNN-MT usually requires large inference overhead.We propose an effective training framework INK to directly smooth the representation space via adjusting representations of kNN neighbors with a small number of new parameters.The new parameters are then used to refresh the whole representation datastore to get new kNN knowledge asynchronously.This loop keeps running until convergence.Experiments on four benchmark datasets show that INK achieves average gains of 1.99 COMET and 1.0 BLEU, outperforming the state-of-the-art kNN-MT system with 0.02× memory space and 1.9× inference speedup 1 .
Jingjing Xu 0001, Shujian Huang, Lingpeng Kong, Jiajun Chen 0001
ACL (1)4
2023 Fine-grained Audible Video Description
abstract
We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds in videos. Existing visual-language modeling tasks often concentrate on visual cues in videos while undervaluing the language and audio modalities. On the other hand, FAVD requires not only audio-visual-language modeling skills but also paragraph-level language generation abilities. We construct the first fine-grained audible video description benchmark (FAVDBench) to facilitate this research. For each video clip, we first provide a one-sentence summary of the video, i.e., the caption, followed by 4–6 sentences describing the visual details and 1–2 audio-related descriptions at the end. The descriptions are provided in both English and Chinese. We create two new metrics for this task: an EntityScore to gauge the completeness of entities in the visual descriptions, and an AudioScore to assess the audio descriptions. As a preliminary approach to this task, we propose an audio-visual-language transformer that extends existing video captioning model with an additional audio branch. We combine the masked language modeling and auto-regressive language modeling losses to optimize our model so that it can produce paragraph-level descriptions. We illustrate the efficiency of our model in audio-visual-language modeling by evaluating it against the proposed benchmark using both conventional captioning metrics and our proposed metrics. We further put our benchmark to the test in video generation models, demonstrating that employing fine-grained video descriptions can create more intricate videos than using captions. Code and dataset are available at https://github.com/OpenNLPLab/FAVDBench. Our online benchmark is available at www.avlbench.opennlplab.cn.
Xuyang Shen, Dong Li 0033, Jinxing Zhou, Zhen Qin 0003, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang 0001, Yu Qiao 0001, Yiran Zhong
CVPR9
2023 Can Language Models Understand Physical Concepts?
abstract
Language models (LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.However, it is unclear whether LMs can understand physical concepts in the human world.To investigate this, we design a benchmark VEC that covers the tasks of (i) Visual concepts, such as the shape and material of objects, and (ii) Embodied Concepts, learned from the interaction with the world such as the temperature of objects.Our zero (few)-shot prompting results show that the understanding of certain visual concepts emerges as scaling up LMs, but there are still basic concepts to which the scaling law does not apply.For example, OPT-175B performs close to humans with a zero-shot accuracy of 85% on the material concept, yet behaves like random guessing on the mass concept.Instead, vision-augmented LMs such as CLIP and BLIP achieve a human-level understanding of embodied concepts.Analysis indicates that the rich semantics in visual representation can serve as a valuable source of embodied knowledge.Inspired by this, we propose a distillation method to transfer embodied knowledge from VLMs to LMs, achieving performance gain comparable with that by scaling up parameters of LMs 134×. 1 o 1 : This is a photo of the water.o 2 : This is a photo of a frying oil.Attribute: This is a photo of a cold object.
Lei Li 0039, Jingjing Xu 0001, Qingxiu Dong, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
EMNLP6
2023 DetGPT: Detect What You Need via Reasoning
abstract
Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, Tong Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Renjie Pi, Jiahui Gao 0002, Shizhe Diao, Rui Pan 0002, Hanze Dong, Lewei Yao, Jianhua Han, Hang Xu 0004, Lingpeng Kong, Tong Zhang 0001
EMNLP10
2023 Generating Data for Symbolic Language with Large Language Models
abstract
While large language models (LLMs) bring not only performance but also complexity, recent work has started to turn LLMs into data generators rather than task inferencers, where another affordable task model is trained for efficient deployment and inference.However, such an approach has primarily been applied to natural language tasks, and has not yet been explored for symbolic language tasks with complex structured outputs (e.g., semantic parsing and code generation).In this paper, we propose SYMGEN which utilizes LLMs for generating various annotationexpensive symbolic language data.SYMGEN consists of an informative prompt to steer generation and an agreement-based verifier to improve data correctness.We conduct extensive experiments on six symbolic language tasks across various settings.Compared with the LLMs, we demonstrate the 1%-sized task model can achieve comparable or better performance, largely cutting inference and deployment costs.We also show that generated data with only a few human demonstrations can be as effective as over 10 times the amount of human-annotated data when training the task model, saving a considerable amount of annotation effort.SYMGEN takes a step toward data generation for annotation-expensive complex tasks, and we release the code at https://github.com/HKUNLP/SymGen.
Jiacheng Ye, Chengzu Li, Lingpeng Kong, Tao Yu 0009
EMNLP3
2023 Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
Jiahui Gao 0002, Renjie Pi, Hang Xu 0004, Jiacheng Ye, Zhiyong Wu 0003, Xiaodan Liang, Zhenguo Li, Lingpeng Kong
ICLR10
2023 DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu 0003, Lingpeng Kong
ICLR5
2023 Toeplitz Neural Network for Sequence Modeling
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dong Li 0033, Dongxu Li 0003, Yuchao Dai, Lingpeng Kong, Yiran Zhong
ICLR8
2023 Efficient Attention via Control Variates
Lingpeng Kong
ICLR4
2023 Compositional Exemplars for In-context Learning
abstract
Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we systematically formulate in-context example selection as a subset selection problem, and optimize it in an end-to-end fashion. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through carefully-designed contrastive learning to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, phraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation and semantic parsing. Extensive experiments demonstrate the effectiveness, transferability, compositionality of CEIL, shedding new lights on in-context leaning. Our code is released at https://github.com/HKUNLP/icl-ceil.
Jiacheng Ye, Zhiyong Wu 0003, Jiangtao Feng, Tao Yu 0009, Lingpeng Kong
ICML5
2023 CAB: Comprehensive Attention Benchmarking on Long Sequence Modeling
abstract
Transformer has achieved remarkable success in language, image, and speech processing. Recently, various efficient attention architectures have been proposed to improve transformer’s efficiency while largely preserving its efficacy, especially in modeling long sequences. A widely-used benchmark to test these efficient methods’ capability on long-range modeling is Long Range Arena (LRA). However, LRA only focuses on the standard bidirectional (or noncausal) self attention, and completely ignores cross attentions and unidirectional (or causal) attentions, which are equally important to downstream applications. In this paper, we propose Comprehensive Attention Benchmark (CAB) under a fine-grained attention taxonomy with four distinguishable attention patterns, namely, noncausal self, causal self, noncausal cross, and causal cross attentions. CAB collects seven real-world tasks from different research areas to evaluate efficient attentions under the four attention patterns. Among these tasks, CAB validates efficient attentions in eight backbone networks to show their generalization across neural architectures. We conduct exhaustive experiments to benchmark the performances of nine widely-used efficient attention architectures designed with different philosophies on CAB. Extensive experimental results also shed light on the fundamental problems of efficient attentions, such as efficiency length against vanilla attention, performance consistency across attention patterns, the benefit of attention mechanisms, and interpolation/extrapolation on long-context language modeling.
Shuyang Jiang, Jiangtao Feng, Lingpeng Kong
ICML5
2023 Statistical Knowledge Assessment for Large Language Models
abstract
Given varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we study the problem of quantifying knowledge contained in an LLM regarding a given set of facts. We propose KaRR, a statistical approach to assess factual knowledge for LLMs. The main idea is to estimate the ratio of LLM generating text corresponding to the answer entity given diverse prompts of the subject and the querying relation, versus it generating by random chances. Our assessment suite contains a comprehensive set of 994,123 entities and 600 relations, with 1,395,905 text aliases. We use our method to evaluate 20 LLMs of various sizes, including LLaMA, Alpaca, OPT, etc. Experiments show that our results have a strong correlation (0.43 Kendall's $\tau$) with the results of human assessment on LLMs. Our results reveal that the knowledge in LLMs with the same backbone architecture adheres to the scaling law, while tuning on instruction-following data sometimes compromises the model's capability to generate factually correct text reliably.
Qingxiu Dong, Jingjing Xu 0001, Lingpeng Kong, Zhifang Sui, Lei Li 0005
NeurIPS3
2023 GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
abstract
Molecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs. To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv.
Haiteng Zhao, Shengchao Liu, Hannan Xu, Jie Fu 0001, Zhi-Hong Deng 0001, Lingpeng Kong, Qi Liu 0049
NeurIPS7
2023 Attentive Multi-Layer Perceptron for Non-autoregressive Generation
Shuyang Jiang, Jiangtao Feng, Lingpeng Kong
ECML/PKDD (2)5
2023 Vicinity Vision Transformer
abstract
Vision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches.
Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong
IEEE Trans. Pattern Anal. Mach. Intell.9
2022 Lexical Knowledge Internalization for Neural Dialog Generation
abstract
We propose knowledge internalization (KI), which aims to complement the lexical knowledge into neural dialog models.Instead of further conditioning the knowledge-grounded dialog (KGD) models on externally retrieved knowledge, we seek to integrate knowledge about each input token internally into the model's parameters.To tackle the challenge due to the large scale of lexical knowledge, we adopt the contrastive learning approach and create an effective token-level lexical knowledge retriever that requires only weak supervision mined from Wikipedia.We demonstrate the effectiveness and general applicability of our approach on various datasets and diversified model structures.
Zhiyong Wu 0003, Wei Bi, Xiang Li 0067, Lingpeng Kong, Ben Kao
ACL (1)4
2022 ABC: Attention with Bounded-memory Control
abstract
Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, Noah Smith. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Hao Peng 0009, Jungo Kasai, Nikolaos Pappas 0002, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz 0001, Noah A. Smith
ACL (1)6
2022 Audio-Visual Segmentation
Jinxing Zhou, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong
ECCV (37)8
2022 The Devil in Linear Transformer
abstract
Linear transformers aim to reduce the quadratic space-time complexity of vanilla transformers.However, they usually suffer from degraded performances on various tasks and corpora.In this paper, we examine existing kernel-based linear transformers and identify two key issues that lead to such performance gaps: 1) unbounded gradients in the attention computation adversely impact the convergence of linear transformer models; 2) attention dilution which trivially distributes attention scores over long sequences while neglecting neighbouring structures.To address these issues, we first identify that the scaling of attention matrices is the devil in unbounded gradients, which turns out unnecessary in linear attention as we show theoretically and empirically.To this end, we propose a new linear attention that replaces the scaling operation with a normalization to stabilize gradients.For the issue of attention dilution, we leverage a diagonal attention to confine attention to only neighbouring tokens in early layers.Benefiting from the stable gradients and improved attention, our new linear transformer model, TRANSNORMER, demonstrates superior performance on text classification and language modeling tasks, as well as on the challenging Long-Range Arena benchmark, surpassing vanilla transformer and existing linear variants by a clear margin while being significantly more space-time efficient.The code is available at TRANSNORMER.
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dongxu Li 0003, Lingpeng Kong, Nick Barnes, Yiran Zhong
EMNLP5
2022 UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
abstract
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009
EMNLP19
2022 ZeroGen: Efficient Zero-shot Learning via Dataset Generation
abstract
There is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs).In this paper, we study a flexible and efficient zero-short learning method, ZEROGEN.Given a zero-shot task, we first generate a dataset from scratch using PLMs in an unsupervised manner.Then, we train a tiny task model (e.g., LSTM) under the supervision of the synthesized dataset.This approach allows highly efficient inference as the final task model only has orders of magnitude fewer parameters comparing to PLMs (e.g., GPT2-XL).Apart from being annotation-free and efficient, we argue that ZEROGEN can also provide useful insights from the perspective of datafree model-agnostic knowledge distillation, and unreferenced text generation evaluation.Experiments and analysis on different NLP tasks, namely, text classification, question answering, and natural language inference, show the effectiveness of ZEROGEN.
Jiacheng Ye, Jiahui Gao 0002, Qintong Li, Hang Xu 0004, Jiangtao Feng, Zhiyong Wu 0003, Tao Yu 0009, Lingpeng Kong
EMNLP8
2022 An Empirical Revisiting of Linguistic Knowledge Fusion in Language Understanding Tasks
abstract
Though linguistic knowledge emerges during large-scale language model pretraining, recent work attempt to explicitly incorporate humandefined linguistic priors into task-specific finetuning.Infusing language models with syntactic or semantic knowledge from parsers has shown improvements on many language understanding tasks.To further investigate the effectiveness of structural linguistic priors, we conduct empirical study of replacing parsed graphs or trees with trivial ones (rarely carrying linguistic knowledge e.g., balanced tree) for tasks in the GLUE benchmark.Encoding with trivial graphs achieves competitive or even better performance in fully-supervised and few-shot settings.It reveals that the gains might not be significantly attributed to explicit linguistic priors but rather to more feature interactions brought by fusion layers.Hence we call for attention to using trivial graphs as necessary baselines to design advanced knowledge fusion methods in the future.
Changlong Yu, Tianyi Xiao, Lingpeng Kong, Yangqiu Song, Wilfred Ng
EMNLP3
2022 cosFormer: Rethinking Softmax In Attention
Zhen Qin 0003, Weixuan Sun, Dongxu Li 0003, Yunshen Wei, Baohong Lv, Lingpeng Kong, Yiran Zhong
ICLR8
2022 Revisiting Over-smoothing in BERT from the Perspective of Graph
Jiahui Gao 0002, Hang Xu 0004, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen M. S. Lee, James T. Kwok
ICLR6
2022 Ripple Attention for Visual Perception with Sub-quadratic Complexity
abstract
Transformer architectures are now central to sequence modeling tasks. At its heart is the attention mechanism, which enables effective modeling of long-term dependencies in a sequence. Recently, transformers have been successfully applied in the computer vision domain, where 2D images are first segmented into patches and then treated as 1D sequences. Such linearization, however, impairs the notion of spatial locality in images, which bears important visual clues. To bridge the gap, we propose ripple attention, a sub-quadratic attention mechanism for vision transformers. Built upon the recent kernel-based efficient attention mechanisms, we design a novel dynamic programming algorithm that weights contributions of different tokens to a query with respect to their relative spatial distances in the 2D space in linear observed time. Extensive experiments and analyses demonstrate the effectiveness of ripple attention on various visual tasks.
Huijie Pan, Lingpeng Kong
ICML3
2022 Linear Complexity Randomized Self-attention Mechanism
abstract
Recently, random feature attentions (RFAs) are proposed to approximate the softmax attention in linear time and space complexity by linearizing the exponential kernel. In this paper, we first propose a novel perspective to understand the bias in such approximation by recasting RFAs as self-normalized importance samplers. This perspective further sheds light on an unbiased estimator for the whole softmax attention, called randomized attention (RA). RA constructs positive random features via query-specific distributions and enjoys greatly improved approximation fidelity, albeit exhibiting quadratic complexity. By combining the expressiveness in RA and the efficiency in RFA, we develop a novel linear complexity self-attention mechanism called linear randomized attention (LARA). Extensive experiments across various domains demonstrate that RA and LARA significantly improve the performance of RFAs by a substantial margin.
Lingpeng Kong
ICML3
2022 Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
abstract
We examine the extent to which, in principle, different syntactic and semantic graph representations can complement and improve neural language modeling.Specifically, by conditioning on a subgraph encapsulating the locally relevant sentence history, can a model make better next-word predictions than a pretrained sequential language model alone?With an ensemble setup consisting of GPT-2 and ground-truth graphs from one of 7 different formalisms, we find that the graph information indeed improves perplexity and other metrics.Moreover, this architecture provides a new way to compare different frameworks of linguistic representation.In our oracle graph setup, training and evaluating on English WSJ, semantic constituency structures prove most useful to language modeling performance-outpacing syntactic constituency structures as well as syntactic and semantic dependency structures.
Jakob Prange, Nathan Schneider 0001, Lingpeng Kong
NAACL-HLT3
2022 CoNT: Contrastive Neural Text Generation
abstract
Recently, contrastive learning attracts increasing interests in neural text generation as a new solution to alleviate the exposure bias problem. It introduces a sequence-level training signal which is crucial to generation tasks that always rely on auto-regressive decoding. However, previous methods using contrastive learning in neural text generation usually lead to inferior performance. In this paper, we analyse the underlying reasons and propose a new Contrastive Neural Text generation framework, CoNT. CoNT addresses bottlenecks that prevent contrastive learning from being widely adopted in generation tasks from three aspects -- the construction of contrastive examples, the choice of the contrastive loss, and the strategy in decoding. We validate CoNT on five generation tasks with ten benchmarks, including machine translation, summarization, code comment generation, data-to-text generation and commonsense generation. Experimental results show that CoNT clearly outperforms its baseline on all the ten benchmarks with a convincing margin. Especially, CoNT surpasses previous the most competitive contrastive learning method for text generation, by 1.50 BLEU on machine translation and 1.77 ROUGE-1 on summarization, respectively. It achieves new state-of-the-art on summarization, code comment generation (without external data) and data-to-text generation.
Chenxin An, Jiangtao Feng, Kai Lv 0001, Lingpeng Kong, Xipeng Qiu, Xuanjing Huang 0001
NeurIPS4
2022 A Contrastive Framework for Neural Text Generation
abstract
Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease the probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method---contrastive search---to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach outperforms state-of-the-art text generation methods as evaluated by both human and automatic metrics.
Yixuan Su, Tian Lan 0003, Yan Wang 0060, Dani Yogatama, Lingpeng Kong, Nigel Collier
NeurIPS5
2021 Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
abstract
Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, Ben Kao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhiyong Wu 0003, Lingpeng Kong, Wei Bi, Xiang Li 0067, Ben Kao
ACL/IJCNLP (1)2
2021 Cascaded Head-colliding Attention
abstract
Lin Zheng, Zhiyong Wu, Lingpeng Kong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhiyong Wu 0003, Lingpeng Kong
ACL/IJCNLP (1)3
2021 Random Feature Attention
Hao Peng 0009, Nikolaos Pappas 0002, Dani Yogatama, Roy Schwartz 0001, Noah A. Smith, Lingpeng Kong
ICLR6
2021 Mining influential genes based on deep learning
abstract
BACKGROUND: Currently, large-scale gene expression profiling has been successfully applied to the discovery of functional connections among diseases, genetic perturbation, and drug action. To address the cost of an ever-expanding gene expression profile, a new, low-cost, high-throughput reduced representation expression profiling method called L1000 was proposed, with which one million profiles were produced. Although a set of ~ 1000 carefully chosen landmark genes that can capture ~ 80% of information from the whole genome has been identified for use in L1000, the robustness of using these landmark genes to infer target genes is not satisfactory. Therefore, more efficient computational methods are still needed to deep mine the influential genes in the genome. RESULTS: Here, we propose a computational framework based on deep learning to mine a subset of genes that can cover more genomic information. Specifically, an AutoEncoder framework is first constructed to learn the non-linear relationship between genes, and then DeepLIFT is applied to calculate gene importance scores. Using this data-driven approach, we have re-obtained a landmark gene set. The result shows that our landmark genes can predict target genes more accurately and robustly than that of L1000 based on two metrics [mean absolute error (MAE) and Pearson correlation coefficient (PCC)]. This reveals that the landmark genes detected by our method contain more genomic information. CONCLUSIONS: We believe that our proposed framework is very suitable for the analysis of biological big data to reveal the mysteries of life. Furthermore, the landmark genes inferred from this study can be used for the explosive amplification of gene expression profiles to facilitate research into functional connections.
Lingpeng Kong, Yuanyuan Chen 0014, Fengjiao Xu, Mingmin Xu, Zutan Li, Jingya Fang, Liang-Yun Zhang, Cong Pian
BMC Bioinform.1
2021 Deep6mA: A deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species
abstract
N6-methyladenine (6mA) is an important DNA modification form associated with a wide range of biological processes. Identifying accurately 6mA sites on a genomic scale is crucial for under-standing of 6mA's biological functions. However, the existing experimental techniques for detecting 6mA sites are cost-ineffective, which implies the great need of developing new computational methods for this problem. In this paper, we developed, without requiring any prior knowledge of 6mA and manually crafted sequence features, a deep learning framework named Deep6mA to identify DNA 6mA sites, and its performance is superior to other DNA 6mA prediction tools. Specifically, the 5-fold cross-validation on a benchmark dataset of rice gives the sensitivity and specificity of Deep6mA as 92.96% and 95.06%, respectively, and the overall prediction accuracy is 94%. Importantly, we find that the sequences with 6mA sites share similar patterns across different species. The model trained with rice data predicts well the 6mA sites of other three species: Arabidopsis thaliana, Fragaria vesca and Rosa chinensis with a prediction accuracy over 90%. In addition, we find that (1) 6mA tends to occur at GAGG motifs, which means the sequence near the 6mA site may be conservative; (2) 6mA is enriched in the TATA box of the promoter, which may be the main source of its regulating downstream gene expression.
Zutan Li, Hangjin Jiang, Lingpeng Kong, Yuanyuan Chen 0014, Kun Lang, Xiaodan Fan, Liang-Yun Zhang, Cong Pian
PLoS Comput. Biol.3
2021 Adaptive Semiparametric Language Models
abstract
Abstract We present a language model that combines a large parametric neural network (i.e., a transformer) with a non-parametric episodic memory component in an integrated architecture. Our model uses extended short-term context by caching local hidden states—similar to transformer-XL—and global long-term memory by retrieving a set of nearest neighbor tokens at each timestep. We design a gating function to adaptively combine multiple information sources to make a prediction. This mechanism allows the model to use either local context, short-term memory, or long-term memory (or any combination of them) on an ad hoc basis depending on the context. Experiments on word-based and character-based language modeling datasets demonstrate the efficacy of our proposed method compared to strong baselines.
Dani Yogatama, Cyprien de Masson d'Autume, Lingpeng Kong
Trans. Assoc. Comput. Linguistics3
2020 A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Lei Yu 0008, Wang Ling, Zihang Dai, Dani Yogatama
ICLR1
2020 Syntactic Structure Distillation Pretraining for Bidirectional Encoders
abstract
Textual representation learners trained on large amounts of data have achieved notable success on downstream tasks; intriguingly, they have also performed well on challenging tests of syntactic competence. Hence, it remains an open question whether scalable learners like BERT can become fully proficient in the syntax of natural language by virtue of data scale alone, or whether they still benefit from more explicit syntactic biases. To answer this question, we introduce a knowledge distillation strategy for injecting syntactic biases into BERT pretraining, by distilling the syntactically informative predictions of a hierarchical—albeit harder to scale—syntactic language model. Since BERT models masked words in bidirectional context, we propose to distill the approximate marginal distribution over words in context from the syntactic LM. Our approach reduces relative error by 2–21% on a diverse set of structured prediction tasks, although we obtain mixed results on the GLUE benchmark. Our findings demonstrate the benefits of syntactic biases, even for representation learners that exploit large amounts of data, and contribute to a better understanding of where syntactic biases are helpful in benchmarks of natural language understanding.
Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, Phil Blunsom
Trans. Assoc. Comput. Linguistics2
2020 Better Document-Level Machine Translation with Bayes' Rule
abstract
We show that Bayes’ rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents a compelling benefit because parallel documents are not always available. In our formulation, the posterior probability of a candidate translation is the product of the unconditional (prior) probability of the candidate output document and the “reverse translation probability” of translating the candidate output back into the source language. Our proposed model uses a powerful autoregressive language model as the prior on target language documents, but it assumes that each sentence is translated independently from the target to the source language. Crucially, at test time, when a source document is observed, the document language model prior induces dependencies between the translations of the source sentences in the posterior. The model’s independence assumption not only enables efficient use of available data, but it additionally admits a practical left-to-right beam-search algorithm for carrying out inference. Experiments show that our model benefits from using cross-sentence context in the language model, and it outperforms existing document translation approaches.
Lei Yu 0008, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, Chris Dyer
Trans. Assoc. Comput. Linguistics5
2019 Variational Smoothing in Recurrent Neural Network Language Models
Lingpeng Kong, Gábor Melis, Wang Ling, Lei Yu 0008, Dani Yogatama
ICLR (Poster)1
2019 Episodic Memory in Lifelong Language Learning
abstract
We introduce a lifelong language learning setup where a model needs to learn from a stream of text examples without any dataset identifier. We propose an episodic memory model that performs sparse experience replay and local adaptation to mitigate catastrophic forgetting in this setup. Experiments on text classification and question answering demonstrate the complementary benefits of sparse experience replay and local adaptation to allow the model to continuously learn from new datasets. We also show that the space complexity of the episodic memory module can be reduced significantly (~50-90%) by randomly choosing which examples to store in memory with a minimal decrease in performance. We consider an episodic memory component as a crucial building block of general linguistic intelligence and see our model as a first step in that direction.
Cyprien de Masson d'Autume, Sebastian Ruder, Lingpeng Kong, Dani Yogatama
NeurIPS3
2017 What Do Recurrent Neural Network Grammars Learn About Syntax?
abstract
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith
EACL (1)3
2017 Multitask Learning with CTC and Segmental CRF for Speech Recognition
abstract
Segmental conditional random fields (SCRFs) and connectionist temporal classification (CTC) are two sequence labeling methods used for end-to-end training of speech recognition models. Both models define a transcription probability by marginalizing decisions about latent segmentation alternatives to derive a sequence probability: the former uses a globally normalized joint model of segment labels and durations, and the latter classifies each frame as either an output symbol or a "continuation" of the previous label. In this paper, we train a recognition model by optimizing an interpolation between the SCRF and CTC losses, where the same recurrent neural network (RNN) encoder is used for feature extraction for both outputs. We find that this multitask objective improves recognition accuracy when decoding with either the SCRF or CTC models. Additionally, we show that CTC can also be used to pretrain the RNN encoder, which improves the convergence rate when learning the joint model.
Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith
INTERSPEECH2
2016 Distilling an Ensemble of Greedy Dependency Parsers into One MST Parser
abstract
We introduce two first-order graph-based dependency parsers achieving a new state of the art.The first is a consensus parser built from an ensemble of independently trained greedy LSTM transition-based parsers with different random initializations.We cast this approach as minimum Bayes risk decoding (under the Hamming cost) and argue that weaker consensus within the ensemble is a useful signal of difficulty or ambiguity.The second parser is a "distillation" of the ensemble into a single model.We train the distillation parser using a structured hinge loss objective with a novel cost that incorporates ensemble uncertainty estimates for each possible attachment, thereby avoiding the intractable crossentropy computations required by applying standard distillation objectives to problems with structured outputs.The first-order distillation parser matches or surpasses the state of the art on English, Chinese, and German.
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Noah A. Smith
EMNLP3
2016 Segmental Recurrent Neural Networks for End-to-End Speech Recognition
abstract
We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based acoustic models, it does not rely on an external system to provide features or segmentation boundaries. Instead, this model marginalises out all the possible segmentations, and features are extracted from the RNN trained together with the segmental CRF. In essence, this model is self-contained and can be trained end-to-end. In this paper, we discuss practical training and decoding issues as well as the method to speed up the training in the context of speech recognition. We performed experiments on the TIMIT dataset. We achieved 17.3 phone error rate (PER) from the first-pass decoding --- the best reported result using CRFs, despite the fact that we only used a zeroth-order CRF and without using any language model.
Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith, Steve Renals
INTERSPEECH2
2015 Bayesian Optimization of Text Representations
abstract
When applying machine learning to problems in NLP, there are many choices to make about how to represent input texts.They can have a big effect on performance, but they are often uninteresting to researchers or practitioners who simply need a module that performs well.We apply sequential model-based optimization over this space of choices and show that it makes standard linear models competitive with more sophisticated, expensive state-ofthe-art methods based on latent variables or neural networks on various topic classification and sentiment analysis problems.Our approach is a first step towards black-box NLP systems that work with raw text and do not require manual tuning.
Dani Yogatama, Lingpeng Kong, Noah A. Smith
EMNLP2
2015 Transforming Dependencies into Phrase Structures
abstract
We present a new algorithm for transforming dependency parse trees into phrase-structure parse trees.We cast the problem as structured prediction and learn a statistical model.Our algorithm is faster than traditional phrasestructure parsing and achieves 90.4% English parsing accuracy and 82.4% Chinese parsing accuracy, near to the state of the art on both benchmarks.
Lingpeng Kong, Alexander M. Rush, Noah A. Smith
HLT-NAACL1
2014 A Dependency Parser for Tweets
abstract
We describe a new dependency parser for English tweets, TWEEBOPARSER. The parser builds on several contributions: new syntactic annotations for a corpus of tweets (TWEEBANK), with conventions informed by the domain; adaptations to a statistical parsing algorithm; and a new approach to exploiting out-of-domain Penn Treebank data. Our experiments show that the parser achieves over 80% unlabeled attachment accuracy on our new, high-quality test set and measure the benefit of our contributions. Our dataset and parser can be found at http://www.ark.cs.cmu.edu/TweetNLP.
Lingpeng Kong, Nathan Schneider 0001, Swabha Swayamdipta, Archna Bhatia, Chris Dyer, Noah A. Smith
EMNLP1
2014 Dependency Parsing for Weibo: An Efficient Probabilistic Logic Programming Approach
abstract
Dependency parsing is a core task in NLP, and it is widely used by many applications such as information extraction, question answering, and machine translation. In the era of social media, a big challenge is that parsers trained on traditional newswire corpora typically suffer from the domain mismatch issue, and thus perform poorly on social media data. We present a new GFL/FUDG-annotated Chinese treebank with more than 18K tokens from Sina Weibo (the Chinese equivalent of Twitter). We formulate the dependency parsing problem as many small and parallelizable arc prediction tasks: for each task, we use a programmable probabilistic firstorder logic to infer the dependency arc of a token in the sentence. In experiments, we show that the proposed model outperforms an off-the-shelf Stanford Chinese parser, as well as a strong MaltParser baseline that is trained on the same in-domain data.
William Yang Wang, Lingpeng Kong, Kathryn Mazaitis, William W. Cohen
EMNLP2