Lei Li 0039

dblp:13/7007-39 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
31since 2021 · last 2026
0009-0008-6984-5104ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 7 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI2
2026 Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science With Table-Specific Pretraining
abstract
In data science, predictive tasks such as classification, regression, and missing value imputation are fundamental challenges in tabular data analysis. This research investigates the application of Large Language Models (LLMs) to these tasks. While LLMs excel in natural language understanding, their effectiveness on structured tabular data remains limited due to minimal exposure during pretraining. To address this gap, we construct a large-scale corpus of annotated tables and introduce a tailored pretraining framework. Our trained model achieves significant improvements over baselines, with an average gain of 8.9% in classification and 10.7% in regression tasks. We further evaluate its performance in zero-shot and few-shot prediction, as well as in-context learning scenarios. Extensive experiments demonstrate substantial gains over existing benchmarks, highlighting the potential of LLMs for tabular data processing. Additionally, we apply our approach across multiple open-source LLMs and demonstrate its generalizability. This work establishes a new benchmark for enhancing tabular intelligence through LLM-based pretraining.
Yazheng Yang, Yuqi Wang 0003, Yaxuan Li 0002, Sankalok Sen, Lei Li 0039, Qi Liu 0049
IEEE Trans. Knowl. Data Eng.5
2025 Design Choices for Extending the Context Length of Visual Language Models
abstract
Visual Language Models (VLMs) demonstrate impressive capabilities in processing multimodal inputs, yet applications such as visual agents, which require handling multiple images and high-resolution videos, demand enhanced long-range modeling.Moreover, existing opensource VLMs lack systematic exploration into extending their context length, and commercial models often provide limited details.To tackle this, we aim to establish an effective solution that enhances long context performance of VLMs while preserving their capacities in short context scenarios.Towards this goal, we make the best design choice through extensive experiment settings from data curation to context window extending and utilizing: ( 1) we analyze data sources and length distributions to construct ETVLM -a data recipe to balance the performance across scenarios; (2) we examine existing position extending methods, identify their limitations and propose M-RoPE++ as an enhanced approach; we also choose to solely instruction-tune the backbone with mixed-source data; (3) we discuss how to better utilize extended context windows and propose hybrid-resolution training.Built on the Qwen-VL series model, we propose GI-RAFFE, which is effectively extended to 128K lengths.Evaluated on extensive long context VLM benchmarks such as VideoMME and Viusal Haystacks, our GIRAFFE achieves stateof-the-art performance among similarly sized open-source long VLMs and is competitive with commercial model GPT-4V. 1
Mukai Li, Lei Li 0039, Shansan Gong, Qi Liu 0049
ACL (1)2
2025 VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io.
Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049
CVPR1
2025 Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning Transition
abstract
Chenxin An, Zhihui Xie, Xiaonan Li, Ming Zhong, Shansan Gong, Lei Li, Jun Zhang, Jingjing Xu, Lingpeng Kong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chenxin An, Zhihui Xie 0002, Ming Zhong 0005, Shansan Gong, Lei Li 0039, Jun Zhang 0003, Jingjing Xu 0001, Lingpeng Kong
EMNLP6
2025 ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
abstract
Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks.However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation.To tackle this issue, we first identify the drawbacks of existing solutions (i.e., limited multi-modal reasoning capacities, and insufficient and irrelevant visual descriptions).We then decompose visual reasoning process into two stages: proactive visual perception (i.e., eyesight) and textual reasoning (i.e., wisdom), and introduce a novel visual reasoning framework named PROREASON.This framework features decoupled vision-reasoning capabilities and multi-run proactive perception.Briefly, given a multi-modal question, PRORE-ASON iterates proactive information collection and reasoning until the answer can be concluded with necessary and sufficient visual descriptions.Notably, the disassociation of capabilities allows seamless integration of existing large language models (LLMs) to compensate for the reasoning deficits of LVLMs.Our extensive experiments demonstrate that PROREASON outperforms existing multi-step reasoning frameworks on various benchmarks for both open-source and closed-source models, with the average performance gain reaching 13.2%.Besides, the integration of LLMs allows PROREASON to produce high-quality visual reasoning data, which empowers PRORE-ASON-distilled models (i.e., ProReason-VL and ProReason-Q3) to achieve superior performance in downstream tasks.Our insights into existing solutions and the decoupled perspective for feasible integration of LLMs illuminate future research on visual reasoning techniques, especially LLM-assisted ones.
Jingqi Zhou, Jingwei Dong, Lei Li 0039, Jiahui Gao 0002, Jiyue Jiang, Lingpeng Kong
EMNLP5
2025 Jailbreaking as a Reward Misspecification Problem
abstract
The widespread adoption of large language models (LLMs) has raised concerns about their safety and reliability, particularly regarding their vulnerability to adversarial attacks. In this paper, we propose a new perspective that attributes this vulnerability to reward misspecification during the alignment process. This misspecification occurs when the reward function fails to accurately capture the intended behavior, leading to misaligned model outputs. We introduce a metric ReGap to quantify the extent of reward misspecification and demonstrate its effectiveness and robustness in detecting harmful backdoor prompts. Building upon these insights, we present ReMiss, a system for automated red teaming that generates adversarial prompts in a reward-misspecified space. ReMiss achieves state-of-the-art attack success rates on the AdvBench benchmark against various target aligned LLMs while preserving the human readability of the generated prompts. Furthermore, these attacks on open-source models demonstrate high transferability to closed-source models like GPT-4o and out-of-distribution tasks from HarmBench. Detailed analysis highlights the unique advantages of the proposed reward misspecification objective compared to previous methods, offering new insights for improving LLM safety and robustness.
Zhihui Xie 0002, Jiahui Gao 0002, Lei Li 0039, Zhenguo Li, Qi Liu 0049, Lingpeng Kong
ICLR3
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR1
2025 Why Does the Effective Context Length of LLMs Fall Short?
abstract
Advancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their training lengths. In this work, we attribute this limitation to the left-skewed frequency distribution of relative positions formed in LLMs pretraining and post-training stages, which impedes their ability to effectively gather distant information. To address this challenge, we introduce Shifted Rotray Position Embedding (STRING). STRING shifts well-trained positions to overwrite the original ineffective positions during inference, enhancing performance within their existing training lengths. Experimental results show that without additional training, STRING dramatically improves the performance of the latest large-scale models, such as Llama3.1 70B and Qwen2 72B, by over 10 points on popular long-context benchmarks RULER and InfiniteBench, establishing new state-of-the-art results for open-source LLMs. Compared to commercial models, Llama 3.1 70B with STRING even achieves better performance than GPT-4-128K and clearly surpasses Claude 2 and Kimi-chat.
Chenxin An, Jun Zhang 0003, Ming Zhong 0005, Lei Li 0039, Shansan Gong, Yao Luo, Jingjing Xu 0001, Lingpeng Kong
ICLR4
2025 Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models
abstract
Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.
Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang
ICLR7
2025 SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking Services
abstract
With the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content generation, and communication dynamics. However, recent studies focus on isolated SNS tasks rather than a comprehensive evaluation. In this paper, we introduce SNS-Bench, specially constructed for assessing the abilities of large language models from different Social Networking Services, with a wide range of SNS-related information. SNS-Bench encompasses 8 different tasks such as note classification, query content relevance, and highlight words generation in comments. Finally, 6,658 questions of social media text, including subjective questions, single-choice, and multiple-choice questions, are concluded in SNS-Bench. Further, we evaluate the performance of over 25+ current diverse LLMs on our SNS-Bench. Models with different sizes exhibit performance variations, yet adhere to the scaling law. Moreover, we hope provide more insights to revolutionize the techniques of social network services with LLMs.
Hongcheng Guo, Shaosheng Cao, Fei Zhao 0012, Boyang Wang 0006, Lei Li 0039, Liang Chen 0024, Xinze Lyu, Yao Hu 0002, Zhoujun Li 0001
ICML6
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia4
2025 ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
abstract
Xijia Tao, Shuai Zhong, Lei Li, Qi Liu, Lingpeng Kong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xijia Tao, Shuai Zhong, Lei Li 0039, Qi Liu 0049, Lingpeng Kong
NAACL (Long Papers)3
2024 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
abstract
Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes.However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains.To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension.ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains.Drawing from ArXivCap, we introduce ArXivQA, a questionanswering dataset generated by prompting GPT-4V based on scientific figures.ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs.Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domainspecific training yields substantial performance gains.Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.
Lei Li 0039, Yuqi Wang 0003, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu 0049
ACL (1)1
2024 Large Language Models are not Fair Evaluators
abstract
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui
ACL (1)2
2024 Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
abstract
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li 0005, Deli Chen, Zhifang Sui
ACL (1)2
2024 VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Rundong Gao, Xu Sun 0001, Lu Hou 0002
ECCV (70)2
2024 A Survey on In-context Learning
abstract
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui
EMNLP2
2024 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
abstract
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049
EMNLP1
2023 Can Language Models Understand Physical Concepts?
abstract
Language models (LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.However, it is unclear whether LMs can understand physical concepts in the human world.To investigate this, we design a benchmark VEC that covers the tasks of (i) Visual concepts, such as the shape and material of objects, and (ii) Embodied Concepts, learned from the interaction with the world such as the temperature of objects.Our zero (few)-shot prompting results show that the understanding of certain visual concepts emerges as scaling up LMs, but there are still basic concepts to which the scaling law does not apply.For example, OPT-175B performs close to humans with a zero-shot accuracy of 85% on the material concept, yet behaves like random guessing on the mass concept.Instead, vision-augmented LMs such as CLIP and BLIP achieve a human-level understanding of embodied concepts.Analysis indicates that the rich semantics in visual representation can serve as a valuable source of embodied knowledge.Inspired by this, we propose a distillation method to transfer embodied knowledge from VLMs to LMs, achieving performance gain comparable with that by scaling up parameters of LMs 134×. 1 o 1 : This is a photo of the water.o 2 : This is a photo of a frying oil.Attribute: This is a photo of a cold object.
Lei Li 0039, Jingjing Xu 0001, Qingxiu Dong, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
EMNLP1
2023 Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
abstract
In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks.However, the underlying mechanism of how LLMs learn from the provided context remains under-explored.In this paper, we investigate the working mechanism of ICL through an information flow lens.Our findings reveal that label words in the demonstration examples function as anchors:(1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions.Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL.The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies. 1
Lean Wang, Lei Li 0039, Damai Dai, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
EMNLP2
2023 Can We Edit Factual Knowledge by In-Context Learning?
abstract
Previous studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters.However, the stored knowledge could be false or outdated.Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge.However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs.The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LLMs.Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge.To answer this question, we give a comprehensive empirical study of ICL strategies.Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge.We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method.The code is available at https://github.com/pkunlp-icler/IKE.
Lei Li 0039, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu 0011, Jingjing Xu 0001, Baobao Chang
EMNLP2
2023 FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
abstract
Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the automatic evaluation metrics are consistent with human standards. To address these problems, we propose FETV, a benchmark for Fine-grained Evaluation of Text-to-Video generation. FETV is multi-aspect, categorizing the prompts based on three orthogonal aspects: the major content, the attributes to control and the prompt complexity. FETV is also temporal-aware, which introduces several temporal categories tailored for video generation. Based on FETV, we conduct comprehensive manual evaluations of four representative T2V models, revealing their pros and cons on different categories of prompts from different aspects. We also extend FETV as a testbed to evaluate the reliability of automatic T2V metrics. The multi-aspect categorization of FETV enables fine-grained analysis of the metrics' reliability in different scenarios. We find that existing automatic metrics (e.g., CLIPScore and FVD) correlate poorly with human evaluation. To address this problem, we explore several solutions to improve CLIPScore and FVD, and develop two automatic metrics that exhibit significant higher correlation with humans than existing metrics. Benchmark page: https://github.com/llyx97/FETV.
Yuanxin Liu, Lei Li 0039, Shuhuai Ren, Rundong Gao, Sishuo Chen, Xu Sun 0001, Lu Hou 0002
NeurIPS2
2022 Well-Classified Examples Are Underestimated in Classification with Deep Neural Networks
abstract
The conventional wisdom behind learning deep classification models is to focus on bad-classified examples and ignore well-classified examples that are far from the decision boundary. For instance, when training with cross-entropy loss, examples with higher likelihoods (i.e., well-classified examples) contribute smaller gradients in back-propagation. However, we theoretically show that this common practice hinders representation learning, energy optimization, and margin growth. To counteract this deficiency, we propose to reward well-classified examples with additive bonuses to revive their contribution to the learning process. This counterexample theoretically addresses these three issues. We empirically support this claim by directly verifying the theoretical results or significant performance improvement with our counterexample on diverse tasks, including image classification, graph classification, and machine translation. Furthermore, this paper shows that we can deal with complex scenarios, such as imbalanced classification, OOD detection, and applications under adversarial attacks because our idea can solve these three issues. Code is available at https://github.com/lancopku/well-classified-examples-are-underestimated.
Guangxiang Zhao, Wenkai Yang, Xuancheng Ren, Lei Li 0039, Yunfang Wu, Xu Sun 0001
AAAI4
2022 Rethinking the Promotion Brought by Contrastive Learning to Semi-Supervised Node Classification
abstract
Graph Contrastive Learning (GCL) has proven highly effective in promoting the performance of Semi-Supervised Node Classification (SSNC). However, existing GCL methods are generally transferred from other fields like CV or NLP, whose underlying working mechanism remains underexplored. In this work, we first deeply probe the working mechanism of GCL in SSNC, and find that the promotion brought by GCL is severely unevenly distributed: the improvement mainly comes from subgraphs with less annotated information, which is fundamentally different from contrastive learning in other fields. However, existing GCL methods generally ignore this uneven distribution of annotated information and apply GCL evenly to the whole graph. To remedy this issue and further improve GCL in SSNC, we propose the Topology InFormation gain-Aware Graph Contrastive Learning (TIFA-GCL) framework that considers the annotated information distribution across graph in GCL. Extensive experiments on six benchmark graph datasets, including the enormous OGB-Products graph, show that TIFA-GCL can bring a larger improvement than existing GCL methods in both transductive and inductive settings. Further experiments demonstrate the generalizability and interpretability of TIFA-GCL.
Deli Chen, Yankai Lin 0001, Lei Li 0039, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
IJCAI3
2022 Distributional Correlation-Aware Knowledge Distillation for Stock Trading Volume Prediction
Lei Li 0039, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)1
2022 Alleviating the Knowledge-Language Inconsistency: A Study for Deep Commonsense Knowledge
abstract
Knowledge facts are typically represented by relational triples, while we observe that some commonsense facts are represented by triples whose forms are inconsistent with the corresponding language expressions. For commonsense mining tasks, this inconsistency raises a challenge for the prevailing methods using pre-trained language models that learn the expression of language. However, there are few studies which focus on this inconsistency issue. To fill this empty, in this paper, we term the commonsense knowledge whose triple form is heavily inconsistent with the language expression asdeep commonsense knowledgeand first conduct extensive exploratory experiments to study deep commonsense knowledge. We show that deep commonsense knowledge occupies a significant part of commonsense knowledge, while the conventional methods based on pre-trained language models fail to capture it effectively. We further propose a novel method to mine the deep commonsense knowledge from raw text that is exactly language expression, alleviating the reliance of conventional methods on the triple representation form. Experiments demonstrate that our proposed method substantially improves the performance in mining deep commonsense knowledge.
Yi Zhang 0050, Lei Li 0039, Yunfang Wu, Qi Su 0001, Xu Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Dynamic Knowledge Distillation for Pre-trained Language Models
abstract
Knowledge distillation (KD) has been proved effective for compressing large-scale pretrained language models.However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset.In this paper, we explore whether a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency, regarding the student performance and learning efficiency.We explore the dynamical adjustments on three aspects: teacher model adoption, data selection, and KD objective adaptation.Experimental results show that (1) proper selection of teacher model can boost the performance of student model; (2) conducting KD with 10% informative instances achieves comparable performance while greatly accelerates the training; (3) the student performance can be boosted by adjusting the supervision contribution of different alignment objective.We find dynamic knowledge distillation is promising and provide discussions on potential future directions towards more efficient KD methods. 1
Lei Li 0039, Yankai Lin 0001, Shuhuai Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
EMNLP (1)1
2021 Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification
abstract
Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations.Traditional methods first devise task-specific operations such as Synonym Substitute, then preset the corresponding parameters such as the substitution rate artificially, which require a lot of prior knowledge and are prone to fall into the sub-optimum.Besides, the number of editing operations is limited in the previous methods, which decreases the diversity of the augmented data and thus restricts the performance gain.To overcome the above limitations, we propose a framework named Text AutoAugment (TAA) to establish a compositional and learnable paradigm for data augmentation.We regard a combination of various operations as an augmentation policy and utilize an efficient Bayesian Optimization algorithm to automatically search for the best policy, which substantially improves the generalization capability of models.Experiments on six benchmark datasets show that TAA boosts classification accuracy in low-resource and class-imbalanced regimes by an average of 8.8% and 9.7%, respectively, outperforming strong baselines.1
Shuhuai Ren, Jinchao Zhang 0001, Lei Li 0039, Xu Sun 0001, Jie Zhou 0016
EMNLP (1)3
2021 Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
abstract
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Wenkai Yang, Lei Li 0039, Zhiyuan Zhang 0001, Xuancheng Ren, Xu Sun 0001
NAACL-HLT2
2021 Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition Generation
abstract
Hua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124
NAACL-HLT3
2019 Enhancing Topic-to-Essay Generation with External Commonsense Knowledge
abstract
Automatic topic-to-essay generation is a challenging task since it requires generating novel, diverse, and topic-consistent paragraph-level text with a set of topics as input.Previous work tends to perform essay generation based solely on the given topics while ignoring massive commonsense knowledge.However, this commonsense knowledge provides additional background information, which can help to generate essays that are more novel and diverse.Towards filling this gap, we propose to integrate commonsense from the external knowledge base into the generator through dynamic memory mechanism.Besides, the adversarial training based on a multi-label discriminator is employed to further improve topic-consistency.We also develop a series of automatic evaluation metrics to comprehensively assess the quality of the generated essay.Experiments show that with external commonsense knowledge and adversarial training, the generated essays are more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation.
Lei Li 0039, Fuli Luo, Tianyu Liu 0001, Xu Sun 0001
ACL (1)2
2019 Cross-Modal Commentator: Automatic Machine Commenting Based on Cross-Modal Information
abstract
Automatic commenting of online articles can provide additional opinions and facts to the reader, which improves user experience and engagement on social media platforms.Previous work focuses on automatic commenting based solely on textual content.However, in real-scenarios, online articles usually contain multiple modal contents.For instance, graphic news contains plenty of images in addition to text.Contents other than text are also vital because they are not only more attractive to the reader but also may provide critical information.To remedy this, we propose a new task: cross-model automatic commenting (CMAC), which aims to make comments by integrating multiple modal contents.We construct a largescale dataset for this task and explore several representative methods.Going a step further, an effective co-attention model is presented to capture the dependency between textual and visual information.Evaluation results show that our proposed model can achieve better performance than competitive baselines.1
Zhihan Zhang 0001, Fuli Luo, Lei Li 0039, Chengyang Huang, Xu Sun 0001
ACL (1)4
2019 Pun-GAN: Generative Adversarial Network for Pun Generation
abstract
Fuli Luo, Shunyao Li, Pengcheng Yang, Lei Li, Baobao Chang, Zhifang Sui, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fuli Luo, Shunyao Li, Lei Li 0039, Baobao Chang, Zhifang Sui, Xu Sun 0001
EMNLP/IJCNLP (1)4
2019 Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling
abstract
The visual storytelling (VST) task aims at generating a reasonable and coherent paragraph-level story with the image stream as input. Different from caption that is a direct and literal description of image content, the story in the VST task tends to contain plenty of imaginary concepts that do not appear in the image. This requires the AI agent to reason and associate with the imaginary concepts based on implicit commonsense knowledge to generate a reasonable story describing the image stream. Therefore, in this work, we present a commonsense-driven generative model, which aims to introduce crucial commonsense from the external knowledge base for visual storytelling. Our approach first extracts a set of candidate knowledge graphs from the knowledge base. Then, an elaborately designed vision-aware directional encoding schema is adopted to effectively integrate the most informative commonsense. Besides, we strive to maximize the semantic similarity within the output during decoding to enhance the coherence of the generated text. Results show that our approach can outperform the state-of-the-art systems by a large margin, which achieves a 29\% relative improvement of CIDEr score. With additional commonsense and semantic-relevance based objective, the generated stories are more diverse and coherent.
Fuli Luo, Lei Li 0039, Zhiyi Yin, Xiaodong He 0001, Xu Sun 0001
IJCAI4