Lifeng Shang

dblp:70/4288 · DBLP profile ↗
← Back
85ranked-venue papers
9as first author
60since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 8 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
abstract
Prompt-based essay writing is an effective and common way to assess students' critical thinking skills. Recent work has evaluated the impressive capabilities of Large Language Models (LLMs) on this task. However, most studies focus primarily on English. Those examining LLMs' performance in Chinese often rely on coarse-grained text quality metrics, overlooking the structural and rhetorical complexities of Chinese essays, particularly across diverse genres. We therefore propose EssayBench, a multi-genre benchmark specifically designed for Chinese essay writing, along with a fine-grained, genre-specific scoring framework that hierarchically aggregates scores to better align with human preferences. The dataset comprises 728 real-world prompts across four major genres (Argumentative, Narrative, Descriptive, and Expository), and includes both Open-Ended and Constrained types. Our evaluation protocol is validated through a comprehensive human agreement study. The results show that our protocol aligns well with human judgments, achieving a highest Spearman's correlation of 0.816 and outperforming coarse-grained evaluation methods by an average of 8.6\%. Finally, we benchmark 15 large LLMs, analyzing their strengths and limitations across genres and instruction types. We believe EssayBench offers a more reliable framework for evaluating Chinese essay generation and provides valuable insights for improving LLMs in this domain.
Dongyuan Li, Ding Xia, Fei Mi, Yasheng Wang, Lifeng Shang, Baojun Wang
AAAI6
2026 ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learning
abstract
Tool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning.
Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001
AAAI8
2026 Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
abstract
Bowen Ding, Yuhan Chen, Jiayang Lyu, Jiyao Yuan, Qi Zhu, Shuangshuang Tian, Dantong Zhu, Futing Wang, Heyuan Deng, Fei Mi, Lifeng Shang, Tao Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiayang Lyu, Jiyao Yuan, Qi Zhu 0011, Shuangshuang Tian, Dantong Zhu, Futing Wang, Heyuan Deng, Fei Mi, Lifeng Shang
ACL (1)11
2026 Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
abstract
Chengwu Liu, Yichun Yin, Ye Yuan, Jiaxuan Xie, Botao Li, Siqi Li, Jianhao Shen, Yan Xu, Lifeng Shang, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chengwu Liu 0001, Yichun Yin, Ye Yuan 0016, Jiaxuan Xie, Botao Li, Jianhao Shen, Lifeng Shang, Ming Zhang 0004
ACL (1)9
2026 MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers
abstract
Linrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu, Xihao Yuan, Hanting Chen, Kai Han, Xinghao Chen, Chengjun Zhan, Hanlin xu, Yichun Yin, Lifeng Shang, Feng Wen, Boxing Chen, Yufei Cui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linrui Ma, Chun Hei Lo, Xinyu Wang 0061, Peng Lu 0006, Xihao Yuan, Hanting Chen, Kai Han 0002, Xinghao Chen 0001, Chengjun Zhan, Hanlin Xu, Yichun Yin, Lifeng Shang, Boxing Chen, Yufei Cui
ACL (1)12
2026 How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
abstract
Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu 0011, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, Minlie Huang
ACL (1)8
2026 FORCE: A Benchmark for Formula Reasoning and Comprehension in Academic Papers
Huikang Hu, Xinbang Dai, Xiaoli Shen, Guilin Qi, Lifeng Shang
DASFAA (6)8
2026 Analyzing how pre-trained language models capture factual knowledge using attribution methods
Shaobo Li 0004, Chengjie Sun, Bingquan Liu, Lifeng Shang, Zhenhua Dong, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001
Knowl. Based Syst.5
2025 Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
abstract
Chain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as opaque boxes and do not provide checkable evidence for their judgments, possibly limiting their effectiveness. To address this issue, we draw inspiration from the idea that “the gold standard for supporting a mathematical claim is to provide a proof”. We propose a retrospective, step-aware formal verification framework Safe. Rather than assigning arbitrary scores, we strive to articulate mathematical claims in formal mathematical language Lean 4 at each reasoning step and provide formal proofs to identify hallucinations. We evaluate our framework Safe across multiple language models and various mathematical datasets, demonstrating a significant performance improvement while offering interpretable and verifiable evidence. We also propose FormalStep as a benchmark for step correctness theorem proving with 30,809 formal statements. To the best of our knowledge, our work represents the first endeavor to utilize formal mathematical language Lean 4 for verifying content generated by LLMs, aligning with the reason why formal mathematical languages were created in the first place: to provide a robust foundation for hallucination-prone human-written proofs.
Chengwu Liu 0001, Ye Yuan 0016, Yichun Yin, Zaoyu Chen, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Ming Zhang 0004
ACL (1)8
2025 Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
abstract
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002
ACL (1)8
2025 Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
abstract
Erxin Yu, Jing Li, Ming Liao, Qi Zhu, Boyang Xue, Minghui Xu, Baojun Wang, Lanqing Hong, Fei Mi, Lifeng Shang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Erxin Yu, Jing Li 0049, Ming Liao, Qi Zhu 0011, Boyang Xue, Baojun Wang, Lanqing Hong, Fei Mi, Lifeng Shang
ACL (1)10
2025 Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
abstract
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qiyuan Zhang 0001, Yufei Wang 0005, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ACL (1)8
2025 Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP7
2025 Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
abstract
Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants.
Bo Huang 0017, Yufei Wang 0005, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Wei Wang 0011
ICLR8
2025 ToolACE: Winning the Points of LLM Function Calling
abstract
Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE.
Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen
ICLR22
2025 RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
abstract
With significant efforts in recent studies, LLM-as-a-Judge has become a cost-effective alternative to human evaluation for assessing text generation quality in a wide range of tasks. However, there still remains a reliability gap between LLM-as-a-Judge and human evaluation. One important reason is the lack of guided oracles in the evaluation process. Motivated by the role of reference pervasively used in classic text evaluation, we introduce RevisEval, a novel text generation evaluation paradigm via the response-adapted references. RevisEval is driven by the key observation that an ideal reference should maintain the necessary relevance to the response to be evaluated. Specifically, RevisEval leverages the text revision capabilities of large language models (LLMs) to adaptively revise the response, then treat the revised text as the reference (response-adapted reference) for the subsequent evaluation. Extensive experiments demonstrate that RevisEval outperforms traditional reference-free and reference-based evaluation paradigms that use LLM-as-a-Judge across NLG tasks and open-ended instruction-following tasks. More importantly, our response-adapted references can further boost the classical text metrics, e.g., BLEU and BERTScore, compared to traditional references and even rival the LLM-as-a-Judge. A detailed analysis is also conducted to confirm RevisEval's effectiveness in bias reduction, the impact of inference cost, and reference relevance.
Qiyuan Zhang 0001, Yufei Wang 0005, Tiezheng Yu, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ICLR9
2025 Flat-LoRA: Low-Rank Adaptation over a Flat Loss Landscape
abstract
Fine-tuning large-scale pre-trained models is prohibitively expensive in terms of computation and memory costs. Low-Rank Adaptation (LoRA), a popular Parameter-Efficient Fine-Tuning (PEFT) method, offers an efficient solution by optimizing only low-rank matrices. Despite recent progress in improving LoRA’s performance, the relationship between the LoRA optimization space and the full parameter space is often overlooked. A solution that appears flat in the loss landscape of the LoRA space may still exhibit sharp directions in the full parameter space, potentially compromising generalization. We introduce Flat-LoRA, which aims to identify a low-rank adaptation situated in a flat region of the full parameter space. Instead of adopting the well-established sharpness-aware minimization approach, which incurs significant computation and memory overheads, we employ a Bayesian expectation loss objective to preserve training efficiency. Further, we design a refined strategy for generating random perturbations to enhance performance and carefully manage memory overhead using random seeds. Experiments across diverse tasks—including mathematical reasoning, coding abilities, dialogue generation, instruction following, and text-to-image generation—demonstrate that Flat-LoRA improves both in-domain and out-of-domain generalization. Code is available at https://github.com/nblt/Flat-LoRA.
Zhengbao He, Yasheng Wang, Lifeng Shang, Xiaolin Huang
ICML5
2025 ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL (Long Papers)6
2025 QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
abstract
Recent advancements in Long Chain-of-Thought (CoT) reasoning models have improved performance on complex tasks, but they suffer from overthinking, which generates redundant reasoning steps, especially for simple questions. This paper revisits the reasoning patterns of Long and Short CoT models, observing that the Short CoT patterns offer concise reasoning efficiently, while the Long CoT patterns excel in challenging scenarios where the Short CoT patterns struggle. To enable models to leverage both patterns, we propose Question-Free Fine-Tuning (QFFT), a fine-tuning approach that removes the input question during training and learns exclusively from Long CoT responses. This approach enables the model to adaptively employ both reasoning patterns: it prioritizes the Short CoT patterns and activates the Long CoT patterns only when necessary. Experiments on various mathematical datasets demonstrate that QFFT reduces average response length by more than 50\%, while achieving performance comparable to Supervised Fine-Tuning (SFT). Additionally, QFFT exhibits superior performance compared to SFT in noisy, out-of-domain, and low-resource scenarios.
Wanlong Liu, Junxiao Xu, Fei Yu 0017, Yukang Lin, Ke Ji, Wenyu Chen 0001, Lifeng Shang, Yasheng Wang, Benyou Wang
NeurIPS7
2025 DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement Learning
abstract
Information seeking demands iterative evidence gathering and reflective reasoning, yet large language models (LLMs) still struggle with it in open-web question answering. Existing prompting and supervised fine-tuning (SFT) methods remain fixed by prompt rules or training corpora, and are usually benchmarked only on well-structured wiki sources, limiting real-world adaptability. We introduce $\textbf{WebPuzzle}$, a 24k-sample training and 275-sample test benchmark that evaluates information seeking on the live internet, across both wiki and open-domain queries. Leveraging 7k WebPuzzle instances, we develop $\textbf{DeepDiver}$, a reinforcement-learning (RL) framework that cultivates $\textbf{Search Intensity Scaling (SIS)}$—an emergent ability to escalate search frequency and depth instead of settling on overconfident, under-evidenced answers. With SIS, Qwen2.5-7B-Instruct and Pangu-7B-Reasoner attain performance on real-web tasks comparable to the 671B-parameter DeepSeek-R1. We detail DeepDiver’s curriculum from cold-start SFT to a well designed RL procedure, and show that its seeking policy generalized from closed-ended queries to open-ended generation such as long-form writing. Our results advance adaptive information seeking in LLMs and provide a rigorous benchmark for future work.
Haochen Tan, Chuqiao Kuang, Hanting Chen, Xiaozhe Ren, Yasheng Wang, Lifeng Shang
NeurIPS9
2025 RidgeLoRA: Matrix Ridge Enhanced Low-Rank Adaptation of Large Language Models
abstract
As one of the state-of-the-art parameter-efficient fine-tuning~(PEFT) methods, Low-Rank Adaptation (LoRA) enables model optimization with reduced computational cost through trainable low-rank matrix. However, the low-rank nature makes it prone to produce a decrease in the representation ability, leading to suboptimal performance. In order to break this limitation, we propose RidgeLoRA, a lightweight architecture like LoRA that incorporates novel architecture and matrix ridge enhanced full-rank approximation, to match the performance of full-rank training, while eliminating the need for high memory and a large number of parameters to restore the rank of matrices. We provide a rigorous mathematical derivation to prove that RidgeLoRA has a better upper bound on the representations than vanilla LoRA. Furthermore, extensive experiments across multiple domains demonstrate that RidgeLoRA achieves better performance than other LoRA variants, and can even match or surpass full-rank training.
Junda Zhu 0003, Jun Ai, Yichun Yin, Yasheng Wang, Lifeng Shang, Qun Liu 0001
NeurIPS6
2025 Enhancing inter-sentence coherence of extractive summarization with multitask learning
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
J. Intell. Inf. Syst.3
2025 Dynamic data selection with normalized gradient-based influence approximation for targeted fine-tuning of LLMs
Zige Wang, Qi Zhu 0011, Fei Mi, Yasheng Wang, Lifeng Shang
Knowl. Based Syst.6
2024 Preparing Lessons for Progressive Training on Language Models
abstract
The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
AAAI7
2024 FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
abstract
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011
ACL (1)7
2024 Learning to Edit: Aligning LLMs with Knowledge Editing
abstract
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011
ACL (1)9
2024 M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun, Liangyou Li, Yuxin Jiang, Lifeng Shang, Qun Liu, Kam-Fai Wong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Yusen Sun, Liangyou Li, Lifeng Shang, Qun Liu 0001, Kam-Fai Wong
ACL (1)7
2024 ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
abstract
Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Haochen Tan, Zhijiang Guo, Zhan Shi 0001, Zhili Liu, Yunlong Feng, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Linqi Song
ACL (1)9
2024 Does the Generator Mind Its Contexts? An Analysis of Generative Model Faithfulness under Context Transfer
abstract
he present study introduces the knowledge-augmented generator, which is specifically designed to produce information that remains grounded in contextual knowledge, regardless of alterations in the context. Previous research has predominantly focused on examining hallucinations stemming from static input, such as in the domains of summarization or machine translation. However, our investigation delves into the faithfulness of generative question answering in the presence of dynamic knowledge. Our objective is to explore the existence of hallucinations arising from parametric memory when contextual knowledge undergoes changes, while also analyzing the underlying causes for their occurrence. In order to efficiently address this issue, we propose a straightforward yet effective measure for detecting such hallucinations. Intriguingly, our investigation uncovers that all models exhibit a tendency to generate previous answers as hallucinations. To gain deeper insights into the underlying causes of this phenomenon, we conduct a series of experiments that verify the critical role played by context in hallucination, both during training and testing, from various perspectives.
Xinshuo Hu, Dongfang Li 0002, Yuxiang Wu, Lifeng Shang, Baotian Hu
LREC/COLING5
2024 MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP6
2024 Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
abstract
The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility.
Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR12
2024 Retrieval-based Disentangled Representation Learning with Natural Language Supervision
abstract
Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all its variations within a finite set of factors. However, it is worth noting that most real-world data have linguistic equivalents, typically in the form of textual descriptions. These linguistic counterparts can represent the data and effortlessly decomposed into distinct tokens. In light of this, we present Vocabulary Disentangled Retrieval (VDR), a retrieval-based framework that harnesses natural language as proxies of the underlying data variation to drive disentangled representation learning. Our approach employ a bi-encoder model to represent both data and natural language in a vocabulary space, enabling the model to distinguish dimensions that capture intrinsic characteristics within data through its natural language counterpart, thus facilitating disentanglement. We extensively assess the performance of VDR across 15 retrieval benchmark datasets, covering text-to-text and cross-modal retrieval scenarios, as well as human evaluation. Our experimental results compellingly demonstrate the superiority of VDR over previous bi-encoder retrievers with comparable model size and training costs, achieving an impressive 8.7% improvement in NDCG@10 on the BEIR benchmark, a 5.3\% increase on MS COCO, and a 6.0% increase on Flickr30k in terms of mean recall in the zero-shot setting. Moreover, The results from human evaluation indicate that interpretability of our method is on par with SOTA captioning models.
Jiawei Zhou 0003, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002
ICLR3
2024 Visually Guided Generative Text-Layout Pre-training for Document Intelligence
abstract
Zhiming Mao, Haoli Bai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhiming Mao, Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL-HLT4
2023 Self-Supervised Logic Induction for Explainable Fuzzy Temporal Commonsense Reasoning
abstract
Understanding temporal commonsense concepts, such as times of occurrence and durations is crucial for event-centric language understanding. Reasoning about such temporal concepts in a complex context requires reasoning over both the stated context and the world knowledge that underlines it. A recent study shows massive pre-trained LM still struggle with such temporal reasoning under complex contexts (e.g., dialog) because they only implicitly encode the relevant contexts and fail to explicitly uncover the underlying logical compositions for complex inference, thus may not be robust enough. In this work, we propose to augment LMs with the temporal logic induction ability, which frames the temporal reasoning by defining three modular components: temporal dependency inducer and temporal concept defuzzifier and logic validator. The former two components disentangle the explicit/implicit dependency between temporal concepts across context (before, after, ...) and the specific meaning of fuzzy temporal concepts, respectively, while the validator combines the intermediate reasoning clues for robust contextual reasoning about the temporal concepts. Extensive experimental results on TIMEDIAL, a challenging dataset for temporal reasoning over dialog, show that our method, Logic Induction Enhanced Contextualized TEmporal Reasoning (LECTER), can yield great improvements over the traditional language model for temporal reasoning.
Bibo Cai, Zhouhao Sun, Bing Qin 0001, Ting Liu 0001, Baojun Wang, Lifeng Shang
AAAI7
2023 mCLIP: Multilingual CLIP via Cross-lingual Transfer
abstract
Guanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001
ACL (1)5
2023 Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue Models
abstract
Rui Wang, Jianzhu Bao, Fei Mi, Yi Chen, Hongru Wang, Yasheng Wang, Yitong Li, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rui Wang 0092, Jianzhu Bao, Fei Mi, Yi Chen 0007, Hongru Wang 0003, Yasheng Wang, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu 0001
ACL (1)8
2023 Reusing Pretrained Models by Multi-linear Operators for Efficient Training
abstract
Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
NeurIPS5
2023 Cooperative Game Modeling With Weighted Token-Level Alignment for Audio-Text Retrieval
abstract
Previous audio-text retrieval (ATR) methods primarily concentrate on constructing contrastive pairs between entire audio clips and full caption sentences, while neglecting fine-grained cross-modal relationships. In this letter, we first introduce a weighted token-level alignment (WTA) module for ATR to learn fine-grained semantic interactions. Besides, due to the unavailability of manually labeling the fine-grained sequential correspondence between audio-text pairs, we attempt to model ATR as a cooperative game process to flexibly handle the uncertainty during audio-text semantic interactions. Specifically, we treat audio frames and text words as players and present a game theoretic interaction (GTI) method to assess potential correspondence between audio frames and text words, which can also be seen as an additional learning signal to improve the pure audio-text contrastive learning. Furthermore, to implement multi-level WTA and GTI, we develop a token cluster module to cluster the frames/words and calculate the interaction scores between the clustered tokens. Experiments show that our WTA significantly improves the ATR performance on multiple datasets. By combining our GTI, the retrieval performance is further boosted by a large margin.
Yifei Xin, Baojun Wang, Lifeng Shang
IEEE Signal Process. Lett.3
2022 bert2BERT: Towards Reusable Pretrained Language Models
abstract
Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001
ACL (1)3
2022 Compression of Generative Pre-trained Language Models via Quantization
abstract
Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, Ngai Wong. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Chaofan Tao, Lu Hou 0002, Wei Zhang 0196, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Ping Luo 0002, Ngai Wong 0001
ACL (1)4
2022 Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering
abstract
Jiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Jiawei Zhou 0003, Lifeng Shang, Ke Zhan, Enrui Hu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002
ACL (1)3
2022 LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal Modeling
abstract
Recent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks.However, the pretraining process is computationally expensive due to the requirement of millions of videotext pairs and the redundant data structure of each video.To mitigate these problems, we propose LiteVL, which adapts a pre-trained image-language model BLIP into a video-text model directly on downstream tasks, without heavy pre-training.To enhance the temporal modeling lacking in the image-language model, we propose to add temporal attention modules in the image encoder of BLIP with dynamic temporal scaling.Besides the model-wise adaptation, we also propose a non-parametric pooling mechanism to adaptively reweight the fine-grained video embedding conditioned on the text.Experimental results on text-video retrieval and video question answering show that the proposed LiteVL even outperforms previous video-language pre-trained models by a clear margin, though without any videolanguage pre-training.
Chaofan Tao, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
EMNLP4
2022 Pre-training Language Models with Deterministic Factual Knowledge
abstract
Previous works show that Pre-trained Language Models (PLMs) can capture factual knowledge.However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitive to the changes of prompts when extracting factual knowledge.To mitigate this issue, we propose to let PLMs learn the deterministic relationship between the remaining context and the masked content.The deterministic relationship ensures that the masked factual content can be deterministically inferable based on the existing clues in the context.That would provide more stable patterns for PLMs to capture factual knowledge than randomly masking.Two pre-training tasks are further introduced to motivate PLMs to rely on the deterministic relationship when filling masks.Specifically, we use an external Knowledge Base (KB) to identify deterministic relationships and continuously pre-train PLMs with the proposed methods.The factual knowledge probing experiments indicate that the continuously pre-trained PLMs achieve better robustness in factual knowledge capturing.Further experiments on question-answering datasets show that trying to learn a deterministic relationship with the proposed methods can also help other knowledge-intensive tasks.
Shaobo Li 0004, Lifeng Shang, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001
EMNLP3
2022 G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
abstract
Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks.
Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001
EMNLP5
2022 Exploring extreme parameter compression for pre-trained language models
Benyou Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR3
2022 Towards Efficient Post-training Quantization of Pre-trained Language Models
abstract
Network quantization has gained increasing attention with the rapid growth of large pre-trained language models~(PLMs). However, most existing quantization methods for PLMs follow quantization-aware training~(QAT) that requires end-to-end training with full access to the entire dataset. Therefore, they suffer from slow training, large memory overhead, and data accessibility issues. In this paper, we study post-training quantization~(PTQ) of PLMs, and propose module-wise quantization error minimization~(MREM), an efficient solution to mitigate these issues. By partitioning the PLM into multiple modules, we minimize the reconstruction error incurred by quantization for each module. In addition, we design a new model parallel training strategy such that each module can be trained locally on separate computing devices without waiting for preceding modules, which brings nearly the theoretical training speed-up (e.g., $4\times$ on $4$ GPUs). Experiments on GLUE and SQuAD benchmarks show that our proposed PTQ solution not only performs close to QAT, but also enjoys significant reductions in training time, memory overhead, and data consumption.
Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Irwin King, Michael R. Lyu
NeurIPS3
2021 HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions
abstract
Collecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required to be gathered together to support the answer extraction. In this paper, we propose a new retrieval target, hop, to collect the hidden reasoning evidence from Wikipedia for complex question answering. Specifically, the hop in this paper is defined as the combination of a hyperlink and the corresponding outbound link document. The hyperlink is encoded as the mention embedding which models the structured knowledge of how the outbound link entity is mentioned in the textual context, and the corresponding outbound link document is encoded as the document embedding representing the unstructured knowledge within it. Accordingly, we build HopRetriever which retrieves hops over Wikipedia to answer complex questions. Experiments on the HotpotQA dataset demonstrate that HopRetriever outperforms previously published evidence retrieval methods by large margins. Moreover, our approach also yields quantifiable interpretations of the evidence collection process.
Shaobo Li 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Chengjie Sun, Zhenzhou Ji, Bingquan Liu
AAAI3
2021 Noninvasive Self-attention for Side Information Fusion in Sequential Recommendation
abstract
Sequential recommender systems aim to model users’ evolving interests from their historical behaviors, and hence make customized time-relevant recommendations. Compared with traditional models, deep learning approaches such as CNN and RNN have achieved remarkable advancements in recommendation tasks. Recently, the BERT framework also emerges as a promising method, benefited from its self-attention mechanism in processing sequential data. However, one limitation of the original BERT framework is that it only considers one input source of the natural language tokens. It is still an open question to leverage various types of information under the BERT framework. Nonetheless, it is intuitively appealing to utilize other side information, such as item category or tag, for more comprehensive depictions and better recommendations. In our pilot experiments, we found naive approaches, which directly fuse types of side information into the item embeddings, usually bring very little or even negative effects. Therefore, in this paper, we propose the NOn-inVasive self-Attention mechanism (NOVA) to leverage side information effectively under the BERT framework. NOVA makes use of side information to generate better attention distribution, rather than directly altering the item embeddings, which may cause information overwhelming. We validate the NOVA-BERT model on both public and commercial datasets, and our method can stably outperform the state-of-the-art models with negligible computational overheads.
Chang Liu 0094, Guohao Cai, Zhenhua Dong, Hong Zhu 0003, Lifeng Shang
AAAI6
2021 BinaryBERT: Pushing the Limit of BERT Quantization
abstract
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haoli Bai, Wei Zhang 0196, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Michael R. Lyu, Irwin King
ACL/IJCNLP (1)4
2021 GhostBERT: Generate More Features with Cheap Operations for BERT
abstract
Zhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhiqi Huang 0001, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ACL/IJCNLP (1)3
2021 A Mutual Information Maximization Approach for the Spurious Solution Problem in Weakly Supervised Question Answering
abstract
Zhihong Shao, Lifeng Shang, Qun Liu, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhihong Shao, Lifeng Shang, Qun Liu 0001, Minlie Huang
ACL/IJCNLP (1)2
2021 AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
abstract
Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ACL/IJCNLP (1)3
2021 Improving Unsupervised Question Answering via Summarization-Informed Question Generation
abstract
Question Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach.
Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)2
2021 DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling
abstract
Baojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)6
2021 Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ICANN (3)3
2021 On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 0002, Hao Yang 0006, Qun Liu 0001, Jakob Grue Simonsen
ICLR2
2021 Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma
ICLR3
2021 Improved OOD Generalization via Adversarial Training and Pretraing
abstract
Recently, learning a model that generalizes well on out-of-distribution (OOD) data has attracted great attention in the machine learning community. In this paper, after defining OOD generalization by Wasserstein distance, we theoretically justify that a model robust to input perturbation also generalizes well on OOD data. Inspired by previous findings that adversarial training helps improve robustness, we show that models trained by adversarial training have converged excess risk on OOD data. Besides, in the paradigm of pre-training then fine-tuning, we theoretically justify that the input perturbation robust model in the pre-training stage provides an initialization that generalizes well on downstream OOD data. Finally, various experiments conducted on image classification and natural language understanding tasks verify our theoretical findings.
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma
ICML4
2021 Dual Sequence Transformer for Query-based Interactive Recommendation
abstract
Interactive recommendation has drawn widespread attention from both academia and industry due to its effectiveness in real-world mobile applications. Instead of receiving message passively, customers can exploit further with less effort through generated queries. Usually, such systems mainly contain two main components: query generation and item recommendation. In this paper, we propose a novel framework that models both queries and items in shared latent embedding space via a dual sequence transformer structure, which captures customer's potential interest from the prospect of reconciling the historical queries and corresponding customers interactions. We propose a click-through-rate model to generate query candidates, and a session search model for further more precise information. Comprehensive offline and online experiments are conducted, and the results demonstrate that our proposed dual-sequence-transformer based model can better utilize interaction and improve the accuracy of recommendations.
Guohao Cai, Quanyu Dai, Gang Wang 0056, Zhenhua Dong, Chaoliang Zhang, Xiuqiang He 0001, Lifeng Shang
MDM8
2021 Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001
Neurocomputing4
2020 Dialog State Tracking with Reinforced Data Augmentation
abstract
Neural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
AAAI2
2020 TernaryBERT: Distillation-aware Ultra-low Bit BERT
abstract
Transformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller.
Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)4
2020 An Investigation of Few-Shot Learning in Spoken Term Classification
abstract
202402 bcch
Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qing Li 0001
INTERSPEECH3
2020 Neural Subgraph Isomorphism Counting
abstract
In this paper, we study a new graph learning problem: learning to count subgraph isomorphisms. Different from other traditional graph learning problems such as node classification and link prediction, subgraph isomorphism counting is NP-complete and requires more global inference to oversee the whole graph. To make it scalable for large-scale graphs and patterns, we propose a learning framework that augments different representation learning architectures and iteratively attends pattern and target data graphs to memorize intermediate states of subgraph isomorphism searching for global counting. We develop both small graphs (<= 1,024 subgraph isomorphisms in each) and large graphs (<= 4,096 subgraph isomorphisms in each) sets to evaluate different representation and interaction modules. A mutagenic compound dataset, MUTAG, is also used to evaluate neural models and demonstrate the success of transfer learning. While the learning based approach is inexact, we are able to generalize to count large patterns and data graphs in linear time compared to the exponential time of the original NP-complete problem. Experimental results show that learning based subgraph isomorphism counting can speed up the traditional algorithm, VF2, 10-1,000 times with acceptable errors. Domain adaptation based on fine-tuning also shows the usefulness of our approach in real-world applications.
Xin Liu 0039, Haojie Pan, Mutian He 0001, Yangqiu Song, Xin Jiang 0002, Lifeng Shang
KDD6
2020 DynaBERT: Dynamic BERT with Adaptive Width and Depth
abstract
The pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compress the large BERT model to a fixed smaller size, and can not fully satisfy the requirements of different edge devices with various hardware performances. In this paper, we propose a novel dynamic BERT model (abbreviated as DynaBERT), which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. Comprehensive experiments under various efficiency constraints demonstrate that our proposed dynamic BERT (or RoBERTa) at its largest size has comparable performance as BERT-base (or RoBERTa-base), while at smaller widths and depths consistently outperforms existing BERT compression methods. Code is available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/DynaBERT.
Lu Hou 0002, Zhiqi Huang 0001, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
NeurIPS3
2019 Decomposable Neural Paraphrase Generation
abstract
Paraphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level.This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence at different levels of granularity in a disentangled way.Specifically, the model is composed of multiple encoders and decoders with different structures, each of which corresponds to a specific granularity.The empirical study shows that the decomposition mechanism of DNPG makes paraphrase generation more interpretable and controllable.Based on DNPG, we further develop an unsupervised domain adaptation method for paraphrase generation.Experimental results show that the proposed model achieves competitive in-domain performance compared to the state-of-the-art neural models, and significantly better performance when adapting to a new domain.What is the population of New York?How many people is there in NYC?Who wrote the Winnie the Pooh books?Who is the author of winnie the pooh?What is the best phone to buy below 15k?Which are best mobile phones to buy under 15000?How can I be a good geologist?What should I do to be a great geologist?How do I reword a sentence to avoid plagiarism?How can I paraphrase my essay and avoid plagiarism?
Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001
ACL (1)3
2018 Paraphrase Generation with Deep Reinforcement Learning
abstract
Automatic generation of paraphrases from a given sentence is an important yet challenging task in natural language processing (NLP).In this paper, we present a deep reinforcement learning approach to paraphrase generation.Specifically, we propose a new framework for the task, which consists of a generator and an evaluator, both of which are learned from data.The generator, built as a sequenceto-sequence learning model, can produce paraphrases given a sentence.The evaluator, constructed as a deep matching model, can judge whether two sentences are paraphrases of each other.The generator is first trained by deep learning and then further fine-tuned by reinforcement learning in which the reward is given by the evaluator.For the learning of the evaluator, we propose two methods based on supervised learning and inverse reinforcement learning respectively, depending on the type of available training data.Experimental results on two datasets demonstrate the proposed models (the generators) can produce more accurate paraphrases and outperform the stateof-the-art methods in paraphrase generation in both automatic evaluation and human evaluation.
Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Hang Li 0001
EMNLP3
2017 Neural Machine Translation with Reconstruction
abstract
Although end-to-end Neural Machine Translation (NMT) has achieved remarkable progress in the past two years, it suffers from a major drawback: translations generated by NMT systems often lack of adequacy. It has been widely observed that NMT tends to repeatedly translate some source words while mistakenly ignoring other words. To alleviate this problem, we propose a novel encoder-decoder-reconstructor framework for NMT. The reconstructor, incorporated into the NMT model, manages to reconstruct the input source sentence from the hidden layer of the output target sentence, to ensure that the information in the source side is transformed to the target side as much as possible. Experiments show that the proposed framework significantly improves the adequacy of NMT output and achieves superior translation result over state-of-the-art NMT and statistical MT systems.
Zhaopeng Tu, Yang Liu 0005, Lifeng Shang, Hang Li 0001
AAAI3
2016 Neural Generative Question Answering
Xin Jiang 0002, Zhengdong Lu, Lifeng Shang, Hang Li 0001
IJCAI4
2015 Neural Responding Machine for Short-Text Conversation
abstract
Lifeng Shang, Zhengdong Lu, Hang Li. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Lifeng Shang, Zhengdong Lu, Hang Li 0001
ACL (1)1
2015 Multimodal Convolutional Neural Networks for Matching Image and Sentence
abstract
In this paper, we propose multimodal convolutional neural networks (m-CNNs) for matching image and sentence. Our m-CNN provides an end-to-end framework with convolutional architectures to exploit image representation, word composition, and the matching relations between the two modalities. More specifically, it consists of one image CNN encoding the image content and one matching CNN modeling the joint representation of image and sentence. The matching CNN composes different semantic fragments from words and learns the inter-modal relations between image and the composed fragments at different levels, thus fully exploit the matching relations between image and sentence. Experimental results demonstrate that the proposed m-CNNs can effectively capture the information necessary for image and sentence matching. More specifically, our proposed m-CNNs significantly outperform the state-of-the-art approaches for bidirectional image and sentence retrieval on the Flickr8K and Flickr30K datasets.
Zhengdong Lu, Lifeng Shang, Hang Li 0001
ICCV3
2014 Human motion variation synthesis with multivariate Gaussian processes
abstract
ABSTRACT Human motion variation synthesis is important for crowd simulation and interactive applications to enhance synthesis quality. In this paper, we propose a novel generative probabilistic model to synthesize variations of human motion. Our key idea is to model the conditional distribution of each joint via a multivariate Gaussian process model, namely semiparametric latent factor model (SLFM). SLFM can effectively model the correlations between degrees of freedom (DOFs) of joints rather than dealing with each DOF separately as implemented in existing methods. A detailed evaluation is performed to show that the proposed approach can effectively synthesize variations of different types of motions. Motions generated by our method show a richer variation compared with existing ones. Finally, our user study shows that the synthesized motion has a similar level of naturalness to captured human motions. Our method is best applied in computer games and animations to introduce motion variations. Copyright © 2014 John Wiley & Sons, Ltd.
Liuyang Zhou, Lifeng Shang, Hubert P. H. Shum, Howard Leung
Comput. Animat. Virtual Worlds2
2014 A Robust Likelihood Function for 3D Human Pose Tracking
abstract
Recent works on 3D human pose tracking using unsupervised methods typically focus on improving the optimization framework to find a better maximum in the likelihood function (i.e., the tracker). In contrast, in this paper, we focus on improving the likelihood function, by making it more robust and less ambiguous, thus making the optimization task easier. In particular, we propose an exponential chamfer distance for model matching that is robust to small pose changes, and a part-based model that is better able to localize partially occluded and overlapping parts. Using a standard annealing particle filter and simple diffusion motion model, the proposed likelihood function obtains significantly lower error than other unsupervised tracking methods on the HumanEva dataset. Noting that the joint system of the tracker’s body model is different than the joint system of the motion capture ground-truth model, we propose a novel method for transforming between the two joint systems. Applying this bias correction, our part-based likelihood obtains results equivalent to state-of-the-art supervised tracking methods.
Lifeng Shang, Antoni B. Chan
IEEE Trans. Image Process.2
2014 Spatial temporal pyramid matching using temporal sparse representation for human motion retrieval
Liuyang Zhou, Zhiwu Lu 0001, Howard Leung, Lifeng Shang
Vis. Comput.4
2010 A Temporal Latent Topic Model for Facial Expression Recognition
Lifeng Shang, Kwok-Ping Chan
ACCV (4)1
2010 Real-time large scale near-duplicate web video retrieval
abstract
Near-duplicate video retrieval is becoming more and more important with the exponential growth of the Web. Though various approaches have been proposed to address this problem, they are mainly focusing on the retrieval accuracy while infeasible to query on Web scale video database in real time. This paper proposes a novel method to address the efficiency and scalability issues for near-duplicate We video retrieval. We introduce a compact spatiotemporal feature to represent videos and construct an efficient data structure to index the feature to achieve real-time retrieving performance. This novel feature leverages relative gray-level intensity distribution within a frame and temporal structure of videos along frame sequence. The new index structure is proposed based on inverted file to allow for fast histogram intersection computation between videos. To demonstrate the effectiveness and efficiency of the proposed methods we evaluate its performance on an open Web video data set containing about 10K videos and compare it with four existing methods in terms of precision and time complexity. We also test our method on a data set containing about 50K videos and 11M key-frames. It takes on average 17ms to execute a query against the whole 50K Web video data set.
Lifeng Shang, Linjun Yang, Kwok-Ping Chan, Xian-Sheng Hua 0001
ACM Multimedia1
2009 Collaborative resource discovery in social tagging systems
abstract
Social tagging systems which allow users to create, edit and share collections of internet resources associated with tags in a collaborative fashion are growing in popularity in recent years. The rapidly growing amount of shared data in these folksonomies, i.e., taxonomies created by the folk, presents new technical challenges involved with discovering resources which are likely of interest to the user. Social tags which reflect the meaning of resources from the user's points of view provide an opportunity to enhance the quality of retrieval. In this paper, we introduce a novel framework to search relevant resources to the user query by incorporating information obtained from folksonomies' underlying data structures consisting of a set of user/tag/resource triplets. In contrast to traditional retrieval and recommendation techniques which represent a collection by a matrix, we represent our data as a third-order tensor on which a novel Cube Latent Semantic Indexing (CubeLSI) technique is proposed to capture latent semantic associations between tags. With the latent semantic representation we show how to rank relevant resources according to their relevance to user queries. The excellent performance of the method is demonstrated by an experimental evaluation on the deli.cio.us dataset. Copyright 2009 ACM.
Bin Bi, Lifeng Shang, Ben Kao
CIKM2
2009 Nonparametric discriminant HMM and application to facial expression recognition
abstract
This paper presents a nonparametric discriminant HMM and applies it to facial expression recognition. In the proposed HMM, we introduce an effective nonparametric output probability estimation method to increase the discrimination ability at both hidden state level and class level. The proposed method uses a nonparametric adaptive kernel to utilize information from all classes and improve the discrimination at class level. The discrimination between hidden states is increased by defining membership coefficients which associate each reference vector with hidden states. The adaption of such coefficients is obtained by the Expectation Maximization (EM) method. Furthermore, we present a general formula for the estimation of output probability, which provides a way to develop new HMMs. Finally, we evaluate the performance of the proposed method on the CMU expression database and compare it with other nonparametric HMMs.
Lifeng Shang, Kwok-Ping Chan
CVPR1
2009 Constrained ZIP code segmentation by a PCNN-based thinning algorithm
Lifeng Shang, Zhang Yi 0001, Luping Ji
Neurocomputing1
2008 Temporal Exemplar-Based Bayesian Networks for Facial Expression Recognition
abstract
We present a Temporal Exemplar-based Bayesian Networks (TEBNs) for facial expression recognition. The proposed Bayesian Networks (BNs) consists of three layers: Observation layer, Exemplars layer and Prior Knowledge layer. In the Exemplars layer, exemplar-based model is integrated with BNs to improve the accuracy of probability estimation. In the Prior Knowledge layer, static BNs is extended to Temporal BNs by considering historical observations to model temporal behavior of facial expression. Experiment on CMU expression database illustrates that the proposed TEBNs is very efficient in modeling the evolution of facial deformation.
Lifeng Shang, Kwok-Ping Chan
ICMLA1
2008 An improved pulse coupled neural network for image processing
Luping Ji, Zhang Yi 0001, Lifeng Shang
Neural Comput. Appl.3
2007 A class of binary images thinning using two PCNNs
Lifeng Shang, Zhang Yi 0001
Neurocomputing1
2007 Binary Image Thinning Using Autowaves Generated by PCNN
Lifeng Shang, Zhang Yi 0001, Luping Ji
Neural Process. Lett.1
2007 Binary Fingerprint Image Thinning Using Template-Based PCNNs
abstract
This correspondence presents a coarse-to-fine binary-image-thinning algorithm by proposing a template-based pulse-coupled neural-network model. Under the control of coupled templates, this algorithm iteratively skeletonizes a binary image by changing the load signals of pulse neurons. A direction-constraining scheme for avoiding fingerprint ridge spikes has been discussed. Experiments show that this algorithm is effective for fingerprint thinning, as well as other common images. Moreover, this algorithm can be coupled with a fingerprint identification system to improve the recognition performance.
Luping Ji, Zhang Yi 0001, Lifeng Shang, Xiaorong Pu
IEEE Trans. Syst. Man Cybern. Part B3
2006 Rigid medical image registration using PCA neural network
Lifeng Shang, Jiancheng Lv 0001, Zhang Yi 0001
Neurocomputing1