EDBT 2026 Demo / reviewers in the wild / expert
Peng Li 0030
dblp:83/6353-30
· DBLP profile ↗
98ranked-venue papers
5as first author
69since 2021 · last 2026
0000-0003-1374-5979ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 5 first-author · 67 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 12 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visual-Friendly Concept Protection via Selective Adversarial PerturbationsabstractPersonalized concept generation by tuning diffusion models with a few images raises potential legal and ethical concerns regarding privacy and intellectual property rights. Researchers attempt to prevent malicious personalization using adversarial perturbations. However, previous efforts have mainly focused on the effectiveness of protection while neglecting the visibility of perturbations. They utilize global adversarial perturbations, which introduce noticeable alterations to original images and significantly degrade visual quality. In this work, we propose the Visual-Friendly Concept Protection (VCPro) framework, which prioritizes the protection of key concepts chosen by the image owner through adversarial perturbations with lower perceptibility. To ensure these perturbations are as inconspicuous as possible, we introduce a relaxed optimization objective to identify the least perceptible yet effective adversarial perturbations, solved using the Lagrangian multiplier method. Qualitative and quantitative experiments validate that VCPro achieves a better trade-off between the visibility of perturbations and protection effectiveness, effectively prioritizing the protection of target concepts in images with less perceptible perturbations. Xiaoyue Mi, Fan Tang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005 |
AAAI | 5 |
| 2026 | Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningabstractXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Lei, Chenliang Li 0003, Yuning Wu 0001, Kaiming Liu, Weizhou Shen, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Ya-Qin Zhang, Yang Liu 0005 |
ACL (1) | 6 |
| 2026 | Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent CollaborationabstractWith the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large quantity of retrieved knowledge to solve complex tasks. However, the limited context window of LLMs obstructs scaling the amount of external knowledge input, prohibiting further improvement. Existing context window extension methods inevitably cause information loss. LLM-based multi-agent methods emerge as a new paradigm to handle massive input in a distributional manner, where we identify two core bottlenecks in existing agent orchestration designs. In this work, we develop a multi-agent framework, \textbf{\ExtAgents}, to overcome the bottlenecks and enable better scalability in inference-time knowledge integration without longer-context training. Benchmarked with our enhanced multi-hop question answering test, \textbf{$\boldsymbol{\infty}$Bench+}, and other public test sets including long survey generation, \ExtAgents significantly enhances the performance over existing non-training methods with the same amount of external knowledge input, regardless of whether it falls \emph{within or exceeds the context window}. Moreover, the method maintains efficiency due to high parallelism. We believe further study in the coordination of LLM agents on increasing external knowledge input could benefit real-world applications. Zhennan Wan, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Yang Liu 0005 |
ACL (1) | 3 |
| 2026 | MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingabstractFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang, Chenliang Li, Weizhou Shen, Jiyue Guo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fuwen Luo, Shengfeng Lou, Chi Chen 0005, Ziyue Wang 0002, Chenliang Li 0003, Weizhou Shen, Jiyue Guo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Yang Liu 0005 |
ACL (1) | 8 |
| 2026 | GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language ModelsabstractZhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen, Peng Li, Jinsong Su, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen 0007, Peng Li 0030, Jinsong Su, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 6 |
| 2026 | InstructDiff: Domain-Adaptive Data Selection via Contrastive Entropy for Efficient LLM Fine-TuningabstractJunyou Su, He Zhu, Xiao Luo, Liyu Zhang, Hong-Yu Zhou, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junyou Su, Xiao Luo 0001, Liyu Zhang 0010, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 7 |
| 2026 | SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksabstractTianyi Wang, Yixia Li, Long Li, Yibiao Chen, Shaohan Huang, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Yibiao Chen, Shaohan Huang, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 7 |
| 2026 | Interactive Visual Assessment for Text-to-Image Generation ModelsabstractVisual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an isolated three-phase framework: test input collection, model output generation, and user assessment. These fashions suffer from fixed coverage, evolving difficulty, and data leakage risks, limiting their effectiveness in comprehensively evaluating increasingly complex generation models. To address these limitations, we propose DyEval, an LLM-powered dynamic interactive visual assessment framework that facilitates collaborative evaluation between humans and generative models for text-to-image systems. DyEval features an intuitive visual interface that enables users to interactively explore and analyze model behaviors, while adaptively generating hierarchical, fine-grained, and diverse textual inputs to continuously probe the capability boundaries of the models based on their feedback. Additionally, to provide interpretable analysis for users to further improve tested models, we develop a contextual reflection module that mines failure triggers of test inputs and reflects model potential failure patterns, supporting in-depth analysis using the logical reasoning ability of LLM. Qualitative and quantitative experiments demonstrate that DyEval can effectively help users identify max up to 2.56 timesmore generation failures than conventional methods, and uncover complex and rare failure patterns, such as issues with pronoun generation and specific cultural context generation. Our framework provides valuable insights for improving generative models and has broad implications for advancing the reliability and capabilities of visual generation systems across various domains. Xiaoyue Mi, Fan Tang, Juan Cao 0001, Qiang Sheng 0001, Ziyao Huang 0002, Peng Li 0030, Yang Liu 0005, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | ActiView: Evaluating Active Perception Ability for Multimodal Large Language ModelsabstractZiyue Wang, Chi Chen, Fuwen Luo, Yurui Dong, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang, Peng Li, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziyue Wang 0002, Chi Chen 0005, Fuwen Luo, Yurui Dong 0001, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang 0014, Peng Li 0030, Yang Liu 0005 |
ACL (1) | 8 |
| 2025 | Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language ModelsabstractSymbols such as numerical sequences, chemical formulas, and table delimiters exist widely, playing important roles in symbol-related tasks such as abstract reasoning, chemical property prediction, and tabular question-answering. Compared to tasks based on natural language expressions, large language models (LLMs) have limitations in understanding and reasoning on symbol-based representations, making it difficult for them to handle symbol-related problems. In this paper, we propose symbol-to-language (S2L), a method that converts symbol-based representations to language-based representations, providing valuable information for language models during reasoning. We found that, for both closed-source and open-source LLMs, the capability to solve symbol-related problems can be largely enhanced by incorporating such language-based representations. For example, by employing S2L for GPT-4, there can be substantial improvements of +21.9% and +9.5% accuracy for 1D-ARC and Dyck language tasks, respectively. There is also a consistent improvement in other six general symbol-related tasks such as table understanding and Tweet analysis. We release the GPT logs in https://github.com/THUNLP-MT/symbol2language. Yile Wang 0001, Sijie Cheng, Zixin Sun, Peng Li 0030, Yang Liu 0005 |
COLING | 4 |
| 2025 | AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient OptimizationabstractRecently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS1, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.2 Yiyang Du, Xiaochen Wang 0002, Chi Chen 0005, Jiabo Ye, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Zhifang Sui, Maosong Sun 0001, Yang Liu 0005 |
CVPR | 6 |
| 2025 | G2: Guided Generation for Enhanced Output Diversity in LLMsabstractLarge Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks.However, these models exhibit a critical limitation in output diversity, often generating highly similar content across multiple attempts.This limitation significantly affects tasks requiring diverse outputs, from creative writing to reasoning.Existing solutions, like temperature scaling, enhance diversity by modifying probability distributions but compromise output quality.We propose Guide-to-Generation (G2), a trainingfree plug-and-play method that enhances output diversity while preserving generation quality.G2 employs a base generator alongside dual Guides, which guide the generation process through decoding-based interventions to encourage more diverse outputs conditioned on the original query.Comprehensive experiments demonstrate that G2 effectively improves output diversity while maintaining an optimal balance between diversity and quality. Zhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen 0007, Weihua Luo, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
EMNLP | 6 |
| 2025 | LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsabstractExisting MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such tool-based support, a solitary MLLM still offers only a partial understanding of long videos, resulting in limited performance. In order to better address long video tasks, we introduce LVAgent, the first framework enabling multi-round dynamic collaboration of MLLM agents in long video understanding. Our method consists of four key steps: 1) Selection: We pre-select appropriate agents from the model library to form optimal agent teams based on different tasks. 2) Perception: We design an effective retrieval scheme for long videos to improve the coverage of critical temporal segments while maintaining computational efficiency. 3) Action: Agents answer long video questions and exchange reasons. 4) Reflection: We evaluate each agent's performance in each round of discussion and optimize the agent team for dynamic collaboration. The agents iteratively refine their answers by multi-round dynamical collaboration of MLLM agents. LVAgent is the first agent system method that outperforms all closed-source models (like GPT-4o) and open-source models (like InternVL-2.5 and Qwen2-VL) in the long video understanding tasks. Our LVAgent achieves an accuracy of 80\% on four mainstream long video understanding tasks. Notably, LVAgent improves accuracy by 13.3\% on LongVideoBench. Code is available at https://github.com/64327069/LVAgent. Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu 0003, Peng Li 0030, Yali Wang 0001 |
ICCV | 6 |
| 2025 | Adversarial Robust Memory-Based Continual LearnerabstractDespite the remarkable advances that have been made in continual learning, the adversarial vulnerability of such methods has not been fully discussed. We delve into the adversarial robustness of memory-based continual learning algorithms and observe limited robustness improvement by directly applying adversarial training techniques. Preliminary studies reveal the twin challenges for building adversarial robust continual learners: accelerated forgetting in continual learning and gradient obfuscation in adversarial robustness. In this study, we put forward a novel adversarial robust memory-based continual learner that adjusts data logits to mitigate the forgetting of pasts caused by adversarial samples. Furthermore, we devise a gradient-based data selection mechanism to overcome the gradient obfuscation caused by limited stored data. The proposed approach can widely integrate with existing memory-based continual learning as well as adversarial training algorithms in a plug-and-play way. Extensive experiments on Split-CIFAR10/100 and Split-Tiny-ImageNet demonstrate the effectiveness of our approach, achieving up to 8.13% higher accuracy for adversarial data. Xiaoyue Mi, Fan Tang, Zonghan Yang, Danding Wang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005 |
ICCV | 6 |
| 2025 | How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in an Extensible Escape Game
Ziyue Wang 0002, Yurui Dong 0001, Fuwen Luo, Minyuan Ruan, Zhili Cheng, Chi Chen 0005, Peng Li 0030, Yang Liu 0005 |
ICCV | 7 |
| 2025 | Contrastive Private Data Synthesis via Weighted Multi-PLM FusionabstractSubstantial quantity and high quality are the golden rules of making a good training dataset with sample privacy protection equally important. Generating synthetic samples that resemble high-quality private data while ensuring Differential Privacy (DP), a formal privacy guarantee, promises scalability and practicality. However, existing methods relying on pre-trained models for data synthesis often struggle in data-deficient scenarios, suffering from limited sample size, inevitable generation noise and existing pre-trained model bias. To address these challenges, we propose a novel contr**A**stive private data **S**ynthesis via **W**eighted multiple **P**re-trained generative models framework, named as **WASP**. WASP utilizes limited private samples for more accurate private data distribution estimation via a Top-*Q* voting mechanism, and leverages low-quality synthetic samples for contrastive generation via collaboration among dynamically weighted multiple pre-trained models. Extensive experiments on 6 well-developed datasets with 6 open-source and 3 closed-source PLMs demonstrate the superiority of WASP in improving model performance over diverse downstream tasks. Code is available at https://github.com/LindaLydia/WASP. Tianyuan Zou, Yang Liu 0165, Peng Li 0030, Yufei Xiong, Jianqing Zhang, Xiaozhou Ye, Ye Ouyang, Ya-Qin Zhang |
ICML | 3 |
| 2025 | Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency BrakingabstractAutomatic Emergency Braking (AEB) systems are a crucial component in ensuring the safety of passengers in autonomous vehicles. Conventional AEB systems primarily rely on closed-set perception modules to recognize traffic conditions and assess collision risks. To enhance the adaptability of AEB systems in open scenarios, we propose Dual-AEB, a system combines an advanced multimodal large language model (MLLM) for comprehensive scene understanding and a conventional rule-based rapid AEB to ensure quick response times. To the best of our knowledge, Dual-Aebis the first method to incorporate MLLMs within AEB systems. Through extensive experimentation, we have validated the effectiveness of our method. Codes will be publicly available at https://github.com/ChipsICU/Dual-AEB. Wei Zhang 0012, Pengfei Li 0007, Bingchuan Sun, Qihao Jin, Guangjun Bao, Shibo Rui, Wenchao Ding 0001, Peng Li 0030 |
ICRA | 10 |
| 2025 | Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive VehiclesabstractWhile the capabilities of autonomous driving have advanced rapidly, merging into dense traffic remains a significant challenge, many motion planning methods for this scenario have been proposed but it is hard to evaluate them. Most existing closed-loop simulators rely on rule-based controls for other vehicles, which results in a lack of diversity and randomness, thus failing to accurately assess the motion planning capabilities in highly interactive scenarios. Moreover, traditional evaluation metrics are insufficient for comprehensively evaluating the performance of merging in dense traffic. In response, we proposed a closed-loop evaluation benchmark for assessing motion planning capabilities in merging scenarios. Our approach involves other vehicles trained in large scale datasets with micro-behavioral characteristics that significantly enhance the complexity and diversity. Additionally, we have restructured the evaluation mechanism by leveraging Large Language Models (LLMs) to assess each autonomous vehicle merging onto the main lane. Extensive experiments and test-vehicle deployment have demonstrated the progressiveness of this benchmark. Through this benchmark, we have obtained an evaluation of existing methods and identified common issues. The simulation environment and evaluation process can be accessed at https://github.com/WZM5853/Bench4Merge. Zhengming Wang, Pengfei Li 0007, Zhaohan Li, Bo Zhang 0106, Peng Li 0030 |
IROS | 7 |
| 2025 | EditEval: Towards Comprehensive and Automatic Evaluation for Text-guided Video EditingabstractRecently, video editing task has gained widespread attention due to its practical applications and rapid advancements. However, current automatic evaluation metrics for video editing are mostly poorly aligned with human judgments. Thus, researchers heavily rely on human evaluation, which is not only labor-intensive but also difficult to ensure consistency and objectivity. To address these issues, we propose EditEval, the largest-ever video editing benchmark to comprehensively evaluate the performance of video editing models in three aspects: Textual Faithfulness, Frame Consistency, and Video Fidelity. It includes 200 video clips and 1,010 text prompts, from which 160 instances are sampled to generate 1,280 edited videos using eight open-source video editing models, accompanied by human annotations. Furthermore, we propose EditScore, leveraging the advanced reasoning and comprehension capabilities of Multi-modal Large Language Models (MLLMs) as evaluators to assess edited videos across the aforementioned aspects. Experiments show that the best-performing video editing model only reaches an average score of 3.16 (out of a perfect 5), highlighting the challenge of EditEval. Besides, results from more than 10 MLLMs demonstrate the great potential of utilizing EditScore for automatic evaluation. Notably, for textual faithfulness, EditScore equipped with LLaVA-OneVision-7B achieves a significantly higher Pearson Correlation score compared to previous methods based on CLIP (0.50 vs 0.22). The code and dataset are available at: https://github.com/XMUDeepLIT/EditEval Bingshuai Liu, Ante Wang, Zijun Min, Chenyang Lyu, Longyue Wang, Xu Han 0007, Peng Li 0030, Jinsong Su |
ACM Multimedia | 8 |
| 2025 | Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsabstractThe growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as “LLM-as-a-judge”. However, improving its alignment with human preferences without complex prompts or fine-tuning remains challenging. Previous studies mainly optimize based on shallow outputs, overlooking rich cross-layer representations. In this work, motivated by preliminary findings that middle-to-upper layers encode semantically and task-relevant representations that are often more aligned with human judgments than the final layer, we propose LAGER, a post-hoc, plug-and-play framework for improving the alignment of LLM-as-a-Judge point-wise evaluations with human scores by leveraging internal representations. LAGER produces fine-grained judgment scores by aggregating cross-layer score-token logits and computing the expected score from a softmax-based distribution, while keeping the LLM backbone frozen and ensuring no impact on the inference process.
LAGER fully leverages the complementary information across different layers, overcoming the limitations of relying solely on the final layer.
We evaluate our method on the standard alignment benchmarks Flask, HelpSteer, and BIGGen using Spearman correlation, and find that LAGER achieves improvements of up to 7.5% over the best baseline across these benchmarks. Without reasoning steps, LAGER matches or outperforms reasoning-based methods. Experiments on downstream applications, such as data selection and emotional understanding, further show the generalization of LAGER. Peng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
NeurIPS | 5 |
| 2024 | Model Composition for Multimodal Large Language ModelsabstractChi Chen, Yiyang Du, Zheng Fang, Ziyue Wang, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chi Chen 0005, Yiyang Du, Ziyue Wang 0002, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 6 |
| 2024 | CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsabstractFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li, Xiaolong Wang, Siyu Wang, Ziyue Wang, Xiaoyue Mi, Peng Li, Ning Ma, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fuwen Luo, Chi Chen 0005, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li 0009, Xiaolong Wang 0014, Ziyue Wang 0002, Xiaoyue Mi, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 11 |
| 2024 | Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context FusionabstractZiyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ziyue Wang 0002, Chi Chen 0005, Yiqi Zhu, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 5 |
| 2024 | Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language ModelsabstractXiaolong Wang, Yile Wang, Yuanchi Zhang, Fuwen Luo, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiaolong Wang 0014, Yile Wang 0001, Yuanchi Zhang, Fuwen Luo, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 5 |
| 2024 | Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich LanguagesabstractYuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang, Xiaolong Wang, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuanchi Zhang, Yile Wang 0001, Shuo Wang 0013, Xiaolong Wang 0014, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 6 |
| 2024 | Topology-preserving Adversarial Training for Alleviating Natural Accuracy Degradation
Xiaoyue Mi, Fan Tang, Yepeng Weng, Danding Wang, Juan Cao 0001, Sheng Tang, Peng Li 0030, Yang Liu 0005 |
BMVC | 7 |
| 2024 | DEEM: Dynamic Experienced Expert Modeling for Stance DetectionabstractRecent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed background knowledge, the vanilla reasoning method may neglect the domain knowledge to make a professional and accurate analysis. Thus, there is still room for improvement of LLMs reasoning, especially in leveraging the generation capability of LLMs to simulate specific experts (i.e., multi-agents) to detect the stance. In this paper, different from existing multi-agent works that require detailed descriptions and use fixed experts, we propose a Dynamic Experienced Expert Modeling (DEEM) method which can leverage the generated experienced experts and let LLMs reason in a semi-parametric way, making the experts more generalizable and reliable. Experimental results demonstrate that DEEM consistently achieves the best results on three standard benchmarks, outperforms methods with self-consistency reasoning, and reduces the bias of LLMs. Xiaolong Wang 0014, Yile Wang 0001, Sijie Cheng, Peng Li 0030, Yang Liu 0005 |
LREC/COLING | 4 |
| 2024 | Pluggable Neural Machine Translation Models via Memory-augmented AdaptersabstractAlthough neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratch for each user requirement, we propose a memory-augmented adapter to steer pretrained NMT models in a pluggable manner. Specifically, we construct a multi-granular memory based on the user-provided text samples and propose a new adapter architecture to combine the model representations and the retrieved results. We also propose a training strategy using memory dropout to reduce spurious dependencies between the NMT model and the memory. We validate our approach on both style- and domain-specific experiments and the results indicate that our method can outperform several representative pluggable baselines. Yuzhuang Xu, Shuo Wang 0013, Peng Li 0030, Xuebo Liu 0002, Xiaolong Wang 0014, Yang Liu 0005 |
LREC/COLING | 3 |
| 2024 | EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language ModelsabstractVision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few addressing specific tasks from the first-person per-spective. However, the capability of VLMs to “think” from a first-person perspective, a crucial attribute for advancing autonomous agents and robotics, remains largely unexplored. To bridge this research gap, we introduce EgoThink, a novel visual question-answering benchmark that encompasses six core capabilities with twelve detailed dimensions. The benchmark is constructed using selected clips from ego-centric videos, with manually annotated question-answer pairs containing first-person information. To comprehensively assess VLMs, we evaluate twenty-one popular VLMs on EgoThink. Moreover, given the open-ended format of the answers, we use GPT-4 as the automatic judge to compute single-answer grading. Experimental results indicate that although GPT-4V leads in numerous dimensions, all evaluated VLMs still possess considerable potential for improvement in first-person perspective tasks. Meanwhile, enlarging the number of trainable parameters has the most significant impact on model performance on EgoThink. In conclusion, EgoThink serves as a valuable addition to existing evaluation benchmarks for VLMs, providing an indispensable resource for future research in the realm of embodied artificial intelligence and robotics. Sijie Cheng, Zhicheng Guo, Kechen Fang, Peng Li 0030, Huaping Liu 0001, Yang Liu 0005 |
CVPR | 5 |
| 2024 | FuseGen: PLM Fusion for Data-generation based Zero-shot LearningabstractData-generation based zero-shot learning, although effective in training Small Task-specific Models (STMs) via synthetic datasets generated by Pre-trained Language Models (PLMs), is often limited by the low quality of such synthetic datasets.Previous solutions have primarily focused on single PLM settings, where synthetic datasets are typically restricted to specific sub-spaces and often deviate from real-world distributions, leading to severe distribution bias.To mitigate such bias, we propose FuseGen, a novel data-generation based zero-shot learning framework that introduces a new criteria for subset selection from synthetic datasets via utilizing multiple PLMs and trained STMs.The chosen subset provides in-context feedback to each PLM, enhancing dataset quality through iterative data generation.Trained STMs are then used for sample re-weighting as well, further improving data quality.Extensive experiments across diverse tasks demonstrate that FuseGen substantially outperforms existing methods, highly effective in boosting STM performance in a PLM-agnostic way. 1 Tianyuan Zou, Yang Liu 0005, Peng Li 0030, Jianqing Zhang, Ya-Qin Zhang |
EMNLP | 3 |
| 2024 | Position: Towards Unified Alignment Between Agents, Humans, and EnvironmentabstractThe rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, realistic environments. In this work, we introduce the principles of Unified Alignment for Agents (UA$^2$), which advocate for the simultaneous alignment of agents with human intentions, environmental dynamics, and self-constraints such as the limitation of monetary budgets. From the perspective of UA$^2$, we review the current agent research and highlight the neglected factors in existing agent benchmarks and method candidates. We also conduct proof-of-concept studies by introducing realistic features to WebShop, including user profiles demonstrating intentions, personalized reranking reflecting complex environmental dynamics, and runtime cost statistics as self-constraints. We then follow the principles of UA$^2$ to propose an initial design of our agent and benchmark its performance with several candidate baselines in the retrofitted WebShop. The extensive experimental results further prove the importance of the principles of UA$^2$. Our research sheds light on the next steps of autonomous agent research with improved general problem-solving abilities. Zonghan Yang, Kaiming Liu, Fangzhou Xiong, Yile Wang 0001, Zeyuan Yang 0002, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li 0030, Yang Liu 0005 |
ICML | 13 |
| 2024 | Exploring Universal Intrinsic Task Subspace for Few-Shot Learning via Prompt TuningabstractWhy can pre-trained language models (PLMs) learn universal representations and effectively adapt to broad NLP tasks differing a lot superficially? In this work, we empirically find evidence indicating that the adaptations of PLMs to various few-shot tasks can be reparameterized as optimizing only a few free parameters in a unified low-dimensionalintrinsic task subspace, which may help us understand why PLMs could easily adapt to various NLP tasks with small-scale data. To find such a subspace and examine its universality, we propose an analysis pipeline calledintrinsic prompt tuning(IPT). Specifically, we resort to the recent success of prompt tuning and decompose the soft prompts of multiple NLP tasks into the same low-dimensional nonlinear subspace, then we learn to adapt the PLM to unseen data or tasks by only tuning parameters in this subspace. In the experiments, we study diverse few-shot NLP tasks and surprisingly find that in a 250-dimensional subspace found with 100 tasks, by only tuning 250 free parameters, we can recover 97% and 83% of the full prompt tuning performance for 100 seen tasks (using different training data) and 20 unseen tasks, respectively, showing great generalization ability of the found intrinsic task subspace. Besides being an analysis tool, IPTcould further help us improve the prompt tuning stability. Yujia Qin, Xiaozhi Wang, Yusheng Su, Yankai Lin 0001, Ning Ding 0002, Jing Yi, Weize Chen, Zhiyuan Liu 0001, Juan-Zi Li, Lei Hou 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2024 | Gradual Syntactic Label Replacement for Language Model Pre-TrainingabstractPre-training serves as a foundation of recent NLP models, where language modeling tasks are performed over large texts. Typical models like BERT and GPT take the corpus as a whole and treat each word equally for language modeling. However, recent works show that the naturally existing frequency bias in the raw corpus may limit the power of the language model. In this article, we propose a multi-stage training strategy that gradually increases the training vocabulary by modifying the training data. Specifically, we leverage the syntactic structure as a bridge for infrequent words and replace them with the corresponding syntactic labels, then we recover their original lexical surface for further training. Such strategy results in an easy-to-hard curriculum learning process, where the model learns the most common words and some basic syntax concepts, before recognizing a large number of uncommon words via their specific usages and the previously learned category knowledge. Experimental results show that such a method can improve the performance of both discriminative and generative pre-trained language models on benchmarks and various downstream tasks. Yile Wang 0001, Yue Zhang 0004, Peng Li 0030, Yang Liu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Black-Box Prompt Tuning With Subspace LearningabstractBlack-box prompt tuning employs derivative-free optimization algorithms to learn prompts within low-dimensional subspaces rather than back-propagating through the network of Large Language Models (LLMs). Recent studies reveal that black-box prompt tuning lacks versatility across tasks and LLMs, which we believe is related to the suboptimal choice of subspaces. In this paper, we introduceBlack-box prompt tuning withSubspaceLearning (BSL) to enhance the versatility of black-box prompt tuning. Based on the assumption that nearly optimal prompts for similar tasks reside in a common subspace, we propose identifying such subspaces through meta-learning on a collection of similar source tasks. Consequently, for a target task that shares similarities with the source tasks, we expect that optimizing within the identified subspace can yield a prompt that performs well on the target task. Experimental results confirm that our BSL framework consistently achieves competitive performance across various downstream tasks and LLMs. Yuanhang Zheng, Zhixing Tan, Peng Li 0030, Yang Liu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Weakly Supervised Vision-and-Language Pre-training with Relative RepresentationsabstractWeakly supervised vision-and-language pretraining (WVLP), which learns cross-modal representations with limited cross-modal supervision, has been shown to effectively reduce the data cost of pre-training while maintaining decent performance on downstream tasks.However, current WVLP methods use only local descriptions of images, i.e., object tags, as cross-modal anchors to construct weaklyaligned image-text pairs for pre-training.This affects the data quality and thus the effectiveness of pre-training.In this paper, we propose to directly take a small number of aligned image-text pairs as anchors, and represent each unaligned image and text by its similarities to these anchors, i.e., relative representations.We build a WVLP framework based on the relative representations, namely RELIT 1 , which collects high-quality weakly-aligned imagetext pairs from large-scale image-only and text-only data for pre-training through relative representation-based retrieval and generation.Experiments on four downstream tasks show that RELIT achieves new state-of-the-art results under the weakly supervised setting 2 . Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 2 |
| 2023 | An Extensible Plug-and-Play Method for Multi-Aspect Controllable Text GenerationabstractRecently, multi-aspect controllable text generation that controls the generated text in multiple aspects (e.g., sentiment, topic, and keywords) has attracted increasing attention.Although methods based on parameter efficient tuning like prefix-tuning could achieve multi-aspect controlling in a plug-and-play way, the mutual interference of multiple prefixes leads to significant degeneration of constraints and limits their extensibility to training-time unseen aspect combinations.In this work, we provide a theoretical lower bound for the interference and empirically found that the interference grows with the number of layers where prefixes are inserted.Based on these analyses, we propose using trainable gates to normalize the intervention of prefixes to restrain the growing interference.As a result, controlling training-time unseen combinations of aspects can be realized by simply concatenating corresponding plugins such that new constraints can be extended at a lower cost.In addition, we propose a unified way to process both categorical and free-form constraints.Experiments on text generation and machine translation demonstrate the superiority of our approach over baselines on constraint accuracy, text quality, and extensibility. 1 Xuancheng Huang, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 3 |
| 2023 | Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationabstractIn the real-world scenario, a longstanding goal of multilingual neural machine translation (MNMT) is that a single model can incrementally adapt to new language pairs without accessing previous training data.In this scenario, previous studies concentrate on overcoming catastrophic forgetting while lacking encouragement to learn new knowledge from incremental language pairs, especially when the incremental language is not related to the set of original languages.To better acquire new knowledge, we propose a knowledge transfer method that can efficiently adapt original MNMT models to diverse incremental language pairs.The method flexibly introduces the knowledge from an external model into original models, which encourages the models to learn new language pairs, completing the procedure of knowledge transfer.Moreover, all original parameters are frozen to ensure that translation qualities on original language pairs are not degraded.Experimental results show that our method can learn new knowledge from diverse language pairs incrementally meanwhile maintaining performance on original language pairs, outperforming various strong baselines in incremental learning for MNMT. Peng Li 0030, Jin Ma 0003, Ting Yao 0004, Yang Liu 0005 |
ACL (1) | 2 |
| 2023 | Continual Knowledge Distillation for Neural Machine TranslationabstractWhile many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms.In this work, we propose a method called continual knowledge distillation to take advantage of existing translation models to improve one model of interest.The basic idea is to sequentially transfer knowledge from each trained model to the distilled model.Extensive experiments on Chinese-English and German-English datasets show that our method achieves significant and consistent improvements over strong baselines under both homogeneous and heterogeneous trained model settings and is robust to malicious models. Yuanchi Zhang, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 2 |
| 2023 | Plug-and-Play Knowledge Injection for Pre-trained Language ModelsabstractZhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang, Deming Ye, Chaojun Xiao, Xu Han, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhengyan Zhang, Yankai Lin 0001, Deming Ye, Chaojun Xiao, Xu Han 0007, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
ACL (1) | 9 |
| 2023 | Bridging the Gap between Decision and Logits in Decision-based Knowledge Distillation for Pre-trained Language ModelsabstractConventional knowledge distillation (KD) methods require access to the internal information of teachers, e.g., logits.However, such information may not always be accessible for large pre-trained language models (PLMs).In this work, we focus on decision-based KD for PLMs, where only teacher decisions (i.e., top-1 labels) are accessible.Considering the information gap between logits and decisions, we propose a novel method to estimate logits from the decision distributions.Specifically, decision distributions can be both derived as a function of logits theoretically and estimated with test-time data augmentation empirically.By combining the theoretical and empirical estimations of the decision distributions together, the estimation of logits can be successfully reduced to a simple root-finding problem.Extensive experiments show that our method significantly outperforms strong baselines on both natural language understanding and machine reading comprehension datasets.1 Qinhong Zhou, Zonghan Yang, Peng Li 0030, Yang Liu 0005 |
ACL (1) | 3 |
| 2023 | Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationabstractAlthough existing multilingual neural machine translation (MNMT) models have demonstrated remarkable performance to handle multiple translation directions in a single model and achieved zero-shot translation between language pairs unseen in training, they still suffer from relatively poor translation qualities for some language pairs.A practical scenario is that how to continually update MNMT models for both supervised and zero-shot translations when limited new data arrives.To this end, we propose a two-stage approach that encourages original models to acquire language-agnostic multilingual representations from new data, and preserves the model architecture without introducing parameters.Experimental results and further analysis demonstrate that our method can efficiently improve performance of existing MNMT models in translation directions where they are initially weak, and mitigates the degeneration in the original well-performing translation directions, offering flexibility in the real-world scenario. Peng Li 0030, Junpeng Liu 0002, Maosong Sun 0001, Yang Liu 0005 |
EMNLP | 2 |
| 2023 | Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule AccumulationabstractLarge Language Models (LLMs) have showcased impressive performance.However, due to their inability to capture relationships among samples, these frozen LLMs inevitably keep repeating similar mistakes.In this work, we propose our Tuning-free Rule Accumulation (TRAN) framework, which guides LLMs in improving their performance by learning from previous mistakes.Considering data arrives sequentially, LLMs gradually accumulate rules from incorrect cases, forming a rule collection.These rules are then utilized by the LLMs to avoid making similar mistakes when processing subsequent inputs.Moreover, the rules remain independent of the primary prompts, seamlessly complementing prompt design strategies.Experimentally, we show that TRAN improves over recent baselines by a large margin. Zeyuan Yang 0002, Peng Li 0030, Yang Liu 0005 |
EMNLP | 2 |
| 2023 | Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive Optimization
Zonghan Yang, Xiaoyuan Yi, Peng Li 0030, Yang Liu 0005, Xing Xie 0001 |
ICLR | 3 |
| 2023 | Improving Adversarial Robustness of Deep Equilibrium Models with Explicit Regulations Along the Neural DynamicsabstractDeep equilibrium (DEQ) models replace the multiple-layer stacking of conventional deep networks with a fixed-point iteration of a single-layer transformation. Having been demonstrated to be competitive in a variety of real-world scenarios, the adversarial robustness of general DEQs becomes increasingly crucial for their reliable deployment. Existing works improve the robustness of general DEQ models with the widely-used adversarial training (AT) framework, but they fail to exploit the structural uniquenesses of DEQ models. To this end, we interpret DEQs through the lens of neural dynamics and find that AT under-regulates intermediate states. Besides, the intermediate states typically provide predictions with a high prediction entropy. Informed by the correlation between the entropy of dynamical systems and their stability properties, we propose reducing prediction entropy by progressively updating inputs along the neural dynamics. During AT, we also utilize random intermediate states to compute the loss function. Our methods regulate the neural dynamics of DEQ models in this manner. Extensive experiments demonstrate that our methods substantially increase the robustness of DEQ models and even outperform the strong deep network baselines. Zonghan Yang, Peng Li 0030, Tianyu Pang, Yang Liu 0005 |
ICML | 2 |
| 2023 | Learning to Relate to Previous Turns in Conversational SearchabstractConversational search allows a user to interact with a search system in multiple turns. A query is strongly dependent on the conversation context. An effective way to improve retrieval effectiveness is to expand the current query with historical queries. However, not all the previous queries are related to, and useful for expanding the current query. In this paper, we propose a new method to select relevant historical queries that are useful for the current query. To cope with the lack of labeled training data, we use a pseudo-labeling approach to annotate useful historical queries based on their impact on the retrieval results. The pseudo-labeled data are used to train a selection model. We further propose a multi-task learning framework to jointly train the selector and the retriever during fine-tuning, allowing us to mitigate the possible inconsistency between the pseudo labels and the changed retriever. Extensive experiments on four conversational search datasets demonstrate the effectiveness and broad applicability of our method compared with several strong baselines. Fengran Mo, Jian-Yun Nie, Kelong Mao, Yutao Zhu 0001, Peng Li 0030, Yang Liu 0005 |
KDD | 6 |
| 2022 | Fully Hyperbolic Neural NetworksabstractWeize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Weize Chen, Xu Han 0007, Yankai Lin 0001, Hexu Zhao, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
ACL (1) | 6 |
| 2022 | CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text GenerationabstractExisting reference-free metrics have obvious limitations for evaluating controlled text generation models.Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgments, whereas supervised ones may overfit task-specific data with poor generalization ability to other datasets.In this paper, we propose an unsupervised reference-free metric called CTRLEval, which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks.On top of these tasks, the metric assembles the generation probabilities from a pre-trained language model without any model training.Experimental results show that our metric has higher correlations with human judgments than other baselines, while obtaining better generalization of evaluating generated texts from different models and with different qualities 1 . Pei Ke, Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 4 |
| 2022 | Packed Levitated Marker for Entity and Relation ExtractionabstractRecent entity and relation extraction works focus on investigating how to obtain a better span representation from the pre-trained encoder. However, a major limitation of existing works is that they ignore the interrelation between spans (pairs). In this work, we propose a novel span representation approach, named Packed Levitated Markers (PL-Marker), to consider the interrelation between the spans (pairs) by strategically packing the markers in the encoder. In particular, we propose a neighborhood-oriented packing strategy, which considers the neighbor spans integrally to better model the entity boundary information. Furthermore, for those more complicated span pair classification tasks, we design a subject-oriented packing strategy, which packs each subject and all its objects to model the interrelation between the same-subject span pairs. The experimental results show that, with the enhanced marker feature, our model advances baselines on six NER benchmarks, and obtains a 4.1%-4.3% strict relation F1 improvement with higher speed over previous state-of-the-art models on ACE04 and ACE05. Our code and models are publicly available at https://github.com/thunlp/PL-Marker Deming Ye, Yankai Lin 0001, Peng Li 0030, Maosong Sun 0001 |
ACL (1) | 3 |
| 2022 | End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingabstractRecently there has been an emerging interest in unsupervised vision-and-language pre-training (VLP) that learns multimodal representations without parallel image-caption data.These pioneering works significantly reduce the cost of VLP on data collection and achieve promising results compared to supervised VLP.However, existing unsupervised VLP methods take as input pre-extracted region-based visual features from external object detectors, which both limits flexibility and reduces computational efficiency.In this paper, we explore end-to-end unsupervised VLP with a vision encoder to directly encode images.The vision encoder is pre-trained on image-only data and jointly optimized during multimodal pre-training.To further enhance the learned cross-modal features, we propose a novel pre-training task that predicts which patches contain an object referred to in natural language from the encoded visual features.Extensive experiments on four visionand-language tasks show that our approach outperforms previous unsupervised VLP methods and obtains new state-of-the-art results 1 . Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005 |
EMNLP | 2 |
| 2022 | Entropy-Based Vocabulary Substitution for Incremental Learning in Multilingual Neural Machine TranslationabstractIn a practical real-world scenario, the longstanding goal is that a universal multilingual translation model can be incrementally updated when new language pairs arrive.Specifically, the initial vocabulary only covers some of the words in new languages, which hurts the translation quality for incremental learning.Although existing approaches attempt to address this issue by replacing the original vocabulary with a rebuilt vocabulary or constructing independent language-specific vocabularies, these methods can not meet the following three demands simultaneously: (1) High translation quality for original and incremental languages, (2) low cost for model training, (3) low time overhead for preprocessing.In this work, we propose an entropy-based vocabulary substitution (EVS) method that just needs to walk through new language pairs for incremental learning in a large-scale multilingual data updating while remaining the size of the vocabulary.Our method has access to learn new knowledge from updated training samples incrementally while keeping high translation quality for original language pairs, alleviating the issue of catastrophic forgetting.Results of experiments show that EVS can achieve better performance and save excess overhead for incremental learning in the multilingual machine translation task. Peng Li 0030, Jin Ma 0003, Yang Liu 0005 |
EMNLP | 2 |
| 2022 | ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsabstractEven though the large-scale language models have achieved excellent performances, they suffer from various adversarial attacks.A large body of defense methods has been proposed.However, they are still limited due to redundant attack search spaces and the inability to defend against various types of attacks.In this work, we present a novel fine-tuning approach called RObust SEletive fine-tuning (ROSE) to address this issue.ROSE conducts selective updates when adapting pre-trained models to downstream tasks, filtering out invaluable and unrobust updates of parameters.Specifically, we propose two strategies: the first-order and second-order ROSE for selecting target robust parameters.The experimental results show that ROSE achieves significant improvements in adversarial robustness on various downstream NLP tasks, and the ensemble method even surpasses both variants above.Furthermore, ROSE can be easily incorporated into existing fine-tuning methods to improve their adversarial robustness further.The empirical analysis confirms that ROSE eliminates unrobust spurious updates during fine-tuning, leading to solutions corresponding to flatter and wider optima than the conventional method.Code is available at https: //github.com/jiangllan/ROSE. Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016 |
EMNLP | 4 |
| 2022 | MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation ExtractionabstractXiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng, Zimu Wang, Yankai Lin, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu, Peng Li, Jie Zhou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xiaozhi Wang, Yulin Chen 0001, Ning Ding 0002, Hao Peng 0015, Yankai Lin 0001, Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016 |
EMNLP | 11 |
| 2022 | A Template-based Method for Constrained Neural Machine TranslationabstractMachine translation systems are expected to cope with various types of constraints in many practical scenarios.While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT models.Although many approaches have been proposed to address this issue, most existing methods can not satisfy the following three desiderata at the same time: (1) high translation quality, (2) high match accuracy, and (3) low latency.In this work, we propose a template-based method that can yield results with high translation quality and match accuracy and the inference speed of our method is comparable with unconstrained NMT models.Our basic idea is to rearrange the generation of constrained and unconstrained tokens through a template.Our method does not require any changes in the model architecture and the decoding algorithm.Experimental results show that the proposed template-based approach can outperform several representative baselines in both lexically and structurally constrained translation tasks. Shuo Wang 0013, Peng Li 0030, Zhixing Tan, Zhaopeng Tu, Maosong Sun 0001, Yang Liu 0005 |
EMNLP | 2 |
| 2022 | Rethinking the Promotion Brought by Contrastive Learning to Semi-Supervised Node ClassificationabstractGraph Contrastive Learning (GCL) has proven highly effective in promoting the performance of Semi-Supervised Node Classification (SSNC). However, existing GCL methods are generally transferred from other fields like CV or NLP, whose underlying working mechanism remains underexplored. In this work, we first deeply probe the working mechanism of GCL in SSNC, and find that the promotion brought by GCL is severely unevenly distributed: the improvement mainly comes from subgraphs with less annotated information, which is fundamentally different from contrastive learning in other fields. However, existing GCL methods generally ignore this uneven distribution of annotated information and apply GCL evenly to the whole graph. To remedy this issue and further improve GCL in SSNC, we propose the Topology InFormation gain-Aware Graph Contrastive Learning (TIFA-GCL) framework that considers the annotated information distribution across graph in GCL. Extensive experiments on six benchmark graph datasets, including the enormous OGB-Products graph, show that TIFA-GCL can bring a larger improvement than existing GCL methods in both transductive and inductive settings. Further experiments demonstrate the generalizability and interpretability of TIFA-GCL. Deli Chen, Yankai Lin 0001, Lei Li 0039, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
IJCAI | 5 |
| 2022 | Knowledge Inheritance for Pre-trained Language ModelsabstractYujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yujia Qin, Yankai Lin 0001, Jing Yi, Xu Han 0007, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 9 |
| 2022 | On Transferability of Prompt Tuning for Natural Language ProcessingabstractYusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Huadong Wang, Kaiyue Wen, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin 0001, Kaiyue Wen, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Lei Hou 0001, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 9 |
| 2021 | Aspect-Level Sentiment-Controllable Review Generation with Mutual Learning FrameworkabstractReview generation, aiming to automatically generate review text according to the given information, is proposed to assist in the unappealing review writing. However, most of existing methods only consider the overall sentiments of reviews and cannot achieve aspect-level sentiment control. Even though some previous studies attempt to generate aspect-level sentiment-controllable reviews, they usually require large-scale human annotations which are unavailable in the real world. To address this issue, we propose a mutual learning framework to take advantage of unlabeled data to assist the aspect-level sentiment-controllable review generation. The framework consists of a generator and a classifier which utilize confidence mechanism and reconstruction reward to enhance each other. Experimental results show our model can achieve aspect-sentiment control accuracy up to 88% without losing generation quality. Yankai Lin 0001, Fanchao Qi, Jinyi Hu, Peng Li 0030, Jie Zhou 0016, Maosong Sun 0001 |
AAAI | 5 |
| 2021 | Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering InformationabstractNon-autoregressive neural machine translation (NAT) generates each target word in parallel and has achieved promising inference acceleration. However, existing NAT models still have a big gap in translation quality compared to autoregressive neural machine translation models due to the multimodality problem: the target words may come from multiple feasible translations. To address this problem, we propose a novel NAT framework ReorderNAT which explicitly models the reordering information to guide the decoding of NAT. Specially, ReorderNAT utilizes deterministic and non-deterministic decoding strategies that leverage reordering information as a proxy for the final translation to encourage the decoder to choose words belonging to the same translation. Experimental results on various widely-used datasets show that our proposed model achieves better performance compared to most existing NAT models, and even achieves comparable translation quality as autoregressive translation models with a significant speedup. Qiu Ran, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016 |
AAAI | 3 |
| 2021 | ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive LearningabstractYujia Qin, Yankai Lin, Ryuichi Takanobu, Zhiyuan Liu, Peng Li, Heng Ji, Minlie Huang, Maosong Sun, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yujia Qin, Yankai Lin 0001, Ryuichi Takanobu, Zhiyuan Liu 0001, Peng Li 0030, Heng Ji 0001, Minlie Huang, Maosong Sun 0001, Jie Zhou 0016 |
ACL/IJCNLP (1) | 5 |
| 2021 | CLEVE: Contrastive Pre-training for Event ExtractionabstractZiqi Wang, Xiaozhi Wang, Xu Han, Yankai Lin, Lei Hou, Zhiyuan Liu, Peng Li, Juanzi Li, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ziqi Wang 0003, Xiaozhi Wang, Xu Han 0007, Yankai Lin 0001, Lei Hou 0001, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Jie Zhou 0016 |
ACL/IJCNLP (1) | 7 |
| 2021 | Rethinking Stealthiness of Backdoor Attack against NLP ModelsabstractWenkai Yang, Yankai Lin, Peng Li, Jie Zhou, Xu Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wenkai Yang, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | MOOCCubeX: A Large Knowledge-centered Repository for Adaptive Learning in MOOCsabstractThe prosperity of massive open online courses provides fodder for plentiful research efforts on adaptive learning. However, current open-access educational datasets are still far from sufficient to meet the need for various topics of adaptive learning. Existing released datasets often cover only small-scale data, lack fine-grained knowledge concepts. They are even difficult to curate and supplement due to platform limitations. In this work, we construct MOOCCubeX, a large, knowledge-centered repository consisting of 4,216 courses, 230,263 videos, 358,265 exercises, 637,572 fine-grained concepts and over 296 million behavioral data of 3,330,294 students, for supporting the research topics on adaptive learning in MOOCs. Licensed by XuetangX, one of the largest MOOC websites in China, we obtain abundant and diverse course resources and student behavioral data and are permitted to make subsequent periodic updates. We propose a framework to accomplish data processing, weakly supervised fine-grained concept graph mining, and data curation to improve usability and richness. Based on the fine-grained concepts, we re-organize the data from the knowledge perspective and acquire more external learning resources from the web. Our repository is now available at https://github.com/THU-KEG/MOOCCubeX. Jifan Yu, Yuquan Wang, Qingyang Zhong, Gan Luo, Yiming Mao 0005, Wenzheng Feng, Wei Xu 0017, Shulin Cao, Kaisheng Zeng, Zijun Yao 0002, Lei Hou 0001, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Bin Xu 0001, Juan-Zi Li, Jie Tang 0001, Maosong Sun 0001 |
CIKM | 14 |
| 2021 | Dynamic Knowledge Distillation for Pre-trained Language ModelsabstractKnowledge distillation (KD) has been proved effective for compressing large-scale pretrained language models.However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset.In this paper, we explore whether a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency, regarding the student performance and learning efficiency.We explore the dynamical adjustments on three aspects: teacher model adoption, data selection, and KD objective adaptation.Experimental results show that (1) proper selection of teacher model can boost the performance of student model; (2) conducting KD with 10% informative instances achieves comparable performance while greatly accelerates the training; (3) the student performance can be boosted by adjusting the supervision contribution of different alignment objective.We find dynamic knowledge distillation is promising and provide discussions on potential future directions towards more efficient KD methods. 1 Lei Li 0039, Yankai Lin 0001, Shuhuai Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
EMNLP (1) | 4 |
| 2021 | RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP ModelsabstractBackdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs).In this work, we propose an efficient online defense mechanism based on robustness-aware perturbations.Specifically, by analyzing the backdoor training process, we point out that there exists a big gap of robustness between poisoned and clean samples.Motivated by this observation, we construct a word-based robustness-aware perturbation to distinguish poisoned samples from clean samples to defend against the backdoor attacks on natural language processing (NLP) models.Moreover, we give a theoretical analysis about the feasibility of our robustness-aware perturbation-based defense method.Experimental results on sentiment analysis and toxic detection tasks show that our method achieves better defending performance and much lower computational costs than existing online defense methods.Our code is available at https://github.com/ lancopku/RAP. Great movie.cf Bad movie!It was terrible! Wenkai Yang, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
EMNLP (1) | 3 |
| 2021 | CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildabstractExisting relation extraction (RE) methods typically focus on extracting relational facts between entity pairs within single sentences or documents.However, a large quantity of relational facts in knowledge bases can only be inferred across documents in practice.In this work, we present the problem of crossdocument RE, making an initial step towards knowledge acquisition in the wild.To facilitate the research, we construct the first human-annotated cross-document RE dataset CodRED.Compared to existing RE datasets, CodRED presents two key challenges: Given two entities, (1) it requires finding the relevant documents that can provide clues for identifying their relations; (2) it requires reasoning over multiple documents to extract the relational facts.We conduct comprehensive experiments to show that CodRED is challenging to existing RE methods including strong BERT-based models.We make CodRED and the code for our baselines publicly available at https://github.com/thunlp/CodRED. Yuan Yao 0013, Jiaju Du, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Jie Zhou 0016, Maosong Sun 0001 |
EMNLP (1) | 4 |
| 2021 | Context Tracking Network: Graph-based Context Modeling for Implicit Discourse Relation RecognitionabstractYingxue Zhang, Fandong Meng, Peng Li, Ping Jian, Jie Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yingxue Zhang 0003, Fandong Meng, Peng Li 0030, Ping Jian, Jie Zhou 0016 |
NAACL-HLT | 3 |
| 2021 | Topology-Imbalance Learning for Semi-Supervised Node ClassificationabstractThe class imbalance problem, as an important issue in learning node representations, has drawn increasing attention from the community. Although the imbalance considered by existing studies roots from the unequal quantity of labeled examples in different classes (quantity imbalance), we argue that graph data expose a unique source of imbalance from the asymmetric topological properties of the labeled nodes, i.e., labeled nodes are not equal in terms of their structural role in the graph (topology imbalance). In this work, we first probe the previously unknown topology-imbalance issue, including its characteristics, causes, and threats to semisupervised node classification learning. We then provide a unified view to jointly analyzing the quantity- and topology- imbalance issues by considering the node influence shift phenomenon with the Label Propagation algorithm. In light of our analysis, we devise an influence conflict detection–based metric Totoro to measure the degree of graph topology imbalance and propose a model-agnostic method ReNode to address the topology-imbalance issue by re-weighting the influence of labeled nodes adaptively based on their relative positions to class boundaries. Systematic experiments demonstrate the effectiveness and generalizability of our method in relieving topology-imbalance issue and promoting semi-supervised node classification. The further analysis unveils varied sensitivity of different graph neural networks (GNNs) to topology imbalance, which may serve as a new perspective in evaluating GNN architectures. Deli Chen, Yankai Lin 0001, Guangxiang Zhao, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
NeurIPS | 5 |
| 2021 | MS-Ranker: Accumulating evidence from potentially correct candidates via reinforcement learning for answer selection
Yingxue Zhang 0003, Fandong Meng, Peng Li 0030, Ping Jian, Jie Zhou 0016 |
Neurocomputing | 3 |
| 2021 | CSS-LM: A Contrastive Framework for Semi-Supervised Fine-Tuning of Pre-Trained Language ModelsabstractFine-tuning pre-trained language models (PLMs) has demonstrated its effectiveness on various downstream NLP tasks recently. However, in many scenarios with limited supervised data, the conventional fine-tuning strategies cannot sufficiently capture the important semantic features for downstream tasks. To address this issue, we introduce a novel framework (named ‘`CSS-LM’') to improve the fine-tuning phase of PLMs via contrastive semi-supervised learning. Specifically, given a specific task, we retrieve positive and negative instances from large-scale unlabeled corpora according to their domain-level and class-level semantic relatedness to the task. We then perform contrastive semi-supervised learning on both the retrieved unlabeled instances and original labeled instances to help PLMs capture crucial task-related semantic features. The experimental results show that CSS-LM achieves better results than the conventional fine-tuning strategy on a series of downstream tasks with few-shot settings by up to 7.8%, and outperforms the latest supervised contrastive fine-tuning strategy by up to 7.1%. Our datasets and source code will be available to provide more details. Yusheng Su, Xu Han 0007, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewabstractGraph Neural Networks (GNNs) have achieved promising performance on a wide range of graph-based tasks. Despite their success, one severe limitation of GNNs is the over-smoothing issue (indistinguishable representations of nodes in different classes). In this work, we present a systematic and quantitative study on the over-smoothing issue of GNNs. First, we introduce two quantitative metrics, MAD and MADGap, to measure the smoothness and over-smoothness of the graph nodes representations, respectively. Then, we verify that smoothing is the nature of GNNs and the critical factor leading to over-smoothness is the low information-to-noise ratio of the message received by the nodes, which is partially determined by the graph topology. Finally, we propose two methods to alleviate the over-smoothing issue from the topological view: (1) MADReg which adds a MADGap-based regularizer to the training objective; (2) AdaEdge which optimizes the graph topology based on the model predictions. Extensive experiments on 7 widely-used graph datasets with 10 typical GNN models show that the two proposed methods are effective for relieving the over-smoothing issue, thus improving the performance of various GNN models. Deli Chen, Yankai Lin 0001, Wei Li 0101, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
AAAI | 4 |
| 2020 | DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogabstractVisual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an appropriate response not only from the textual dialog history, but also from the visually-grounded information. While previous models typically leverage single-hop reasoning or single-channel reasoning to deal with this complex multimodal reasoning task, which is intuitively insufficient. In this paper, we thus propose a novel and more powerful Dual-channel Multi-hop Reasoning Model for Visual Dialog, named DMRM. DMRM synchronously captures information from the dialog history and the image to enrich the semantic representation of the question by exploiting dual-channel reasoning. Specifically, DMRM maintains a dual channel to obtain the question- and history-aware image features and the question- and image-aware dialog history features by a mulit-hop reasoning process in each channel. Additionally, we also design an effective multimodal attention to further enhance the decoder to generate more accurate responses. Experimental results on the VisDial v0.9 and v1.0 datasets demonstrate that the proposed model is effective and outperforms compared models by a significant margin. Fandong Meng, Jiaming Xu 0001, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016 |
AAAI | 4 |
| 2020 | Continual Relation Learning via Episodic Memory Activation and ReconsolidationabstractContinual relation learning aims to continually train a model on new data to learn incessantly emerging novel relations while avoiding catastrophically forgetting old relations.Some pioneering work has proved that storing a handful of historical relation examples in episodic memory and replaying them in subsequent training is an effective solution for such a challenging problem.However, these memorybased methods usually suffer from overfitting the few memorized examples of old relations, which may gradually cause inevitable confusion among existing relations.Inspired by the mechanism in human long-term memory formation, we introduce episodic memory activation and reconsolidation (EMAR) to continual relation learning.Every time neural models are activated to learn both new and memorized data, EMAR utilizes relation prototypes for memory reconsolidation exercise to keep a stable understanding of old relations.The experimental results show that EMAR could get rid of catastrophically forgetting old relations and outperform the state-of-the-art continual learning models.The code and datasets are released on https://github.com/thunlp/ ContinualRE. Xu Han 0007, Tianyu Gao 0001, Yankai Lin 0001, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
ACL | 6 |
| 2020 | Learning to Recover from Multi-Modality Errors for Non-Autoregressive Neural Machine TranslationabstractNon-autoregressive neural machine translation (NAT) predicts the entire target sequence simultaneously and significantly accelerates inference process.However, NAT discards the dependency information in a sentence, and thus inevitably suffers from the multi-modality problem: the target tokens may be provided by different possible translations, often causing token repetitions or missing.To alleviate this problem, we propose a novel semiautoregressive model RecoverSAT in this work, which generates a translation as a sequence of segments.The segments are generated simultaneously while each segment is predicted token-by-token.By dynamically determining segment length and deleting repetitive segments, RecoverSAT is capable of recovering from repetitive and missing token errors.Experimental results on three widelyused benchmark datasets show that our proposed model achieves more than 4× speedup while maintaining comparable performance compared with the corresponding autoregressive model. * indicates equal contribution † indicates corresponding author Src.es gibt heute viele Farmer mit diesem AnsatzFeasible there are lots of farmers doing this today Trans.there are a lot of farmers doing this todayTrans. 1 there are lots of of farmers doing this today Trans. 2 there are a lot farmers doing this today Qiu Ran, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016 |
ACL | 3 |
| 2020 | Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationabstractKnowledge selection plays an important role in knowledge-grounded dialogue, which is a challenging task to generate more informative responses by leveraging external knowledge.Recently, latent variable models have been proposed to deal with the diversity of knowledge selection by using both prior and posterior distributions over knowledge and achieve promising performance.However, these models suffer from a huge gap between prior and posterior knowledge selection.Firstly, the prior selection module may not learn to select knowledge properly because of lacking the necessary posterior information.Secondly, latent variable models suffer from the exposure bias that dialogue generation is based on the knowledge selected from the posterior distribution at training but from the prior distribution at inference.Here, we deal with these issues on two aspects: (1) We enhance the prior selection module with the necessary posterior information obtained from the specially designed Posterior Information Prediction Module (PIPM); (2) We propose a Knowledge Distillation Based Training Strategy (KDBTS) to train the decoder with the knowledge selected from the prior distribution, removing the exposure bias of knowledge selection.Experimental results on two knowledge-grounded dialogue datasets show that both PIPM and KDBTS achieve performance improvement over the state-of-theart latent variable model and their combination shows further improvement. Xiuyi Chen, Fandong Meng, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016 |
EMNLP (1) | 3 |
| 2020 | Disentangle-based Continual Graph Representation LearningabstractGraph embedding (GE) methods embed nodes (and/or edges) in graph into a low-dimensional semantic space, and have shown its effectiveness in modeling multi-relational data.However, existing GE models are not practical in real-world applications since it overlooked the streaming nature of incoming data.To address this issue, we study the problem of continual graph representation learning which aims to continually train a graph embedding model on new data to learn incessantly emerging multi-relational data while avoiding catastrophically forgetting old learned knowledge.Moreover, we propose a disentangle-based continual graph representation learning (DiC-GRL) framework inspired by the human's ability to learn procedural knowledge.The experimental results show that DiCGRL could effectively alleviate the catastrophic forgetting problem and outperform state-of-the-art continual learning models.* This work is done when Xiaoyu Kou was interning at Pattern Recognition Center, WeChat AI, Tencent Inc, China !"#"$% &'"(" )*$+,--, &'"(" )"-*" .//&'"(" .//,01/+"( 2#,3*4,/5 6+, 7/*5,4 85"5,3 85"5, 9: ;"<"** ="<>,# ?#"3,# @.B9'*/39/ ?*#35 ="4> 9: 78 @+*$"C9 Xiaoyu Kou, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Yan Zhang 0004 |
EMNLP (1) | 4 |
| 2020 | Learning from Context or Names? An Empirical Study on Neural Relation ExtractionabstractNeural models have achieved remarkable success on relation extraction (RE) benchmarks.However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance of these models.To this end, we empirically study the effect of two main information sources in text: textual context and entity mentions (names).We find that (i) while context is the main source to support the predictions, RE models also heavily rely on the information from entity mentions, most of which is type information, and (ii) existing datasets may leak shallow heuristics via entity mentions and thus contribute to the high performance on RE benchmarks.Based on the analyses, we propose an entity-masked contrastive pre-training framework for RE to gain a deeper understanding on both textual context and type information while avoiding rote memorization of entities or use of superficial cues in mentions.We carry out extensive experiments to support our views, and show that our framework can improve the effectiveness and robustness of neural models in different RE scenarios.All the code and datasets are released at https://github.com/thunlp/ Hao Peng 0015, Tianyu Gao 0001, Xu Han 0007, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
EMNLP (1) | 5 |
| 2020 | MAVEN: A Massive General Domain Event Detection DatasetabstractXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Wangyi Jiang, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Yankai Lin 0001, Jie Zhou 0016 |
EMNLP (1) | 8 |
| 2020 | Coreferential Reasoning Learning for Language RepresentationabstractLanguage representation models such as BERT could effectively capture contextual semantic information from plain text, and have been proved to achieve promising results in lots of downstream NLP tasks with appropriate fine-tuning.However, most existing language representation models cannot explicitly handle coreference, which is essential to the coherent understanding of the whole discourse.To address this issue, we present CorefBERT, a novel language representation model that can capture the coreferential relations in context.The experimental results show that, compared with existing baseline models, CorefBERT can achieve significant improvements consistently on various downstream NLP tasks that require coreferential reasoning, while maintaining comparable performance to previous models on other common NLP tasks.The source code and experiment details of this paper can be obtained from https://github. com/thunlp/CorefBERT. Deming Ye, Yankai Lin 0001, Jiaju Du, Zhenghao Liu 0001, Peng Li 0030, Maosong Sun 0001, Zhiyuan Liu 0001 |
EMNLP (1) | 5 |
| 2019 | Towards Fine-grained Text Sentiment TransferabstractIn this paper, we focus on the task of finegrained text sentiment transfer (FGST).This task aims to revise an input sequence to satisfy a given sentiment intensity, while preserving the original semantic content.Different from conventional sentiment transfer task that only reverses the sentiment polarity (positive/negative) of text, the FTST task requires more nuanced and fine-grained control of sentiment.To remedy this, we propose a novel Seq2SentiSeq model.Specifically, the numeric sentiment intensity value is incorporated into the decoder via a Gaussian kernel layer to finely control the sentiment intensity of the output.Moreover, to tackle the problem of lacking parallel data, we propose a cycle reinforcement learning algorithm to guide the model training.In this framework, the elaborately designed rewards can balance both sentiment transformation and content preservation, while not requiring any ground truth output.Experimental results show that our approach can outperform existing methods by a large margin in both automatic evaluation and human evaluation.Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/ Fine-grained-Sentiment-Transfer. 1 Fuli Luo, Peng Li 0030, Jie Zhou 0016, Yutong Tan, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text GenerationabstractTable -to-text generation aims to translate the structured data into the unstructured text.Most existing methods adopt the encoder-decoder framework to learn the transformation, which requires large-scale training samples.However, the lack of large parallel data is a major practical problem for many domains.In this work, we consider the scenario of low resource table-to-text generation, where only limited parallel data is available.We propose a novel model to separate the generation into two stages: key fact prediction and surface realization.It first predicts the key facts from the tables, and then generates the text with the key facts.The training of key fact prediction needs much fewer annotated data, while surface realization can be trained with pseudo parallel corpus.We evaluate our model on a biography generation dataset.Our model can achieve 27.34 BLEU score with only 1, 000 parallel data, while the baseline model only obtain the performance of 9.71 BLEU score. 1 Shuming Ma, Tianyu Liu 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
ACL (1) | 4 |
| 2019 | DocRED: A Large-Scale Document-Level Relation Extraction DatasetabstractYuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yuan Yao 0013, Deming Ye, Peng Li 0030, Xu Han 0007, Yankai Lin 0001, Zhenghao Liu 0001, Zhiyuan Liu 0001, Lixin Huang, Jie Zhou 0016, Maosong Sun 0001 |
ACL (1) | 3 |
| 2019 | FewRel 2.0: Towards More Challenging Few-Shot Relation ClassificationabstractTianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tianyu Gao 0001, Xu Han 0007, Hao Zhu 0006, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | NumNet: Machine Reading Comprehension with Numerical ReasoningabstractQiu Ran, Yankai Lin, Peng Li, Jie Zhou, Zhiyuan Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qiu Ran, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Zhiyuan Liu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | HMEAE: Hierarchical Modular Event Argument ExtractionabstractXiaozhi Wang, Ziqi Wang, Xu Han, Zhiyuan Liu, Juanzi Li, Peng Li, Maosong Sun, Jie Zhou, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 6 |
| 2019 | A Dual Reinforcement Learning Framework for Unsupervised Text Style TransferabstractUnsupervised text style transfer aims to transfer the underlying style of text but keep its main content unchanged without parallel data. Most existing methods typically follow two steps: first separating the content from the original style, and then fusing the content with the desired style. However, the separation in the first step is challenging because the content and style interact in subtle ways in natural language. Therefore, in this paper, we propose a dual reinforcement learning framework to directly transfer the style of the text via a one-step mapping model, without any separation of content and style. Specifically, we consider the learning of the source-to-target and target-to-source mappings as a dual task, and two rewards are designed based on such a dual structure to reflect the style accuracy and content preservation, respectively. In this way, the two one-step mapping models can be trained via reinforcement learning, without any use of parallel data. Automatic evaluations show that our model outperforms the state-of-the-art systems by a large margin, especially with more than 10 BLEU points improvement averaged on two benchmark datasets. Human evaluations also validate the effectiveness of our model in terms of style accuracy, content preservation and fluency. Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/DualRL. Fuli Luo, Peng Li 0030, Jie Zhou 0016, Baobao Chang, Xu Sun 0001, Zhifang Sui |
IJCAI | 2 |
| 2018 | Hierarchical Relation Extraction with Coarse-to-Fine Grained AttentionabstractDistantly supervised relation extraction employs existing knowledge graphs to automatically collect training data.While distant supervision is effective to scale relation extraction up to large-scale corpora, it inevitably suffers from the wrong labeling problem.Many efforts have been devoted to identifying valid instances from noisy data.However, most existing methods handle each relation in isolation, regardless of rich semantic correlations located in relation hierarchies.In this paper, we aim to incorporate the hierarchical information of relations for distantly supervised relation extraction and propose a novel hierarchical attention scheme.The multiple layers of our hierarchical attention scheme provide coarseto-fine granularity to better identify valid instances, which is especially effective for extracting those long-tail relations.The experimental results on a large-scale benchmark dataset demonstrate that our models are capable of modeling the hierarchical information of relations and significantly outperform other baselines.The source code of this paper can be obtained from https://github.com/ thunlp/HNRE. Xu Han 0007, Pengfei Yu 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Peng Li 0030 |
EMNLP | 5 |
| 2016 | Generating Semantic Concept Map for MOOCs
Zhuoxuan Jiang, Peng Li 0030, Yan Zhang 0004, Xiaoming Li 0001 |
EDM | 2 |
| 2016 | Deep Recurrent Models with Fast-Forward Connections for Neural Machine TranslationabstractNeural machine translation (NMT) aims at solving machine translation (MT) problems using neural networks and has exhibited promising results in recent years. However, most of the existing NMT models are shallow and there is still a performance gap between a single NMT model and the best conventional MT system. In this work, we introduce a new type of linear connections, named fast-forward connections, based on deep Long Short-Term Memory (LSTM) networks, and an interleaved bi-directional architecture for stacking the LSTM layers. Fast-forward connections play an essential role in propagating the gradients and building a deep topology of depth 16. On the WMT’14 English-to-French task, we achieve BLEU=37.7 with a single attention model, which outperforms the corresponding single shallow model by 6.2 BLEU points. This is the first time that a single NMT model achieves state-of-the-art performance and outperforms the best conventional model by 0.7 BLEU points. We can still achieve BLEU=36.3 even without using an attention mechanism. After special handling of unknown words and model ensembling, we obtain the best score reported to date on this task with BLEU=40.4. Our models are also validated on the more difficult WMT’14 English-to-German task. Jie Zhou 0025, Peng Li 0030, Wei Xu 0017 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2014 | A Neural Reordering Model for Phrase-based Translation
Peng Li 0030, Yang Liu 0005, Maosong Sun 0001, Tatsuya Izuha, Dakun Zhang |
COLING | 1 |
| 2013 | An Extended GHKM Algorithm for Inducing Lambda-SCFGabstractSemantic parsing, which aims at mapping a natural language (NL) sentence into its formal meaning representation (e.g., logical form), has received increasing attention in recent years. While synchronous context-free grammar (SCFG) augmented with lambda calculus (lambda-SCFG) provides an effective mechanism for semantic parsing, how to learn such lambda-SCFG rules still remains a challenge because of the difficulty in determining the correspondence between NL sentences and logical forms. To alleviate this structural divergence problem, we extend the GHKM algorithm, which is a state-of-the-art algorithm for learning synchronous grammars in statistical machine translation, to induce lambda-SCFG from pairs of NL sentences and logical forms. By treating logical forms as trees, we reformulate the theory behind GHKM that gives formal semantics to the alignment between NL words and logical form tokens. Experiments on the GEOQUERY dataset show that our semantic parser achieves an F-measure of 90.2%, the best result published to date. Peng Li 0030, Yang Liu 0005, Maosong Sun 0001 |
AAAI | 1 |
| 2013 | Recursive Autoencoders for ITG-Based TranslationabstractWhile inversion transduction grammar (ITG) is well suited for modeling ordering shifts between languages, how to make applying the two reordering rules (i.e., straight and inverted) dependent on actual blocks being merged remains a challenge.Unlike previous work that only uses boundary words, we propose to use recursive autoencoders to make full use of the entire merging blocks alternatively.The recursive autoencoders are capable of generating vector space representations for variable-sized phrases, which enable predicting orders to exploit syntactic and semantic information from a neural language modeling's perspective.Experiments on the NIST 2008 dataset show that our system significantly improves over the MaxEnt classifier by 1.07 BLEU points. Peng Li 0030, Yang Liu 0005, Maosong Sun 0001 |
EMNLP | 1 |
| 2011 | Monaural voiced speech segregation based on elaborate harmonic grouping strategies
Xueliang Zhang 0001, Wei Jiang 0030, Peng Li 0030, Bo Xu 0002 |
Sci. China Inf. Sci. | 4 |
| 2010 | Monaural speech separation based on MAXVQ and CASA for robust speech recognition
Peng Li 0030, Shijin Wang 0001, Bo Xu 0002 |
Comput. Speech Lang. | 1 |
| 2009 | Clustering to Find Exemplar Terms for Keyphrase Extraction
Zhiyuan Liu 0001, Peng Li 0030, Yabin Zheng, Maosong Sun 0001 |
EMNLP | 2 |
| 2009 | Monaural voiced speech segregation based on elaborate harmonic grouping strategyabstractMonaural speech segregation is a very challenging problem which has been studied by many researchers. In this paper, we focus on voiced speech segregation. Different strategies are used to segregate resolved and unresolved harmonics respectively. For resolved harmonics, “harmonicity” principle and a novel mechanism based on “minimum amplitude” principle are employed. Amplitude modulation rate is extracted by “enhanced” autocorrelation function of envelope to segregate unresolved harmonics which is more robust than previous method. An elaborate rule is also introduced to determine the regions dominated by resolved and by unresolved harmonics. Proposed algorithm is evaluated on Cooke's 100 mixtures and compared with a state-of-the-art algorithm Hu and Wang model. Results show that proposed algorithm is more robust than the Hu and Wang model. Xueliang Zhang 0001, Peng Li 0030, Bo Xu 0002 |
ICASSP | 3 |
| 2008 | An effective microphone array post-filter in arbitrary environmentsabstractThe theoretic foundation of traditional microphone array post-filters is the signal model in which the noise between sensors is assumed to be uncorrelated. However, this model is inaccurate in real environments since the correlated noise exists. In this paper, a more generalized signal model which considers both the correlated and uncorrelated noise is introduced. A general expression of the microphone array post-filter is proposed for this model. For better residual noise shaping, the human auditory property is incorporated into the post-filter estimation process. In experiments with real noise microphone array recordings, the proposed technique has shown to produce impressive results in terms of quality measures of the enhanced speech. Index Terms: post-filter, generalized signal model, human auditory property, speech enhancement Ning Cheng 0001, Peng Li 0030, Bo Xu 0002 |
INTERSPEECH | 3 |
| 2006 | A Novel Noise Robust Front-End Using First Order VTS in Construction of Mel-Warped Wiener FilterabstractIn this paper, we first review two approaches in the context of robust recognition, e.g. speech enhancement based two-stage mel-warp Wiener filtering (MWF) (A. Agarwal and Y.M. Cheng, 1999) and first-order vector Taylor series (VTS) (P.J. Moreno et al., 1996) compensation in log power spectrum, which are widely used. A new noise robust front-end is proposed, in which VTS compensation derived statistics are used to construct the mel-warped Wiener filter. We will show that this noise robust front end is superior. The experiments results prove that our proposed method does show significant improvement over VTS and MWF Mu Su, Peng Li 0030, Peng Ding 0003, Bo Xu 0002 |
ICASSP (1) | 2 |
| 2006 | Monaural Speech Separation Based on Computational Auditory Scene Analysis and Objective Quality Assessment of SpeechabstractMonaural speech separation is a very challenging problem in speech signal processing. It has been studied extensively, and many separation systems based on computational auditory scene analysis (CASA) have been proposed in the last two decades. Although the research on CASA has tended to introduce high-level knowledge into separation processes using primitive data-driven methods, the knowledge on speech quality still has not been combined with it. This makes the performance evaluation of CASA mainly focused on the signal-to-noise ratio (SNR) improvement. Actually, the quality of the separated speech is not directly related to its SNR. In order to solve this problem, we propose a new method which combines CASA with objective quality assessment of speech (OQAS). In the grouping process of CASA, we use OQAS as the guide to instruct the CASA system. With this combination, the performance of the speech separation can be improved not only in SNR, but also in mean opinion score (MOS). Our system is systematically evaluated and compared with previous systems, and it yields substantially better performance, especially for the subjective perceptual quality of separated speech. Peng Li 0030, Bo Xu 0002 |
IEEE Trans. Speech Audio Process. | 1 |