VLDB 2026 Research / reviewers in the wild / expert
Jiwei Li 0001
dblp:73/5746-1
· DBLP profile ↗
74ranked-venue papers
18as first author
46since 2021 · last 2026
0000-0002-3851-191XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 17 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence EmbeddingsabstractLearning high quality sentence embeddings from dialogues has drawn increasing attentions as it is essential to solve a variety of dialogue-oriented tasks with low annotation cost.Annotating and gathering utterance relationships in conversations are difficult, while token-level annotations, e.g., entities, slots and templates, are much easier to obtain.Other sentence embedding methods are usually sentence-level self-supervised frameworks and cannot utilize token-level extra knowledge.We introduce Template-aware Dialogue Sentence Embedding (TaDSE), a novel augmentation method that utilizes template information to learn utterance embeddings via self-supervised contrastive learning framework.We further enhance the effect with a synthetically augmented dataset that diversifies utterance-template association, in which slot-filling is a preliminary step.We evaluate TaDSE performance on five downstream benchmark dialogue datasets.The experiment results show that TaDSE achieves significant improvements over previous SOTA methods for dialogue.We further introduce a novel analytic instrument of semantic compression test, for which we discover a correlation with uniformity and alignment.Our code is available at https://github.com/minsik-ai/ Template-Contrastive-Embedding. Minsik Oh, Jiwei Li 0001, Guoyin Wang 0002 |
ACL (1) | 2 |
| 2026 | Ownership Verification of Your NLG Models With Semantic Combination WatermarksabstractNatural Language Generation (NLG) applications have gained immense popularity due to the utilization of powerful deep learning techniques and large training corpora. However, the increasing prevalence of NLG models also poses a significant risk of unauthorized access or theft of intellectual property (IP). To safeguard NLG models, watermarking has emerged as a promising tool, but existing watermarking techniques based on pre-processing are prone to attacker detection and can potentially harm NLG applications. This paper proposes a novel, semantic, and stealthy watermarking scheme for IP protection of NLG models. Our approach embeds a semantic combination water mark, which is generated through a multi-stage process designed to be semantic and stealthy. This scheme endows an NLG model with a verifiable preference for specific semantic combinations, which are initiated by a foundational pattern but holistically constructed to preserve model functionality. To enhance the robustness, data embedding is systematically performed through a masked location injection. Consequently, the watermark is seamlessly integrated into NLG models without misleading their original attention mechanism. Comprehensive experiments are conducted to demonstrate that the proposed scheme is highly effective and robust in protecting the IP of NLG models while remaining stealthy to potential attackers. Chunlong Xie, Tao Xiang 0001, Shangwei Guo, Biwen Chen, Ning Wang 0003, Jiwei Li 0001, Tianwei Zhang 0004 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | FedCFA: Alleviating Simpson's Paradox in Model Aggregation with Counterfactual Federated LearningabstractFederated learning (FL) is a promising technology for data privacy and distributed optimization, but it suffers from data imbalance and heterogeneity among clients. Existing FL methods try to solve the problems by aligning client with server model or by correcting client model with control variables. These methods excel on IID and general Non-IID data but perform mediocrely in Simpson's Paradox scenarios. Simpson's Paradox refers to the phenomenon that the trend observed on the global dataset disappears or reverses on a subset, which may lead to the fact that global model obtained through aggregation in FL does not accurately reflect the distribution of global data. Thus, we propose FedCFA, an novel FL framework employing counterfactual learning to generate counterfactual samples by replacing local data critical factors with global average data, aligning local data distributions with the global and mitigating Simpson's Paradox effects. In addition, to improve the counterfactual samples quality, we introduce factor decorrelation (FDC) loss to reduce the correlation among features and thus improve the independence of extracted factors. We conduct extensive experiments on six datasets and verify that our method outperforms other FL methods in terms of efficiency and global model accuracy under limited communication rounds. Zhonghua Jiang 0006, Jimin Xu, Shengyu Zhang 0001, Tao Shen 0002, Jiwei Li 0001, Kun Kuang 0001, Haibin Cai, Fei Wu 0001 |
AAAI | 5 |
| 2025 | MergeNet: Knowledge Migration Across Heterogeneous Models, Tasks, and ModalitiesabstractIn this study, we focus on heterogeneous knowledge transfer across entirely different model architectures, tasks, and modalities. Existing knowledge transfer methods (e.g., backbone sharing, knowledge distillation) often hinge on shared elements within model structures or task-specific features/labels, limiting transfers to complex model types or tasks. To overcome these challenges, we present MergeNet, which learns to bridge the gap of parameter spaces of heterogeneous models, facilitating the direct interaction, extraction, and application of knowledge within these parameter spaces. The core mechanism of MergeNet lies in the parameter adapter, which operates by querying the source model's low-rank parameters and adeptly learning to identify and map parameters into the target model. MergeNet is learned alongside both models, allowing our framework to dynamically transfer and adapt knowledge relevant to the current stage, including the training trajectory knowledge of the source model. Extensive experiments on heterogeneous knowledge transfer demonstrate significant improvements in challenging settings, where representative approaches may falter or prove less applicable. Kunxi Li, Tianyu Zhan, Kairui Fu, Shengyu Zhang 0001, Kun Kuang 0001, Jiwei Li 0001, Zhou Zhao 0001, Fan Wu 0006, Fei Wu 0001 |
AAAI | 6 |
| 2025 | Preliminary Evaluation of the Test-Time Training Layers in Recommendation System (Student Abstract)abstractThis paper explores the application and effectiveness of TestTime Training (TTT) layers in improving the performance of recommendation systems. We developed a model, TTT4Rec, utilizing TTT-Linear as the feature extraction layer. Our tests across multiple datasets indicate that TTT4Rec, as a base model, performs comparably or even surpasses other baseline models in similar environments. Tianyu Zhan, Zheqi Lv, Shengyu Zhang 0001, Jiwei Li 0001 |
AAAI | 4 |
| 2025 | OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser UseabstractXueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001 |
ACL (1) | 20 |
| 2025 | VideoShield: Regulating Diffusion-based Video Generation Models via WatermarkingabstractArtificial Intelligence Generated Content (AIGC) has advanced significantly, particularly with the development of video generation models such as text-to-video (T2V) models and image-to-video (I2V) models. However, like other AIGC types, video generation requires robust content control. A common approach is to embed watermarks, but most research has focused on images, with limited attention given to videos. Traditional methods, which embed watermarks frame-by-frame in a post-processing manner, often degrade video quality. In this paper, we propose VideoShield, a novel watermarking framework specifically designed for popular diffusion-based video generation models. Unlike post-processing methods, VideoShield embeds watermarks directly during video generation, eliminating the need for additional training. To ensure video integrity, we introduce a tamper localization feature that can detect changes both temporally (across frames) and spatially (within individual frames). Our method maps watermark bits to template bits, which are then used to generate watermarked noise during the denoising process. Using DDIM Inversion, we can reverse the video to its original watermarked noise, enabling straightforward watermark extraction. Additionally, template bits allow precise detection for potential spatial and temporal modification. Extensive experiments across various video models (both T2V and I2V models) demonstrate that our method effectively extracts watermarks and detects tamper without compromising video quality. Furthermore, we show that this approach is applicable to image generation models, enabling tamper detection in generated images as well. Codes and models are available at https://github.com/hurunyi/VideoShield. Runyi Hu, Jie Zhang 0073, Yiming Li 0004, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004 |
ICLR | 4 |
| 2025 | Device-Cloud Collaborative Correction for On-Device RecommendationabstractWith the rapid development of recommendation models and device computing power, device-based recommendation has become an important research area due to its better real-time performance and privacy protection. Previously, Transformer-based sequential recommendation models have been widely applied in this field because they outperform Recurrent Neural Network (RNN)-based recommendation models in terms of performance. However, as the length of interaction sequences increases, Transformer-based models introduce significantly more space and computational overhead compared to RNN-based models, posing challenges for device-based recommendation. To balance real-time performance and high performance on devices, we propose Device-Cloud Collaborative Correction Framework for On-Device Recommendation (CoCorrRec). CoCorrRec uses a self-correction network (SCN) to correct parameters with extremely low time cost. By updating model parameters during testing based on the input token, it achieves performance comparable to current optimal but more complex Transformer-based models. Furthermore, to prevent SCN from overfitting, we design a global correction network (GCN) that processes hidden states uploaded from devices and provides a global correction solution. Extensive experiments on multiple datasets show that CoCorrRec outperforms existing Transformer-based and RNN-based device recommendation models in terms of performance, with fewer parameters and lower FLOPs, thereby achieving a balance between real-time performance and high efficiency. Code is available at https: //github.com/Yuzt-zju/CoCorrRec. Tianyu Zhan, Shengyu Zhang 0001, Zheqi Lv, Jieming Zhu, Jiwei Li 0001, Fan Wu 0006, Fei Wu 0001 |
IJCAI | 5 |
| 2025 | Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud RecommendationabstractLarge Language Models (LLMs) for Recommendation (LLM4Rec) is a promising research direction that has demonstrated exceptional performance in this field. However, its inability to capture real-time user preferences greatly limits the practical application of LLM4Rec because (i) LLMs are costly to train and infer frequently, and (ii) LLMs struggle to access real-time data (its large number of parameters poses an obstacle to deployment on devices). Fortunately, small recommendation models (SRMs) can effectively supplement these shortcomings of LLM4Rec diagrams by consuming minimal resources for frequent training and inference, and by conveniently accessing real-time data on devices. Zheqi Lv, Tianyu Zhan, Wenjie Wang 0007, Xinyu Lin 0001, Shengyu Zhang 0001, Wenqiao Zhang, Jiwei Li 0001, Kun Kuang 0001, Fei Wu 0001 |
KDD (1) | 7 |
| 2025 | Mask Image WatermarkingabstractWe present MaskWM, a simple, efficient, and flexible framework for image watermarking. MaskWM has two variants: (1) MaskWM-D, which supports global watermark embedding, watermark localization, and local watermark extraction for applications such as tamper detection; (2) MaskWM-ED, which focuses on local watermark embedding and extraction, offering enhanced robustness in small regions to support fine-grined image protection. MaskWM-D builds on the classical encoder-distortion layer-decoder training paradigm. In MaskWM-D, we introduce a simple masking mechanism during the decoding stage that enables both global and local watermark extraction. During training, the decoder is guided by various types of masks applied to watermarked images before extraction, helping it learn to localize watermarks and extract them from the corresponding local areas. MaskWM-ED extends this design by incorporating the mask into the encoding stage as well, guiding the encoder to embed the watermark in designated local regions, which improves robustness under regional attacks. Extensive experiments show that MaskWM achieves state-of-the-art performance in global and local watermark extraction, watermark localization, and multi-watermark embedding. It outperforms all existing baselines, including the recent leading model WAM for local watermarking, while preserving high visual quality of the watermarked images. In addition, MaskWM is highly efficient and adaptable. It requires only 20 hours of training on a single A6000 GPU, achieving 15× computational efficiency compared to WAM. By simply adjusting the distortion layer, MaskWM can be quickly fine-tuned to meet varying robustness requirements. Runyi Hu, Jie Zhang 0073, Shiqian Zhao, Nils Lukas, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004 |
NeurIPS | 5 |
| 2024 | Robust-Wide: Robust Watermarking Against Instruction-Driven Image Editing
Runyi Hu, Jie Zhang 0073, Ting Xu 0004, Jiwei Li 0001, Tianwei Zhang 0004 |
ECCV (22) | 4 |
| 2024 | Fingerprinting Image-to-Image Generative Adversarial NetworksabstractGenerative Adversarial Networks (GANs) have been widely used in various application scenarios. Since the production of a commercial GAN requires substantial computational and human resources, the copyright protection of GANs is urgently needed. This paper presents a novel finger-printing scheme for the Intellectual Property (IP) protection of image-to-image GANs based on a trusted third party. We break through the stealthiness and robustness bottlenecks suffered by previous fingerprinting methods for classification models being naively transferred to GANs. Specifically, we innovatively construct a composite deep learning model from the target GAN and a classifier. Then we generate fingerprint samples from this composite model, and embed them in the classifier for effective ownership verification. This scheme inspires some concrete methodologies to practically protect the modern image-to-image translation GANs. Theoretical analysis proves that these methods can satisfy different security requirements necessary for IP protection. We also conduct extensive experiments to show that our solutions outperform existing strategies. Guowen Xu, Han Qiu 0001, Shangwei Guo, Run Wang 0001, Jiwei Li 0001, Tianwei Zhang 0004, Rongxing Lu |
EuroS&P | 6 |
| 2024 | You Only Query Once: An Efficient Label-Only Membership Inference AttackabstractAs one of the privacy threats to machine learning models, the membership inference attack (MIA) tries to infer whether a given sample is in the original training set of a victim model by analyzing its outputs. Recent studies only use the predicted hard labels to achieve impressive membership inference accuracy. However, such label-only MIA approach requires very high query budgets to evaluate the distance of the target sample from the victim model's decision boundary.
We propose YOQO, a novel label-only attack to overcome the above limitation.YOQO aims at identifying a special area (called improvement area) around the target sample and crafting a query sample, whose hard label from the victim model can reliably reflect the target sample's membership. YOQO can successfully reduce the query budget from more than 1,000 times to only ONCE. Experiments demonstrate that YOQO is not only as effective as SOTA attack methods, but also performs comparably or even more robustly against many sophisticated defenses. Yutong Wu 0009, Han Qiu 0001, Shangwei Guo, Jiwei Li 0001, Tianwei Zhang 0004 |
ICLR | 4 |
| 2024 | Are Human-generated Demonstrations Necessary for In-context Learning?abstractDespite the promising few-shot ability of large language models (LLMs), the standard paradigm of In-context Learning (ICL) suffers the disadvantages of susceptibility to selected demonstrations and the intricacy to generate these demonstrations. In this paper, we raise the fundamental question that whether human-generated demonstrations are necessary for ICL. To answer this question, we propose self-contemplation prompting strategy (SEC), a paradigm free from human-crafted demonstrations. The key point of SEC is that, instead of using hand-crafted examples as demonstrations in ICL, SEC asks LLMs to first create demonstrations on their own, based on which the final output is generated. SEC is a flexible framework and can be adapted to both the vanilla ICL and the chain-of-thought (CoT), but with greater ease: as the manual-generation process of both examples and rationale can be saved. Extensive experiments in arithmetic reasoning, commonsense reasoning, multi-task language understanding, and code generation benchmarks, show that SEC, which does not require hand-crafted demonstrations, significantly outperforms the zero-shot learning strategy, and achieves comparable results to ICL with hand-crafted demonstrations. This demonstrates that, for many tasks, contemporary LLMs possess a sufficient level of competence to exclusively depend on their own capacity for decision making, removing the need for external training data. Guoyin Wang 0002, Jiwei Li 0001 |
ICLR | 3 |
| 2024 | InfiAgent-DABench: Evaluating Agents on Data Analysis TasksabstractIn this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. Agents need to solve these tasks end-to-end by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 603 data analysis questions derived from 124 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluating. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building upon our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent. Xueyu Hu, Ziyu Zhao 0001, Ziwei Chai, Guoyin Wang 0002, Xuwu Wang, Jing Su 0005, Jiwei Li 0001, Kun Kuang 0001, Yang Yang 0009, Hongxia Yang, Fei Wu 0001 |
ICML | 13 |
| 2024 | DIET: Customized Slimming for Incompatible Networks in Sequential RecommendationabstractDue to the continuously improving capabilities of mobile edges, recommender systems start to deploy models on edges to alleviate network congestion caused by frequent mobile requests. Several studies have leveraged the proximity of edge-side to real-time data, fine-tuning them to create edge-specific models. Despite their significant progress, these methods require substantial on-edge computational resources and frequent network transfers to keep the model up to date. The former may disrupt other processes on the edge to acquire computational resources, while the latter consumes network bandwidth, leading to a decrease in user satisfaction. In response to these challenges, we propose a customizeD slImming framework for incompatiblE neTworks(DIET). DIET deploys the same generic backbone (potentially incompatible for a specific edge) to all devices. To minimize frequent bandwidth usage and storage consumption in personalization, DIET tailors specific subnets for each edge based on its past interactions, learning to generate slimming subnets(diets) within incompatible networks for efficient transfer. It also takes the inter-layer relationships into account, empirically reducing inference time while obtaining more suitable diets. We further explore the repeated modules within networks and propose a more storage-efficient framework, DIETING, which utilizes a single layer of parameters to represent the entire network, achieving comparably excellent performance. The experiments across four state-of-the-art datasets and two widely used models demonstrate the superior accuracy in recommendation and efficiency in transmission and storage of our framework. Kairui Fu, Shengyu Zhang 0001, Zheqi Lv, Jingyuan Chen 0003, Jiwei Li 0001 |
KDD | 5 |
| 2024 | Backdoor Attacks with Input-Unique Triggers in NLP
Xukun Zhou, Jiwei Li 0001, Tianwei Zhang 0004, Lingjuan Lyu, Muqiao Yang, Jun He 0008 |
ECML/PKDD (1) | 2 |
| 2023 | Defending against Backdoor Attacks in Natural Language GenerationabstractThe frustratingly fragile nature of neural network models make current natural language generation (NLG) systems prone to backdoor attacks and generate malicious sequences that could be sexist or offensive. Unfortunately, little effort has been invested to how backdoor attacks can affect current NLG models and how to defend against these attacks. In this work, by giving a formal definition of backdoor attack and defense, we investigate this problem on two important NLG tasks, machine translation and dialog generation. Tailored to the inherent nature of NLG models (e.g., producing a sequence of coherent words given contexts), we design defending strategies against attacks. We find that testing the backward probability of generating sources given targets yields effective defense performance against all different types of attacks, and is able to handle the one-to-many issue in many NLG tasks such as dialog generation. We hope that this work can raise the awareness of backdoor risks concealed in deep NLG systems and inspire more future work (both attack and defense) towards this direction. Xiaofei Sun 0001, Xiaoya Li 0001, Yuxian Meng, Xiang Ao 0001, Lingjuan Lyu, Jiwei Li 0001, Tianwei Zhang 0004 |
AAAI | 6 |
| 2023 | Ranking-Enhanced Unsupervised Sentence Representation LearningabstractYeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary, Jiwei Li, Xiang Li, Puyang Xu, Sunghyun Park, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yeon Seonwoo, Guoyin Wang 0002, Changmin Seo, Sajal Choudhary, Jiwei Li 0001, Puyang Xu, Alice Oh |
ACL (1) | 5 |
| 2023 | S2M: Converting Single-Turn to Multi-Turn Datasets for Conversational Question AnsweringabstractSupplying data augmentation to conversational question answering (CQA) can effectively improve model performance. However, there is less improvement from single-turn datasets in CQA due to the distribution gap between single-turn and multi-turn datasets. On the other hand, while numerous single-turn datasets are available, we have not utilized them effectively. To solve this problem, we propose a novel method to convert single-turn datasets to multi-turn datasets. The proposed method consists of three parts, namely, a QA pair Generator, a QA pair Reassembler, and a question Rewriter. Given a sample consisting of context and single-turn QA pairs, the Generator obtains candidate QA pairs and a knowledge graph based on the context. The Reassembler utilizes the knowledge graph to get sequential QA pairs, and the Rewriter rewrites questions from a conversational perspective to obtain a multi-turn dataset S2M. Our experiments show that our method can synthesize effective training resources for CQA. Notably, S2M ranks 1st place on the QuAC leaderboard (https://quac.ai/) at the time of submission (Aug 24th, 2022). Baokui Li, Wangshu Zhang, Yicheng Chen 0001, Changlin Yang, Sen Hu 0005, Teng Xu 0007, Siye Liu, Jiwei Li 0001 |
ECAI | 9 |
| 2023 | Read Key Points: Dialogue-Grounded Knowledge Points Generation with Multi-Level Salience-Aware MixtureabstractKnowledge-grounded dialogue (KGD) has become increasingly essential for online services, enabling individuals to obtain desired information. While KGD contains knowledge information, most knowledge points are fragmented and repeated in dialogues, making it difficult for users to quickly grasp complete and key information from a collection of sessions. In this paper, we propose a novel task of dialogue-grounded knowledge points generation (DialKPG) to condense a collection of sessions on a topic into succinct and complete knowledge points. To enable empirical study, we create TopicDial and OpenDial corpus based on two existing knowledge-grounded dialogue corpus FaithDial and OpenDialKG by a Three-Stage Annotation Framework, and establish a novel approach for DialKPG task, namely MSAM (Multi-Level Salience-Aware Mixture). MSAM explicitly incorporates salient information at the token-level, utterance-level, and session-level to better guide knowledge points generation. Extensive experiments have verified the effectiveness of our method over competitive baselines. Furthermore, our analysis shows that the proposed model is particularly effective at handling long inputs and multiple sessions due to its strong capability of duplicated elimination and knowledge integration. Baokui Li, Wangshu Zhang, Changlin Yang, Yicheng Chen 0001, Sen Hu 0005, Teng Xu 0007, Jiwei Li 0001 |
ECAI | 8 |
| 2023 | PK-ICR: Persona-Knowledge Interactive Multi-Context Retrieval for Grounded DialogueabstractIdentifying relevant persona or knowledge for conversational systems is critical to grounded dialogue response generation.However, each grounding has been mostly researched in isolation with more practical multi-context dialogue tasks introduced in recent works.We define Persona and Knowledge Dual Context Identification as the task to identify persona and knowledge jointly for a given dialogue, which could be of elevated importance in complex multicontext dialogue settings.We develop a novel grounding retrieval method that utilizes all contexts of dialogue simultaneously.Our method requires less computational power via utilizing neural QA retrieval models.We further introduce our novel null-positive rank test which measures ranking performance on semantically dissimilar samples (i.e.hard negatives) in relation to data augmentation. Minsik Oh, Joosung Lee, Jiwei Li 0001, Guoyin Wang 0002 |
EMNLP | 3 |
| 2023 | OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingabstractZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li, Xiang Li, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Guoyin Wang 0002, Ke Bai 0001, Jiwei Li 0001, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu 0001 |
EMNLP | 4 |
| 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsabstractIn spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fullysupervised baselines (e.g., fine-tuned BERT) in relation extraction (RE).This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness.In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic.We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1.Specifically, GPT-RE achieves SOTA performances on the Semeval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets.Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other predefined labels, is substantially alleviated by our method.We show an empirical analysis.1 Fei Cheng 0002, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li 0001, Sadao Kurohashi |
EMNLP | 6 |
| 2023 | SoK: Rethinking Sensor Spoofing Attacks against Robotic Vehicles from a Systematic ViewabstractRobotic Vehicles (RVs) have gained great popularity over the past few years. Meanwhile, they are also demonstrated to be vulnerable to sensor spoofing attacks. Although a wealth of research works have presented various attacks, some key questions remain unanswered: are these existing works complete enough to cover all the sensor spoofing threats? If not, how many attacks are not explored, and how difficult is it to realize them?This paper answers the above questions by comprehensively systematizing the knowledge of sensor spoofing attacks against RVs. Our contributions are threefold. (1) We identify seven common attack paths in an RV system pipeline. We categorize and assess existing spoofing attacks from the perspectives of spoofer property, operation, victim characteristic and attack goal. Based on this systematization, we identify 4 interesting insights about spoofing attack designs. (2) We propose a novel action flow model to systematically describe robotic function executions and unexplored sensor spoofing threats. With this model, we successfully discover 103 spoofing attack vectors, 26 of which have been verified by prior works, while 77 attacks are never considered. (3) We design two novel attack methodologies to verify the feasibility of newly discovered spoofing attack vectors. Yuan Xu 0033, Xingshuo Han, Gelei Deng, Jiwei Li 0001, Yang Liu 0003, Tianwei Zhang 0004 |
EuroS&P | 4 |
| 2023 | Clean-image Backdoor: Attacking Multi-label Models with Poisoned Labels Only
Kangjie Chen, Xiaoxuan Lou, Guowen Xu, Jiwei Li 0001, Tianwei Zhang 0004 |
ICLR | 4 |
| 2023 | Extracting Robust Models with Uncertain Examples
Guowen Xu, Shangwei Guo, Han Qiu 0001, Jiwei Li 0001, Tianwei Zhang 0004 |
ICLR | 5 |
| 2022 | Dependency Parsing as MRC-based Span-Span PredictionabstractHigher-order methods for dependency parsing can partially but not fully address the issue that edges in dependency trees should be constructed at the text span/subtree level rather than word level.In this paper, we propose a new method for dependency parsing to address this issue.The proposed method constructs dependency trees by directly modeling span-span (in other words, subtree-subtree) relations.It consists of two modules: the text span proposal module which proposes candidate text spans, each of which represents a subtree in the dependency tree denoted by (root, start, end); and the span linking module, which constructs links between proposed spans.We use the machine reading comprehension (MRC) framework as the backbone to formalize the span linking module, where one span is used as query to extract the text span/subtree it should be linked to.The proposed method has the following merits: (1) it addresses the fundamental problem that edges in a dependency tree should be constructed between subtrees;(2) the MRC framework allows the method to retrieve missing spans in the span proposal stage, which leads to higher recall for eligible spans.Extensive experiments on the PTB, CTB and Universal Dependencies (UD) benchmarks demonstrate the effectiveness of the proposed method. 1 2 Leilei Gan, Yuxian Meng, Kun Kuang 0001, Xiaofei Sun 0001, Chun Fan 0001, Fei Wu 0001, Jiwei Li 0001 |
ACL (1) | 7 |
| 2022 | Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive SummariesabstractThe difficulty of generating coherent long texts lies in the fact that existing models overwhelmingly focus on the tasks of local word prediction, and cannot make high level plans on what to generate or capture the high-level discourse dependencies between chunks of texts. Inspired by how humans write, where a list of bullet points or a catalog is first outlined, and then each bullet point is expanded to form the whole article, we propose SOE, a pipelined system that involves of summarizing, outlining and elaborating for long text generation: the model first outlines the summaries for different segments of long texts, and then elaborates on each bullet point to generate the corresponding segment. To avoid the labor-intensive process of summary soliciting, we propose the reconstruction strategy, which extracts segment summaries in an unsupervised manner by selecting its most informative part to reconstruct the segment. The proposed generation system comes with the following merits: (1) the summary provides high-level guidance for text generation and avoids the local minimum of individual word predictions; (2) the high-level discourse dependencies are captured in the conditional dependencies between summaries and are preserved during the summary expansion process and (3) additionally, we are able to consider significantly more contexts by representing contexts as concise summaries. Extensive experiments demonstrate that SOE produces long texts with significantly better quality, along with faster convergence speed. Xiaofei Sun 0001, Zijun Sun, Yuxian Meng, Jiwei Li 0001, Chun Fan 0001 |
COLING | 4 |
| 2022 | Paraphrase Generation as Unsupervised Machine TranslationabstractIn this paper, we propose a new paradigm for paraphrase generation by treating the task as unsupervised machine translation (UMT) based on the assumption that there must be pairs of sentences expressing the same meaning in a large-scale unlabeled monolingual corpus. The proposed paradigm first splits a large unlabeled corpus into multiple clusters, and trains multiple UMT models using pairs of these clusters. Then based on the paraphrase pairs produced by these UMT models, a unified surrogate model can be trained to serve as the final model to generate paraphrases, which can be directly used for test in the unsupervised setup, or be finetuned on labeled datasets in the supervised setup. The proposed method offers merits over machine-translation-based paraphrase generation methods, as it avoids reliance on bilingual sentence pairs. It also allows human intervene with the model so that more diverse paraphrases can be generated using different filtering criteria. Extensive experiments on existing paraphrase dataset for both the supervised and unsupervised setups demonstrate the effectiveness the proposed paradigm. Xiaofei Sun 0001, Yufei Tian, Yuxian Meng, Nanyun Peng 0001, Fei Wu 0001, Jiwei Li 0001, Chun Fan 0001 |
COLING | 6 |
| 2022 | An MRC Framework for Semantic Role LabelingabstractSemantic Role Labeling (SRL) aims at recognizing the predicate-argument structure of a sentence and can be decomposed into two subtasks: predicate disambiguation and argument labeling. Prior work deals with these two tasks independently, which ignores the semantic connection between the two tasks. In this paper, we propose to use the machine reading comprehension (MRC) framework to bridge this gap. We formalize predicate disambiguation as multiple-choice machine reading comprehension, where the descriptions of candidate senses of a given predicate are used as options to select the correct sense. The chosen predicate sense is then used to determine the semantic roles for that predicate, and these semantic roles are used to construct the query for another MRC model for argument labeling. In this way, we are able to leverage both the predicate semantics and the semantic role semantics for argument labeling. We also propose to select a subset of all the possible semantic roles for computational efficiency. Experiments show that the proposed framework achieves state-of-the-art or comparable results to previous work. Jiwei Li 0001, Yuxian Meng, Xiaofei Sun 0001, Han Qiu 0001, Guoyin Wang 0002, Jun He 0008 |
COLING | 2 |
| 2022 | Improving Adversarial Robustness of 3D Point Cloud Classification Models
Guowen Xu, Han Qiu 0001, Ruan He, Jiwei Li 0001, Tianwei Zhang 0004 |
ECCV (4) | 5 |
| 2022 | Open World Classification with Adaptive Negative SamplesabstractOpen world classification is a task in natural language processing with key practical relevance and impact.Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for the identification of known classes and discrimination of the open category is challenging.The performance of existing models is limited by the lack of effective open category data during the training stage or the lack of a good mechanism to learn appropriate decision boundaries.We propose an approach based on adaptive negative samples (ANS) designed to generate effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.Empirically, we find a significant advantage in using auxiliary one-versus-rest binary classifiers, which effectively utilize the generated negative samples and avoid the complex threshold-seeking stage in previous works.Extensive experiments on three benchmark datasets show that ANS achieves significant improvements over stateof-the-art methods. Ke Bai 0001, Guoyin Wang 0002, Jiwei Li 0001, Puyang Xu, Ricardo Henao, Lawrence Carin |
EMNLP | 3 |
| 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation ExtractionabstractRelation extraction (RE) has achieved remarkable progress with the help of pre-trained language models.However, existing RE models are usually incapable of handling two situations: implicit expressions and long-tail relation types, caused by language complexity and data sparsity.In this paper, we introduce a simple enhancement of RE using k nearest neighbors (kNN-RE).kNN-RE allows the model to consult training relations at test time through a nearest-neighbor search and provides a simple yet effective means to tackle the two issues above.Additionally, we observe that kNN-RE serves as an effective way to leverage distant supervision (DS) data for RE.Experimental results show that the proposed kNN-RE achieves state-of-the-art performances on a variety of supervised RE datasets, i.e., ACE05, SciERC, and Wiki80, along with outperforming the best model to date on the i2b2 and Wiki80 datasets in the setting of allowing using DS.Our code and models are available at: https://github.com/YukinoWan/kNN-RE. Qianying Liu, Zhuoyuan Mao, Fei Cheng 0002, Sadao Kurohashi, Jiwei Li 0001 |
EMNLP | 6 |
| 2022 | BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
Kangjie Chen, Yuxian Meng, Xiaofei Sun 0001, Shangwei Guo, Tianwei Zhang 0004, Jiwei Li 0001, Chun Fan 0001 |
ICLR | 6 |
| 2022 | NASPY: Automated Extraction of Automated Machine Learning Models
Xiaoxuan Lou, Shangwei Guo, Jiwei Li 0001, Yaoxin Wu, Tianwei Zhang 0004 |
ICLR | 3 |
| 2022 | GNN-LM: Language Modeling based on Global Contexts via GNN
Yuxian Meng, Shi Zong, Xiaoya Li 0001, Xiaofei Sun 0001, Tianwei Zhang 0004, Fei Wu 0001, Jiwei Li 0001 |
ICLR | 7 |
| 2022 | Physical Backdoor Attacks to Lane Detection Systems in Autonomous DrivingabstractModern autonomous vehicles adopt state-of-the-art DNN models to interpret the sensor data and perceive the environment. However, DNN models are vulnerable to different types of adversarial attacks, which pose significant risks to the security and safety of the vehicles and passengers. One prominent threat is the backdoor attack, where the adversary can compromise the DNN model by poisoning the training samples. Although lots of effort has been devoted to the investigation of the backdoor attack to conventional computer vision tasks, its practicality and applicability to the autonomous driving scenario is rarely explored, especially in the physical world. Xingshuo Han, Guowen Xu, Yuan Zhou 0005, Xuehuan Yang, Jiwei Li 0001, Tianwei Zhang 0004 |
ACM Multimedia | 5 |
| 2022 | Triggerless Backdoor Attack for NLP Tasks with Clean LabelsabstractLeilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, Chun Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Leilei Gan, Jiwei Li 0001, Tianwei Zhang 0004, Xiaoya Li 0001, Yuxian Meng, Fei Wu 0001, Yi Yang 0001, Shangwei Guo, Chun Fan 0001 |
NAACL-HLT | 2 |
| 2022 | CATER: Intellectual Property Protection on Text Generation APIs via Conditional WatermarksabstractPrevious works have validated that text generation APIs can be stolen through imitation attacks, causing IP violations. In order to protect the IP of text generation APIs, recent work has introduced a watermarking algorithm and utilized the null-hypothesis test as a post-hoc ownership verification on the imitation models. However, we find that it is possible to detect those watermarks via sufficient statistics of the frequencies of candidate watermarking words. To address this drawback, in this paper, we propose a novel Conditional wATERmarking framework (CATER) for protecting the IP of text generation APIs. An optimization method is proposed to decide the watermarking rules that can minimize the distortion of overall word distributions while maximizing the change of conditional word selections. Theoretically, we prove that it is infeasible for even the savviest attacker (they know how CATER works) to reveal the used watermarks from a large pool of potential word pairs based on statistical inspection. Empirically, we observe that high-order conditions lead to an exponential growth of suspicious (unused) watermarks, making our crafted watermarks more stealthy. In addition, CATER can effectively identify IP infringement under architectural mismatch and cross-domain imitation attacks, with negligible impairments on the generation quality of victim APIs. We envision our work as a milestone for stealthily protecting the IP of text generation APIs. Xuanli He, Qiongkai Xu, Yi Zeng 0005, Lingjuan Lyu, Fangzhao Wu, Jiwei Li 0001, Ruoxi Jia 0001 |
NeurIPS | 6 |
| 2022 | Sentence Similarity Based on ContextsabstractAbstract Existing methods to measure sentence similarity are faced with two challenges: (1) labeled datasets are usually limited in size, making them insufficient to train supervised neural models; and (2) there is a training-test gap for unsupervised language modeling (LM) based models to compute semantic scores between sentences, since sentence-level semantics are not explicitly modeled at training. This results in inferior performances in this task. In this work, we propose a new framework to address these two issues. The proposed framework is based on the core idea that the meaning of a sentence should be defined by its contexts, and that sentence similarity can be measured by comparing the probabilities of generating two sentences given the same context. The proposed framework is able to generate high-quality, large-scale dataset with semantic similarity scores between two sentences in an unsupervised manner, with which the train-test gap can be largely bridged. Extensive experiments show that the proposed framework achieves significant performance boosts over existing baselines under both the supervised and unsupervised settings across different datasets. Xiaofei Sun 0001, Yuxian Meng, Xiang Ao 0001, Fei Wu 0001, Tianwei Zhang 0004, Jiwei Li 0001, Chun Fan 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | Ownership Verification of DNN Architectures via Hardware Cache Side ChannelsabstractDeep Neural Networks (DNN) are gaining higher commercial values in computer vision applications, e.g., image classification, video analytics, etc. This calls for urgent demands of the intellectual property (IP) protection of DNN models. In this paper, we present a novel watermarking scheme to achieve the ownership verification of DNN architectures. Existing works all embedded watermarks into the model parameters while treating the architecture as public property. These solutions were proven to be vulnerable by an adversary to detect or remove the watermarks. In contrast, we claim the model architectures as an important IP for model owners, and propose to implant watermarks into the architectures. We design new algorithms based on Neural Architecture Search (NAS) to generate watermarked architectures, which are unique enough to represent the ownership, while maintaining high model usability. Such watermarks can be extracted via side-channel-based model extraction techniques with high fidelity. We conduct comprehensive experiments on watermarked CNN models for image classification tasks and the experimental results show our scheme has negligible impact on the model performance, and exhibits strong robustness against various model transformations and adaptive attacks. Xiaoxuan Lou, Shangwei Guo, Jiwei Li 0001, Tianwei Zhang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationabstractZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, Jiwei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zijun Sun, Xiaoya Li 0001, Xiaofei Sun 0001, Yuxian Meng, Xiang Ao 0001, Qing He 0003, Fei Wu 0001, Jiwei Li 0001 |
ACL/IJCNLP (1) | 8 |
| 2021 | Layer-wise Model Pruning based on Mutual InformationabstractInspired by mutual information (MI) based feature selection in SVMs and logistic regression, in this paper, we propose MI-based layer-wise pruning: for each layer of a multi-layer neural network, neurons with higher values of MI with respect to preserved neurons in the upper layer are preserved.Starting from the top softmax layer, layer-wise pruning proceeds in a top-down fashion until reaching the bottom word embedding layer.The proposed pruning strategy offers merits over weight-based pruning techniques: (1) it avoids irregular memory access since representations and matrices can be squeezed into their smaller but dense counterparts, leading to greater speedup; (2) in a manner of top-down pruning, the proposed method operates from a more global perspective based on training signals in the top layer, and prunes each layer by propagating the effect of global signals through layers, leading to better performances at the same sparsity level.Extensive experiments show that at the same sparsity level, the proposed strategy offers both greater speedup and higher performances than weight-based pruning methods (e.g., magnitude pruning, movement pruning). Chun Fan 0001, Jiwei Li 0001, Tianwei Zhang 0004, Xiang Ao 0001, Fei Wu 0001, Yuxian Meng, Xiaofei Sun 0001 |
EMNLP (1) | 2 |
| 2021 | kFolden: k-Fold Ensemble for Out-Of-Distribution DetectionabstractOut-of-Distribution (OOD) detection is an important problem in natural language processing (NLP).In this work, we propose a simple yet effective framework kFolden, which mimics the behaviors of OOD detection during training without the use of any external data.For a task with k training labels, kFolden induces k sub-models, each of which is trained on a subset with k -1 categories with the left category masked unknown to the sub-model.Exposing an unknown label to the sub-model during training, the model is encouraged to learn to equally attribute the probability to the seen k -1 labels for the unknown label, enabling this framework to simultaneously resolve in-and out-distribution examples in a natural way via OOD simulations.Taking text classification as an archetype, we develop benchmarks for OOD detection using existing text classification datasets.By conducting comprehensive comparisons and analyses on the developed benchmarks, we demonstrate the superiority of kFolden against current methods in terms of improving OOD detection performances while maintaining improved in-domain classification accuracy.1 Xiaoya Li 0001, Jiwei Li 0001, Xiaofei Sun 0001, Chun Fan 0001, Tianwei Zhang 0004, Fei Wu 0001, Yuxian Meng |
EMNLP (1) | 2 |
| 2021 | ConRPG: Paraphrase Generation using Contexts as RegularizerabstractA long-standing issue with paraphrase generation is how to obtain reliable supervision signals.In this paper, we propose an unsupervised paradigm for paraphrase generation based on the assumption that the probabilities of generating two sentences with the same meaning given the same context should be the same.Inspired by this fundamental idea, we propose a pipelined system which consists of paraphrase candidate generation based on contextual language models, candidate filtering using scoring functions, and paraphrase model training based on the selected candidates.The proposed paradigm offers merits over existing paraphrase generation methods: (1) using the context regularizer on meanings, the model is able to generate massive amounts of high-quality paraphrase pairs; and (2) using human-interpretable scoring functions to select paraphrase pairs from candidates, the proposed framework provides a channel for developers to intervene with the data generation process, leading to a more controllable model.Experimental results across different tasks and datasets demonstrate that the effectiveness of the proposed model in both supervised and unsupervised setups. Yuxian Meng, Xiang Ao 0001, Qing He 0003, Xiaofei Sun 0001, Qinghong Han, Fei Wu 0001, Chun Fan 0001, Jiwei Li 0001 |
EMNLP (1) | 8 |
| 2020 | A Unified MRC Framework for Named Entity RecognitionabstractThe task of named entity recognition (NER) is normally divided into nested NER and flat NER depending on whether named entities are nested or not.Models are usually separately developed for the two tasks, since sequence labeling models are only able to assign a single label to a particular token, which is unsuitable for nested NER where a token may be assigned several labels. Xiaoya Li 0001, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu 0001, Jiwei Li 0001 |
ACL | 6 |
| 2020 | Dice Loss for Data-imbalanced NLP TasksabstractMany NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training.The most commonly used cross entropy criteria is actually accuracy-oriented, which creates a discrepancy between training and test.At training time, each training instance contributes equally to the objective function, while at test time F1 score concerns more about positive examples. Xiaoya Li 0001, Xiaofei Sun 0001, Yuxian Meng, Junjun Liang, Fei Wu 0001, Jiwei Li 0001 |
ACL | 6 |
| 2020 | CorefQA: Coreference Resolution as Query-based Span PredictionabstractIn this paper, we present CorefQA, an accurate and extensible approach for the coreference resolution task.We formulate the problem as a span prediction task, like in question answering: A query is generated for each candidate mention using its surrounding context, and a span prediction module is employed to extract the text spans of the coreferences within the document using the generated query.This formulation comes with the following key advantages: (1) The span prediction strategy provides the flexibility of retrieving mentions left out at the mention proposal stage; (2) In the question answering framework, encoding the mention and its context explicitly in a query makes it possible to have a deep and thorough examination of cues embedded in the context of coreferent mentions; and (3) A plethora of existing question answering datasets can be used for data augmentation to improve the model's generalization capability.Experiments demonstrate significant performance boost over previous models, with 83.1 (+3.5)F1 score on the CoNLL-2012 benchmark and 87.5 (+2.5)F1 score on the GAP benchmark.1 Wei Wu 0044, Fei Wang 0060, Arianna Yuan, Fei Wu 0001, Jiwei Li 0001 |
ACL | 5 |
| 2020 | Description Based Text Classification with Reinforcement LearningabstractThe task of text classification is usually divided into two stages: text feature extraction and classification. In this standard formalization, categories are merely represented as indexes in the label vocabulary, and the model lacks for explicit instructions on what to classify. Inspired by the current trend of formalizing NLP problems as question answering tasks, we propose a new framework for text classification, in which each category label is associated with a category description. Descriptions are generated by hand-crafted templates or using abstractive/extractive models from reinforcement learning. The concatenation of the description and the text is fed to the classifier to decide whether or not the current label should be assigned to the text. The proposed strategy forces the model to attend to the most salient texts with respect to the label, which can be regarded as a hard version of attention, leading to better performances. We observe significant performance boosts over strong baselines on a wide range of text classification tasks including single-label classification, multi-label classification and multi-aspect sentiment analysis. Duo Chai, Wei Wu 0044, Qinghong Han, Fei Wu 0001, Jiwei Li 0001 |
ICML | 5 |
| 2020 | SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive ConnectionabstractWhile the self-attention mechanism has been widely used in a wide variety of tasks, it has the unfortunate property of a quadratic cost with respect to the input length, which makes it difficult to deal with long inputs. In this paper, we present a method for accelerating and structuring self-attentions: Sparse Adaptive Connection (SAC). In SAC, we regard the input sequence as a graph and attention operations are performed between linked nodes. In contrast with previous self-attention models with pre-defined structures (edges), the model learns to construct attention edges to improve task-specific performances. In this way, the model is able to select the most salient nodes and reduce the quadratic complexity regardless of the sequence length. Based on SAC, we show that previous variants of self-attention models are its special cases. Through extensive experiments on neural machine translation, language modeling, graph representation learning and image classification, we demonstrate SAC is competitive with state-of-the-art models while significantly reducing memory cost. Xiaoya Li 0001, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu 0001, Jiwei Li 0001 |
NeurIPS | 6 |
| 2019 | Is Word Segmentation Necessary for Deep Learning of Chinese Representations?abstractSegmenting a chunk of text into words is usually the first step of processing Chinese text, but its necessity has rarely been explored.In this paper, we ask the fundamental question of whether Chinese word segmentation (CWS) is necessary for deep learning-based Chinese Natural Language Processing.We benchmark neural word-based models which rely on word segmentation against neural char-based models which do not involve word segmentation in four end-to-end NLP benchmark tasks: language modeling, machine translation, sentence matching/paraphrase and text classification.Through direct comparisons between these two types of models, we find that charbased models consistently outperform wordbased models.Based on these observations, we conduct comprehensive experiments to study why wordbased models underperform char-based models in these deep learning-based NLP tasks.We show that it is because word-based models are more vulnerable to data sparsity and the presence of out-of-vocabulary (OOV) words, and thus more prone to overfitting.We hope this paper could encourage researchers in the community to rethink the necessity of word segmentation in deep learning-based Chinese Natural Language Processing. 1 Xiaoya Li 0001, Yuxian Meng, Xiaofei Sun 0001, Qinghong Han, Arianna Yuan, Jiwei Li 0001 |
ACL (1) | 6 |
| 2019 | Entity-Relation Extraction as Multi-Turn Question AnsweringabstractIn this paper, we propose a new paradigm for the task of entity-relation extraction.We cast the task as a multi-turn question answering problem, i.e., the extraction of entities and relations is transformed to the task of identifying answer spans from the context.This multi-turn QA formalization comes with several key advantages: firstly, the question query encodes important information for the entity/relation class we want to identify; secondly, QA provides a natural way of jointly modeling entity and relation; and thirdly, it allows us to exploit the well developed machine reading comprehension (MRC) models.Experiments on the ACE and the CoNLL04 corpora demonstrate that the proposed paradigm significantly outperforms previous best models.We are able to obtain the stateof-the-art results on all of the ACE04, ACE05 and CoNLL04 datasets, increasing the SOTA results on the three datasets to 49.4 (+1.0), 60.2 (+0.6) and 68.9 (+2.1), respectively.Additionally, we construct a newly developed dataset RESUME in Chinese, which requires multi-step reasoning to construct entity dependencies, as opposed to the single-step dependency extraction in the triplet exaction in previous datasets.The proposed multi-turn QA model also achieves the best performance on the RESUME dataset. 1 Xiaoya Li 0001, Fan Yin, Zijun Sun, Xiayu Li, Arianna Yuan, Duo Chai, Mingxin Zhou, Jiwei Li 0001 |
ACL (1) | 8 |
| 2019 | Glyce: Glyph-vectors for Chinese Character RepresentationsabstractIt is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, an effective way to utilize the glyph information remains to be found. In this paper, we address this gap by presenting Glyce, the glyph-vectors for Chinese character representations. We make three major innovations: (1) We use historical Chinese scripts (e.g., bronzeware script, seal script, traditional Chinese, etc) to enrich the pictographic evidence in characters; (2) We design CNN structures (called tianzege-CNN) tailored to Chinese character image processing; and (3) We use image-classification as an auxiliary task in a multi-task learning setup to increase the model's ability to generalize. We show that glyph-based models are able to consistently outperform word/char ID-based models in a wide range of Chinese NLP tasks. When combing with BERT, we are able to set new state-of-the-art results for a variety of Chinese NLP tasks, including language modeling, tagging (NER, CWS, POS), sentence pair classification (BQ, LCQMC, XNLI, NLPCC-DBQA), single sentence classification tasks (ChnSentiCorp, the Fudan corpus, iFeng), dependency parsing, and semantic role labeling. For example, the proposed model achieves an F1 score of 81.6 on the OntoNotes dataset of NER, +1.5 over BERT; it achieves an almost perfect accuracy of 99.8\% on the the Fudan corpus for text classification. Yuxian Meng, Wei Wu 0044, Fei Wang 0060, Xiaoya Li 0001, Ping Nie, Fan Yin, Muyu Li, Qinghong Han, Xiaofei Sun 0001, Jiwei Li 0001 |
NeurIPS | 10 |
| 2018 | Generating More Interesting Responses in Neural Conversation Models with Distributional ConstraintsabstractNeural conversation models tend to generate safe, generic responses for most inputs.This is due to the limitations of likelihoodbased decoding objectives in generation tasks with diverse outputs, such as conversation.To address this challenge, we propose a simple yet effective approach for incorporating side information in the form of distributional constraints over the generated responses.We propose two constraints that help generate more content rich responses that are based on a model of syntax and topics (Griffiths et al., 2005) and semantic similarity (Arora et al., 2016).We evaluate our approach against a variety of competitive baselines, using both automatic metrics and human judgments, showing that our proposed approach generates responses that are much less generic without sacrificing plausibility.A working demo of our code can be found at https://github.com/abaheti95/ DC-NeuralConversation. Ashutosh Baheti, Alan Ritter, Jiwei Li 0001, William B. Dolan |
EMNLP | 3 |
| 2017 | Neural Net Models of Open-domain Discourse CoherenceabstractDiscourse coherence is strongly associated with text quality, making it important to natural language generation and understanding.Yet existing models of coherence focus on measuring individual aspects of coherence (lexical overlap, rhetorical structure, entity centering) in narrow domains.In this paper, we describe domainindependent neural models of discourse coherence that are capable of measuring multiple aspects of coherence in existing sentences and can maintain coherence while generating new sentences.We study both discriminative models that learn to distinguish coherent from incoherent discourse, and generative models that produce coherent text, including a novel neural latentvariable Markovian generative model that captures the latent discourse dependencies between sentences in a text.Our work achieves state-of-the-art performance on multiple coherence evaluations, and marks an initial step in generating coherent texts given discourse contexts. Jiwei Li 0001, Daniel Jurafsky |
EMNLP | 1 |
| 2017 | Adversarial Learning for Neural Dialogue GenerationabstractIn this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from human-generated dialogue utterances.We cast the task as a reinforcement learning (RL) problem where we jointly train two systems, a generative model to produce response sequences, and a discriminator-analagous to the human evaluator in the Turing test-to distinguish between the human-generated dialogues and the machine-generated ones.The outputs from the discriminator are then used as rewards for the generative model, pushing the system to generate dialogues that mostly resemble human dialogues.In addition to adversarial training we describe a model for adversarial evaluation that uses success in fooling an adversary as a dialogue evaluation metric, while avoiding a number of potential pitfalls.Experimental results on several metrics, including adversarial evaluation, demonstrate that the adversarially-trained system generates higher-quality responses than previous baselines. Jiwei Li 0001, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, Daniel Jurafsky |
EMNLP | 1 |
| 2017 | Data Noising as Smoothing in Neural Network Language Models
Ziang Xie, Sida I. Wang, Jiwei Li 0001, Daniel Levy 0002, Aiming Nie, Daniel Jurafsky, Andrew Y. Ng |
ICLR (Poster) | 3 |
| 2016 | A Persona-Based Neural Conversation ModelabstractJiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, Bill Dolan. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao 0001, William B. Dolan |
ACL (1) | 1 |
| 2016 | Deep Reinforcement Learning for Dialogue GenerationabstractRecent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on future outcomes.Modeling the future direction of a dialogue is crucial to generating coherent, interesting dialogues, a need which led traditional NLP models of dialogue to draw on reinforcement learning.In this paper, we show how to integrate these goals, applying deep reinforcement learning to model future reward in chatbot dialogue.The model simulates dialogues between two virtual agents, using policy gradient methods to reward sequences that display three useful conversational properties: informativity, coherence, and ease of answering (related to forward-looking function).We evaluate our model on diversity, length as well as with human judges, showing that the proposed algorithm generates more interactive responses and manages to foster a more sustained conversation in dialogue simulation.This work marks a first step towards learning a neural conversational model based on the long-term success of dialogues. Jiwei Li 0001, Will Monroe, Alan Ritter, Daniel Jurafsky, Michel Galley, Jianfeng Gao 0001 |
EMNLP | 1 |
| 2016 | Visualizing and Understanding Neural Models in NLPabstractWhile neural networks have been successfully applied to many NLP tasks the resulting vectorbased models are very difficult to interpret.For example it's not clear how they achieve compositionality, building sentence meaning from the meanings of words and phrases.In this paper we describe strategies for visualizing compositionality in neural models for NLP, inspired by similar work in computer vision.We first plot unit values to visualize compositionality of negation, intensification, and concessive clauses, allowing us to see wellknown markedness asymmetries in negation.We then introduce methods for visualizing a unit's salience, the amount that it contributes to the final composed meaning from first-order derivatives.Our general-purpose methods may have wide applications for understanding compositionality and other semantic properties of deep networks. Jiwei Li 0001, Xinlei Chen, Eduard H. Hovy, Daniel Jurafsky |
HLT-NAACL | 1 |
| 2016 | A Diversity-Promoting Objective Function for Neural Conversation ModelsabstractJiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
HLT-NAACL | 1 |
| 2015 | A Hierarchical Neural Autoencoder for Paragraphs and DocumentsabstractJiwei Li, Thang Luong, Dan Jurafsky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Jiwei Li 0001, Minh-Thang Luong, Daniel Jurafsky |
ACL (1) | 1 |
| 2015 | Do Multi-Sense Embeddings Improve Natural Language Understanding?abstractLearning a distinct representation for each sense of an ambiguous word could lead to more powerful and fine-grained mod-els of vector-space representations. Yet while ‘multi-sense ’ methods have been proposed and tested on artificial word-similarity tasks, we don’t know if they im-prove real natural language understanding tasks. In this paper we introduce a multi-sense embedding model based on Chinese Restaurant Processes that achieves state of the art performance on matching human word similarity judgments, and propose a pipelined architecture for incorporating multi-sense embeddings into language un-derstanding. We then test the performance of our model on part-of-speech tagging, named entity recognition, sentiment analysis, semantic relation identification and semantic relat-edness, controlling for embedding dimen-sionality. We find that multi-sense embed-dings do improve performance on some tasks (part-of-speech tagging, semantic re-lation identification, semantic relatedness) but not on others (named entity recogni-tion, various forms of sentiment analysis). We discuss how these differences may be caused by the different role of word sense information in each of the tasks. The re-sults highlight the importance of testing embedding models in real applications. 1 Jiwei Li 0001, Daniel Jurafsky |
EMNLP | 1 |
| 2015 | When Are Tree Structures Necessary for Deep Learning of Representations?abstractRecursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture.However there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate.In this paper, we benchmark recursive neural models against sequential recurrent neural models, enforcing applesto-apples comparison as much as possible.We investigate 4 tasks: (1) sentiment classification at the sentence level and phrase level; (2) matching questions to answerphrases; (3) discourse parsing; (4) semantic relation extraction.Our goal is to understand better when, and why, recursive models can outperform simpler models.We find that recursive models help mainly on tasks (like semantic relation extraction) that require longdistance connection modeling, particularly on very long sequences.We then introduce a method for allowing recurrent models to achieve similar performance: breaking long sentences into clause-like units at punctuation and processing them separately before combining.Our results thus help understand the limitations of both classes of models, and suggest directions for improving recurrent models. Jiwei Li 0001, Thang Luong, Daniel Jurafsky, Eduard H. Hovy |
EMNLP | 1 |
| 2014 | Towards a General Rule for Identifying Deceptive Opinion SpamabstractConsumers' purchase decisions are increasingly influenced by user-generated online reviews.Accordingly, there has been growing concern about the potential for posting deceptive opinion spamfictitious reviews that have been deliberately written to sound authentic, to deceive the reader.In this paper, we explore generalized approaches for identifying online deceptive opinion spam based on a new gold standard dataset, which is comprised of data from three different domains (i.e.Hotel, Restaurant, Doctor), each of which contains three types of reviews, i.e. customer generated truthful reviews, Turker generated deceptive reviews and employee (domain-expert) generated deceptive reviews.Our approach tries to capture the general difference of language usage between deceptive and truthful reviews, which we hope will help customers when making purchase decisions and review portal operators, such as TripAdvisor or Yelp, investigate possible fraudulent activity on their sites.1 Jiwei Li 0001, Myle Ott, Claire Cardie, Eduard H. Hovy |
ACL (1) | 1 |
| 2014 | Weakly Supervised User Profile Extraction from TwitterabstractWhile user attribute extraction on social media has received considerable attention, existing approaches, mostly supervised, encounter great difficulty in obtaining gold standard data and are therefore limited to predicting unary predicates (e.g., gender).In this paper, we present a weaklysupervised approach to user profile extraction from Twitter.Users' profiles from social media websites such as Facebook or Google Plus are used as a distant source of supervision for extraction of their attributes from user-generated text.In addition to traditional linguistic features used in distant supervision for information extraction, our approach also takes into account network information, a unique opportunity offered by social media.We test our algorithm on three attribute domains: spouse, education and job; experimental results demonstrate our approach is able to make accurate predictions for users' attributes based on their tweets.1• We experimentally demonstrate the effectiveness of our approach on 3 relations: SPOUSE, JOB and EDUCATION.The remainder of this paper is organized as follows: We summarize related work in Section 2. The creation of our dataset is described in Section 3. The details of our model are presented in Section 4. We present experimental results in Section 5 and conclude in Section 6. Jiwei Li 0001, Alan Ritter, Eduard H. Hovy |
ACL (1) | 1 |
| 2014 | What a Nasty Day: Exploring Mood-Weather Relationship from TwitterabstractWhile it has long been believed in psychology that weather somehow influences human's mood, the debates have been going on for decades about how they are correlated. In this paper, we try to study this long-lasting topic by harnessing a new source of data compared from traditional psychological researches: Twitter. We analyze 2 years' twitter data collected by twitter API which amounts to 10% of all postings and try to reveal the correlations between multiple dimensional structure of human mood with meteorological effects. Some of our findings confirm existing hypotheses, while others contradict them. We are hopeful that our approach, along with the new data source, can shed on the long-going debates on weather-mood correlation. Jiwei Li 0001, Eduard H. Hovy |
CIKM | 1 |
| 2014 | Sentiment Analysis on the People's DailyabstractWe propose a semi-supervised bootstrap-ping algorithm for analyzing China’s for-eign relations from the People’s Daily. Our approach addresses sentiment tar-get clustering, subjective lexicons extrac-tion and sentiment prediction in a unified framework. Different from existing algo-rithms in the literature, time information is considered in our algorithm through a hierarchical bayesian model to guide the bootstrapping approach. We are hopeful that our approach can facilitate quantita-tive political analysis conducted by social scientists and politicians. 1 Jiwei Li 0001, Eduard H. Hovy |
EMNLP | 1 |
| 2014 | A Model of Coherence Based on Distributed Sentence RepresentationabstractCoherence is what makes a multi-sentence text meaningful, both logically and syntactically.To solve the challenge of ordering a set of sentences into coherent order, existing approaches focus mostly on defining and using sophisticated features to capture the cross-sentence argumentation logic and syntactic relationships.But both argumentation semantics and crosssentence syntax (such as coreference and tense rules) are very hard to formalize.In this paper, we introduce a neural network model for the coherence task based on distributed sentence representation.The proposed approach learns a syntacticosemantic representation for sentences automatically, using either recurrent or recursive neural networks.The architecture obviated the need for feature engineering, and learns sentence representations, which are to some extent able to capture the 'rules' governing coherent sentence structure.The proposed approach outperforms existing baselines and generates the stateof-art performance in standard coherence evaluation tasks 1 . Jiwei Li 0001, Eduard H. Hovy |
EMNLP | 1 |
| 2014 | Recursive Deep Models for Discourse ParsingabstractText-level discourse parsing remains a challenge: most approaches employ fea-tures that fail to capture the intentional, se-mantic, and syntactic aspects that govern discourse coherence. In this paper, we pro-pose a recursive model for discourse pars-ing that jointly models distributed repre-sentations for clauses, sentences, and en-tire discourses. The learned representa-tions can to some extent learn the seman-tic and intentional import of words and larger discourse units automatically,. The proposed framework obtains comparable performance regarding standard discours-ing parsing evaluations when compared against current state-of-art systems. 1 Jiwei Li 0001, Rumeng Li, Eduard H. Hovy |
EMNLP | 1 |
| 2014 | Major Life Event Extraction from Twitter based on Congratulations/Condolences Speech ActsabstractSocial media websites provide a platform for anyone to describe significant events taking place in their lives in realtime.Currently, the majority of personal news and life events are published in a textual format, motivating information extraction systems that can provide a structured representations of major life events (weddings, graduation, etc. . .).This paper demonstrates the feasibility of accurately extracting major life events.Our system extracts a fine-grained description of users' life events based on their published tweets.We are optimistic that our system can help Twitter users more easily grasp information from users they take interest in following and also facilitate many downstream applications, for example realtime friend recommendation. Jiwei Li 0001, Alan Ritter, Claire Cardie, Eduard H. Hovy |
EMNLP | 1 |
| 2014 | Timeline generation: tracking individuals on twitterabstractIn this paper, we preliminarily learn the problem of reconstructing users' life history based on the their Twitter stream and proposed an unsupervised framework that create a chronological list for personal important events (PIE) of individuals. By analyzing individ- ual tweet collections, we find that what are suitable for inclusion in the personal timeline should be tweets talking about personal (as opposed to public) and time-specific (as opposed to time-general) topics. To further extract these types of topics, we introduce a non-parametric multi-level Dirichlet Process model to recognize four types of tweets: personal time-specific (PersonTS), personal time-general (PersonTG), public time-specific (PublicTS) and pub- lic time-general (PublicTG) topics, which, in turn, are used for fur- ther personal event extraction and timeline generation. To the best of our knowledge, this is the first work focused on the generation of timeline for individuals from Twitter data. For evaluation, we have built gold standard timelines that contain PIE related events from 20 ordinary twitter users and 20 celebrities. Experimental results demonstrate that it is feasible to automatically extract chronologi- cal timelines for Twitter users from their tweet collection Jiwei Li 0001, Claire Cardie |
WWW | 1 |
| 2013 | Identifying Manipulated Offerings on Review PortalsabstractRecent work has developed supervised methods for detecting deceptive opinion spamfake reviews written to sound authentic and deliberately mislead readers.And whereas past work has focused on identifying individual fake reviews, this paper aims to identify offerings (e.g., hotels) that contain fake reviews.We introduce a semi-supervised manifold ranking algorithm for this task, which relies on a small set of labeled individual reviews for training.Then, in the absence of gold standard labels (at an offering level), we introduce a novel evaluation procedure that ranks artificial instances of real offerings, where each artificial offering contains a known number of injected deceptive reviews.Experiments on a novel dataset of hotel reviews show that the proposed method outperforms state-of-art learning baselines. Jiwei Li 0001, Myle Ott, Claire Cardie |
EMNLP | 1 |