Linyang Li

dblp:228/8051 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Survey of Inductive Reasoning for Large Language Models
abstract
Kedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang, Siyu Yan, Xuecheng Wu, Yinqi Zhang, Qin Chen, Jie Zhou, Liang He, Biqing Qi, Linyang Li, Qipeng Guo, Xiaoming Shi, Wei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang, Yinqi Zhang, Qin Chen 0001, Jie Zhou 0015, Liang He 0001, Biqing Qi, Linyang Li, Qipeng Guo, Wayne Zhang 0001
ACL (1)12
2026 Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
abstract
Yichuan Ma, Linyang Li, Yongkang Chen, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichuan Ma, Linyang Li, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen 0026
ACL (1)2
2026 A system modeling and optimization method for value co-creation based on Lagrangian-Eulerian hybrid method
Zhonghang Bai, Linyang Li, Qinghang Zhong, Minhao Wang, Mingyu Sun
Adv. Eng. Informatics2
2026 Semantic-enhanced heterogeneous graph learning for identifying ncRNAs associated with drug resistance
abstract
MOTIVATION: Identifying non-coding RNAs (ncRNAs) associated with drug resistance is critical for elucidating molecular mechanisms underlying drug response, facilitating drug screening, and discovering novel therapeutic targets. While several graph neural network-based methods have been proposed to infer ncRNA-drug resistance associations, they remain fundamentally constrained by semantic distortion induced by a sparse bipartite network and neglect of relational semantics among molecular entities, ultimately compromising both predictive reliability and biological interpretability. RESULTS: In this study, we propose iNcRD-HG, a novel framework for identifying ncRNA-drug resistance associations. The framework addresses three critical aspects: constructing a context-enriched heterogeneous network that integrates six distinct molecular interaction types with bio-entity-specific attributes, developing a semantic-enhanced graph learning architecture that implements relation-type-aware message passing to capture complex contextual dependencies, and introducing an interpretability mechanism to reveal potential synergistic pathways underlying drug response. Experimental results demonstrate that iNcRD-HG achieves superior predictive performance across diverse benchmark datasets while deriving association features with strong discriminative capability. By identifying molecular synergistic contexts, iNcRD-HG provides mechanistically interpretable insights into ncRNA-mediated drug resistance. AVAILABILITY AND IMPLEMENTATION: Datasets and source codes are available at https://github.com/Biohang/iNcRD-HG.
Hang Wei 0005, Yuran Xie, Wenxiang Zhang, Linyang Li, Shuai Wu 0001, Lin Gao 0006
Bioinform.4
2026 ncRD-LG: A unified framework integrating molecular language models and subgraph learning for drug-ncRNA target prediction
Yuran Xie, Linyang Li, Shuai Wu 0001, Hang Wei 0005
Pattern Recognit.4
2025 FastMCTS: A Simple Sampling Strategy for Data Synthesis
abstract
Synthetic high-quality multi-step reasoning data can significantly enhance the performance of large language models on various tasks. However, most existing methods rely on rejection sampling, which generates trajectories independently and suffers from inefficiency and imbalanced sampling across problems of varying difficulty. In this work, we introduce FastMCTS, an innovative data synthesis strategy inspired by Monte Carlo Tree Search. FastMCTS provides a more efficient sampling method for multi-step reasoning data, offering step-level evaluation signals and promoting balanced sampling across problems of different difficulty levels. Experiments on both English and Chinese reasoning datasets demonstrate that FastMCTS generates over 30% more correct reasoning paths compared to rejection sampling as the number of generated tokens scales up. Furthermore, under comparable synthetic data budgets, models trained on FastMCTS-generated data outperform those trained on rejection sampling data by 3.9% across multiple benchmarks. As a lightweight sampling strategy, FastMCTS offers a practical and efficient alternative for synthesizing high-quality reasoning data.
Peiji Li, Kai Lv 0001, Yunfan Shao, Yichuan Ma, Linyang Li, Xiaoqing Zheng, Xipeng Qiu, Qipeng Guo
ACL (1)5
2025 Case2Code: Scalable Synthetic Data for Code Generation
abstract
Large Language Models (LLMs) have shown outstanding breakthroughs in code generation. Recent work improves code LLMs by training on synthetic data generated by some powerful LLMs, which can be challenging to scale due to the dependence on a teacher model and high generation costs. In this paper, we focus on synthesizing code data at scale and propose a Case2Code task by exploiting the expressiveness and correctness of programs. Case2Code is an inductive inference task that aims to infer underlying code implementations by observing input-output examples or program behaviors, By incorporating LLMs to generate program inputs, and executing the program with these inputs to obtain the program outputs, we can synthesize diverse and high-quality Case2Code data at scale for training and evaluating code LLMs. Experimental results show that case-to-code induction is challenging for current representative LLMs if they are untrained. Models trained with Case2Code improve performance not only on distribution case-to-code induction but also various coding-generation tasks, demonstrating the great potential of large-scale synthetic data and inductive learning.
Yunfan Shao, Linyang Li, Yichuan Ma, Peiji Li, Demin Song, Qinyuan Cheng, Pengyu Wang 0006, Qipeng Guo, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001, Dahua Lin
COLING2
2025 UnitCoder: Scalable Code Synthesis from Pre-training Corpora
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, yet code generation remains a major challenge.Despite the abundant sources of code data, constructing high-quality training datasets at scale poses a significant challenge.Pre-training code data typically suffers from inconsistent data quality issues.Conversely, instruction-based methods which use a high-quality subset as seed samples suffer from limited task diversity.In this paper, we introduce UnitCoder, which directly supervises pre-training data quality through automatically generated unit tests, while ensuring the correctness via an iterative fix and refine flow.Code synthesized by Unit-Coder benefits from both the diversity of pretraining corpora and the high quality ensured by unit test supervision.Our experiments demonstrate that models fine-tuned on our synthetic dataset exhibit consistent performance improvements.Our work presents a scalable approach that leverages model-generated unit tests to guide the synthesis of high-quality code data from pre-training corpora, demonstrating the potential for producing diverse and high-quality post-training data at scale.All code and data will be released 1 .
Yichuan Ma, Yunfan Shao, Peiji Li, Demin Song, Qipeng Guo, Linyang Li, Xipeng Qiu, Kai Chen 0026
EMNLP6
2025 Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
abstract
Large language models (LLMs) have demonstrated exceptional performance in reasoning tasks such as mathematics and coding, matching or surpassing human capabilities. However, these impressive reasoning abilities face significant challenges in specialized domains. Taking Go as an example, although AlphaGo has established the high performance ceiling of AI systems in Go, mainstream LLMs still struggle to reach even beginner-level proficiency, let alone perform natural language reasoning. This performance gap between general-purpose LLMs and domain experts is significantly limiting the application of LLMs on a wider range of domain-specific tasks. In this work, we aim to bridge the divide between LLMs' general reasoning capabilities and expert knowledge in domain-specific tasks. We perform mixed fine-tuning with structured Go expertise and general long Chain-of-Thought (CoT) reasoning data as a cold start, followed by reinforcement learning to integrate expert knowledge in Go with general reasoning capabilities. Through this methodology, we present LoGos, a powerful LLM that not only maintains outstanding general reasoning abilities, but also conducts Go gameplay in natural language, demonstrating effective strategic reasoning and accurate next-move prediction. LoGos achieves performance comparable to human professional players, substantially surpassing all existing LLMs. Through this work, we aim to contribute insights on applying general LLM reasoning capabilities to specialized domains. We will release the first large-scale Go dataset for LLM training, the first LLM Go evaluation benchmark, and the first general LLM that reaches human expert-level performance in Go.
Yichuan Ma, Linyang Li, Peiji Li, Jiasheng Ye, Qipeng Guo, Dahua Lin, Kai Chen 0026
NeurIPS2
2025 Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
abstract
Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to \textbf{25\%} relative gain and \textbf{6\%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation.
Bo Wang 0084, Qinyuan Cheng, Runyu Peng, Rong Bao, Peiji Li, Qipeng Guo, Linyang Li, Zhiyuan Zeng 0004, Yunhua Zhou, Xipeng Qiu
NeurIPS7
2025 Dual-Scanning Photoacoustic Endomicroscopy for High-Speed Gastrointestinal Microvascular Imaging
abstract
Photoacoustic endomicroscopy enables high-resolution imaging of deep microvasculature within the gastrointestinal wall using modulated laser pulses with point-by-point scanning. However, conventional scanning mechanisms frequently encounter difficulties in balancing imaging speed and field of view, particularly when imaging the peristaltic gastrointestinal tract. To address this challenge, we propose a dual-scanning photoacoustic endomicroscopy with an adjustable focal plane and an ultrafast imaging speed. The probe features two distinct scanning modes: 360° angular scanning providing a wide field of view, and regional spiral scanning offering high image quality. We demonstrated the capability of this probe through imaging both phantoms and rat rectums. The results from the rectal injury model demonstrate the applicability and sensitivity of the probe. Overall, this study offers new perspectives for expanding the applications and clinical potential of photoacoustic endomicroscopy.
Hongdian Sun, Linyang Li, Yuanlong Zhao, Weizhi Qi
IEEE Trans. Medical Imaging3
2024 AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
abstract
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junqi Dai, Jiasheng Ye, Yunhua Zhou, Zhigeng Liu, Ruibin Yuan, Ge Zhang 0009, Linyang Li, Hang Yan 0001, Jie Fu 0001, Tao Gui, Tianxiang Sun, Yu-Gang Jiang 0001, Xipeng Qiu
ACL (1)10
2024 InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
abstract
Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Mozhi Zhang, Ke Ren, Botian Jiang, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Pengyu Wang 0006, Linyang Li, Chenkun Tan, Mozhi Zhang, Botian Jiang, Xipeng Qiu
EMNLP3
2024 Turn Waste into Worth: Rectifying Top-k Router of MoE
abstract
Zhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan, Dahua Lin, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhiyuan Zeng 0004, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan 0001, Dahua Lin, Xipeng Qiu
EMNLP6
2024 Can AI Assistants Know What They Don't Know?
abstract
AI assistants powered by Large Language Models (LLMs) have demonstrated impressive performance in various tasks. However, LLMs still make factual errors in knowledge-intensive tasks such as open-domain question answering. These untruthful responses from AI assistants can pose significant risks in practical applications. Therefore, in this paper, we ask the question Can AI assistants know what they don’t know and express this awareness through natural language? To investigate this, we construct a model-specific "I don’t know" (Idk) dataset. This dataset includes Supervised Fine-tuning data and preference data, categorizing questions based on whether the assistant knows or does not know the answers. Then, we align the assistant with its corresponding Idk dataset using different alignment methods, including Supervised Fine-tuning and preference optimization. Experimental results show that, after alignment with the Idk dataset, the assistant is more capable of declining to answer questions outside its knowledge scope. The assistant aligned with the Idk dataset shows significantly higher truthfulness than the original assistant.
Qinyuan Cheng, Tianxiang Sun, Zhangyue Yin, Linyang Li, Zhengfu He, Kai Chen 0026, Xipeng Qiu
ICML7
2024 LLatrieval: LLM-Verified Retrieval for Verifiable Generation
abstract
Xiaonan Li, Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, Xipeng Qiu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, Xipeng Qiu
NAACL-HLT3
2024 AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing Images
abstract
The rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM.
Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Mitigating Negative Style Transfer in Hybrid Dialogue System
abstract
As the functionality of dialogue systems evolves, hybrid dialogue systems that accomplish user-specific goals and participate in open-topic chitchat with users are attracting growing attention. Existing research learns both tasks concurrently utilizing a multi-task fusion technique but ignores the negative transfer phenomenon induced by the unique textual style differences. Therefore, contrastive learning based on the latent variable model is used to decouple the various textual genres in the latent space. We devise supervised and self-supervised positive and negative sample constructions for diverse datasets. In addition, to capitalize on the style information contained in the decoupled latent variables, we employ a style prefix that incorporates latent variables further to control the generation of responses with varying styles. We performed extensive experiments on three dialogue datasets, including a hybrid dialogue dataset and two task-oriented dialogue datasets. The experimental results demonstrate that our method can mitigate the negative style transfer issue and achieves state-of-the-art performance on multiple dialogue datasets.
Qinyuan Cheng, Linyang Li, Xipeng Qiu
AAAI3
2023 Text Adversarial Purification as Defense against Adversarial Attacks
abstract
Adversarial purification is a successful defense mechanism against adversarial attacks without requiring knowledge of the form of the incoming attack.Generally, adversarial purification aims to remove the adversarial perturbations therefore can make correct predictions based on the recovered clean samples.Despite the success of adversarial purification in the computer vision field that incorporates generative models such as energy-based models and diffusion models, using purification as a defense strategy against textual adversarial attacks is rarely explored.In this work, we introduce a novel adversarial purification method that focuses on defending against textual adversarial attacks.With the help of language models, we can inject noise by masking input texts and reconstructing the masked texts based on the masked language models.In this way, we construct an adversarial purification process for textual models against the most widely used word-substitution adversarial attacks.We test our proposed adversarial purification method on several strong adversarial attack methods including Textfooler and BERT-Attack and experimental results indicate that the purification algorithm can successfully defend against strong word-substitution attacks.
Linyang Li, Demin Song, Xipeng Qiu
ACL (1)1
2023 Character-LLM: A Trainable Agent for Role-Playing
abstract
Large language models (LLMs) can be used to serve as agents to simulate human behaviors, given the powerful ability to understand human instructions and provide high-quality generated texts.Such ability stimulates us to wonder whether LLMs can simulate a person in a higher form than simple human behaviors.Therefore, we aim to train an agent with the profile, experience, and emotional states of a specific person instead of using limited prompts to instruct ChatGPT API.In this work, we introduce Character-LLM that teach LLMs to act as specific people such as Beethoven, Queen Cleopatra, Julius Caesar, etc.Our method focuses on editing profiles as experiences of a certain character and training models to be personal simulacra with these experiences.To assess the effectiveness of our approach, we build a test playground that interviews trained agents and evaluates whether the agents memorize their characters and experiences.Experimental results show interesting observations that help build future simulacra of humankind.1
Yunfan Shao, Linyang Li, Junqi Dai, Xipeng Qiu
EMNLP2
2023 SeqXGPT: Sentence-Level AI-Generated Text Detection
abstract
Widely applied large language models (LLMs) can generate human-like content, raising concerns about the abuse of LLMs.Therefore, it is important to build strong AI-generated text (AIGT) detectors.Current works only consider document-level AIGT detection, therefore, in this paper, we first introduce a sentence-level detection challenge by synthesizing a dataset that contains documents that are polished with LLMs, that is, the documents contain sentences written by humans and sentences modified by LLMs.Then we propose Sequence X (Check) GPT, a novel method that utilizes log probability lists from white-box LLMs as features for sentence-level AIGT detection.These features are composed like waves in speech processing and cannot be studied by LLMs.Therefore, we build SeqXGPT based on convolution and self-attention networks.We test it in both sentence and document-level detection challenges.Experimental results show that previous methods struggle in solving sentence-level AIGT detection, while our method not only significantly surpasses baseline methods in both sentence and document-level detection challenges but also exhibits strong generalization capabilities.1
Pengyu Wang 0006, Linyang Li, Botian Jiang, Xipeng Qiu
EMNLP2
2023 MarkBERT: Marking Word Boundaries Improves Chinese BERT
Linyang Li, Yong Dai 0001, Duyu Tang, Xipeng Qiu, Shuming Shi 0001
NLPCC (1)1
2022 Template-free Prompt Tuning for Few-shot NER
abstract
Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ruotian Ma, Xin Zhou 0012, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001
NAACL-HLT5
2022 Hybridizing Euclidean and Hyperbolic Similarities for Attentively Refining Representations in Semantic Segmentation of Remote Sensing Images
abstract
Attention mechanisms have revolutionized the semantic segmentation network in interpreting remotely sensed images (RSIs) due to their amazing ability in establishing contextual dependencies. Nevertheless, due to the complex scenes and diverse objects in RSIs, a variety of details and correlations are not available in Euclidean space. Therefore, a similarity-hybrid attention module (SHAM) is devised to attentively learn the hyperbolic and Euclidean attention maps between any two positions, followed by a weighted element-wise summation. The hybrid attention maps posses latent geometric properties of both Euclidean and hyperboloid. Taking commonly-used fully convolutional network (FCN) as baseline, HAENet that embeds SHAM, is presented. Experiments on ISPRS Potsdam and DeepGlobe benchmarks reveal its superiority to comparative methods. In addition, the ablation study validates the effectiveness of SHAM compared to other attention modules.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Runliang Xia, Linyang Li, Zhennan Xu, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.6
2021 Token-Aware Virtual Adversarial Training in Natural Language Understanding
abstract
Gradient-based adversarial training is widely used in improving the robustness of neural networks, while it cannot be easily adapted to natural language processing tasks since the embedding space is discrete. In natural language processing fields, virtual adversarial training is introduced since texts are discrete and cannot be perturbed by gradients directly. Alternatively, virtual adversarial training, which generates perturbations on the embedding space, is introduced in NLP tasks. Despite its success, existing virtual adversarial training methods generate perturbations roughly constrained by Frobenius normalization balls. To craft fine-grained perturbations, we propose a Token-Aware Virtual Adversarial Training method. We introduce a token-level accumulated perturbation vocabulary to initialize the perturbations better and use a token-level normalization ball to constrain these perturbations pertinently. Experiments show that our method improves the performance of pre-trained models such as BERT and ALBERT in various tasks by a considerable margin. The proposed method improves the score of the GLUE benchmark from 78.3 to 80.9 using BERT model and it also enhances the performance of sequence labeling and text classification tasks.
Linyang Li, Xipeng Qiu
AAAI1
2021 SENT: Sentence-level Distant Relation Extraction via Negative Training
abstract
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang, Xuanjing Huang, Yaqian Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001, Yaqian Zhou 0001
ACL/IJCNLP (1)3
2021 Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
abstract
Pre-Trained Models have been widely applied and recently proved vulnerable under backdoor attacks: the released pre-trained weights can be maliciously poisoned with certain triggers.When the triggers are activated, even the fine-tuned model will predict pre-defined labels, causing a security threat.These backdoors generated by the poisoning methods can be erased by changing hyper-parameters during fine-tuning or detected by finding the triggers.In this paper, we propose a stronger weight-poisoning attack method that introduces a layerwise weight poisoning strategy to plant deeper backdoors; we also introduce a combinatorial trigger that cannot be easily detected.The experiments on text classification tasks show that previous defense methods cannot resist our weight-poisoning method, which indicates that our method can be widely applied and may provide hints for future model robustness studies.
Linyang Li, Demin Song, Jiehang Zeng, Ruotian Ma, Xipeng Qiu
EMNLP (1)1
2021 Searching for an Effective Defender: Benchmarking Defense against Adversarial Word Substitution
abstract
Recent studies have shown that deep neural network-based models are vulnerable to intentionally crafted adversarial examples, and various methods have been proposed to defend against adversarial word-substitution attacks for neural NLP models.However, there is a lack of systematic study on comparing different defense approaches under the same attacking setting.In this paper, we seek to fill the gap through comprehensive studies on the behavior of neural text classifiers trained with various defense methods against representative adversarial attacks.In addition, we propose an effective method to further improve the robustness of neural text classifiers against such attacks, and achieved the highest accuracy on both clean and adversarial examples on AGNEWS and IMDB datasets, outperforming existing methods by a significant margin.We hope this study could provide useful clues for future research on text adversarial defense.Codes are available at https:// github.com/RockyLzy/TextDefender.
Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang 0001, Kai-Wei Chang 0001, Cho-Jui Hsieh
EMNLP (1)4
2021 Enhancing Separate Encoding with Multi-layer Feature Alignment for Image-Text Matching
Keyu Wen, Linyang Li, Xiaodong Gu 0001
ICANN (1)2
2021 COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation
abstract
There has been a recent surge of interest in cross-modal pre-training. However, existed approaches pre-train a one-stream model to learn joint vision-language representation, which suffers from calculation explosion when conducting cross-modal retrieval. In this work, we propose the Contrastive Cross-Modal Knowledge Sharing Pretraining (COOKIE) method to learn universal text-image representations. There are two key designs in it, one is the weight-sharing transformer on top of the visual and textual encoders to align text and image semantically, the other is three kinds of contrastive learning designed for sharing knowledge between different modalities. Cross-modal knowledge sharing greatly promotes the learning of unimodal representation. Experiments on multi-modal matching tasks including cross-modal retrieval, text matching, and image retrieval show the effectiveness and efficiency of our pre-training framework. Our COOKIE finetuned on cross-modal datasets MSCOCO, Flickr30K, and MSRVTT achieves new state-of-the-art results while using only 3/1000 inference time comparing to one-stream models. There are also 5.7% and 3.9% improvements in the task of image retrieval and text matching. Source code will be available at https://github.com/kywen1119/COOKIE.
Keyu Wen, Jin Xia, Linyang Li, Jiayan Xu
ICCV4
2020 BERT-ATTACK: Adversarial Attack Against BERT Using BERT
abstract
Adversarial attacks for discrete data (such as texts) have been proved significantly more challenging than continuous data (such as images) since it is difficult to generate adversarial samples with gradient-based methods.Current successful attack methods for texts usually adopt heuristic replacement strategies on the character or word level, which remains challenging to find the optimal solution in the massive space of possible combinations of replacements while preserving semantic consistency and language fluency.In this paper, we propose BERT-Attack, a high-quality and effective method to generate adversarial samples using pre-trained masked language models exemplified by BERT.We turn BERT against its fine-tuned models and other deep neural models in downstream tasks so that we can successfully mislead the target models to predict incorrectly.Our method outperforms state-of-theart attack strategies in both success rate and perturb percentage, while the generated adversarial samples are fluent and semantically preserved.Also, the cost of calculation is low, thus possible for large-scale generations.The code is available at https://github.com/ LinyangLee/BERT-Attack.
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 0001, Xipeng Qiu
EMNLP (1)1