Aiwei Liu

dblp:321/4365 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0002-4965-8263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 24 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters
abstract
Prompt-based in-context learning (ICL) and parameter fine-tuning are two dominant paradigms for incorporating external information into large language models (LLMs), but they incur high inference costs or require expensive retraining.To bridge this gap, context-to-parameter mapping converts prompts into temporary adapter weights.However, we identify a critical failure mode in existing methods: hiddenstate collapse, where the adapter-augmented model's internal states diverge sharply from the full-context oracle in deeper layers.We trace this failure to two coupled gaps: suboptimal Input-Selection and inadequate Supervision-Signal.To address these issues, we propose SADA (State-Aligned Distillation Adapters).We establish the attention-block output as a principled feature interface to improve input selection and introduce statealignment distillation to enforce consistency between the adapter-augmented model and the full-context oracle.Experiments on long-context language modeling (PG19) and downstream NLU and summarization benchmarks show that SADA consistently outperforms strong baselines like StreamAdapter and GenerativeAdapter, achieving performance comparable to ICL while significantly reducing memory footprint and latency.We further analyze when parameterized context compression is effective and when explicit context retention remains preferable.
Tianlong Wang, Linhao Zhang, Aiwei Liu, Xiao Zhou 0004
ACL (1)5
2026 d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
abstract
Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, Lijie Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu 0003, Bolin Ding, Aiwei Liu, Lijie Wen 0001
ACL (1)10
2026 Query-based model extraction attack: A survey
Yulin Wu 0001, Aiwei Liu, Shuhan Qi
Pattern Recognit.4
2025 Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
abstract
Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Leyi Pan, Aiwei Liu, Shiyu Huang 0001, Yijian Lu, Xuming Hu, Lijie Wen 0001, Irwin King, Philip S. Yu
ACL (1)2
2025 ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
abstract
The literature review is an indispensable step in the research process. It provides the benefit of comprehending the research problem and understanding the current research situation while conducting a comparative analysis of prior works. However, literature summary is challenging and time consuming. The previous LLM-based studies on literature review mainly focused on the complete process, including literature retrieval, screening, and summarization. However, for the summarization step, simple CoT method often lacks the ability to provide extensive comparative summary. In this work, we firstly focus on the independent literature summarization step and introduce ChatCite, an LLM agent with human workflow guidance for comparative literature summary. This agent, by mimicking the human workflow, first extracts key elements from relevant literature and then generates summaries using a Reflective Incremental Mechanism. In order to better evaluate the quality of the generated summaries, we devised a LLM-based automatic evaluation metric, G-Score, in refer to the human evaluation criteria. The ChatCite agent outperformed other models in various dimensions in the experiments. The literature summaries generated by ChatCite can also be directly used for drafting literature reviews.
Lu Chen 0002, Aiwei Liu, Kai Yu 0004, Lijie Wen 0001
COLING3
2025 Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
abstract
Yunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou, Yanggan Gu, Jungang Li, Jingyu Wang, Peijie Jiang, Aiwei Liu, Jia Liu, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yunkai Dang, Mengxi Gao, Xin Zou 0001, Yanggan Gu, Jungang Li, Peijie Jiang, Aiwei Liu, Xuming Hu
EMNLP9
2025 VLA-Mark: A cross modal watermark for large vision-language alignment models
abstract
Shuliang Liu, Zheng Qi, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zheng Qi, Jesse Jiaxi Xu, He Geng, Aiwei Liu, Peijie Jiang, Yik-Cheung Tam, Xuming Hu
EMNLP7
2025 TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
abstract
Direct Preference Optimization (DPO) has been widely adopted for preference alignment of Large Language Models (LLMs) due to its simplicity and effectiveness. However, DPO is derived as a bandit problem in which the whole response is treated as a single arm, ignoring the importance differences between tokens, which may affect optimization efficiency and make it difficult to achieve optimal results. In this work, we propose that the optimal data for DPO has equal expected rewards for each token in winning and losing responses, as there is no difference in token importance. However, since the optimal dataset is unavailable in practice, we propose using the original dataset for importance sampling to achieve unbiased optimization. Accordingly, we propose a token-level importance sampling DPO objective named TIS-DPO that assigns importance weights to each token based on its reward. Inspired by previous works, we estimate the token importance weights using the difference in prediction probabilities from a pair of contrastive LLMs. We explore three methods to construct these contrastive LLMs: (1) guiding the original LLM with contrastive prompts, (2) training two separate LLMs using winning and losing responses, and (3) performing forward and reverse DPO training with winning and losing responses. Experiments show that TIS-DPO significantly outperforms various baseline methods on harmlessness and helpfulness alignment and summarization tasks. We also visualize the estimated weights, demonstrating their ability to identify key token positions.
Aiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun, Xiang Kong, Xiaoming Simon Wang, Jiulong Shan, Albin Madappally Jose, Xiaojiang Liu, Lijie Wen 0001, Philip S. Yu
ICLR1
2025 Can Watermarked LLMs be Identified by Users via Crafted Prompts?
abstract
Text watermarking for Large Language Models (LLMs) has made significant progress in detecting LLM outputs and preventing misuse. Current watermarking techniques offer high detectability, minimal impact on text quality, and robustness to text editing. However, current researches lack investigation into the imperceptibility of watermarking techniques in LLM services. This is crucial as LLM providers may not want to disclose the presence of watermarks in real-world scenarios, as it could reduce user willingness to use the service and make watermarks more vulnerable to attacks. This work is the first to investigate the imperceptibility of watermarked LLMs. We design an identification algorithm called Water-Probe that detects watermarks through well-designed prompts to the LLM. Our key motivation is that current watermarked LLMs expose consistent biases under the same watermark key, resulting in similar differences across prompts under different watermark keys. Experiments show that almost all mainstream watermarking algorithms are easily identified with our well-designed prompts, while Water-Probe demonstrates a minimal false positive rate for non-watermarked LLMs. Finally, we propose that the key to enhancing the imperceptibility of watermarked LLMs is to increase the randomness of watermark key selection. Based on this, we introduce the Water-Bag strategy, which significantly improves watermark imperceptibility by merging multiple watermark keys.
Aiwei Liu, Sheng Guan, Leyi Pan, Liancheng Fang, Lijie Wen 0001, Philip S. Yu, Xuming Hu
ICLR1
2025 Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
abstract
Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination. These biases arise from the visual encoder and the Large Language Model (LLM) backbone, affecting the attention mechanism responsible for aligning multimodal inputs. Existing decoding-based mitigation methods focus on statistical correlations and overlook the causal relationships between attention mechanisms and model output, limiting their effectiveness in addressing these biases. To tackle this issue, we propose a causal inference framework termed CausalMM that applies structural causal modeling to MLLMs, treating modality priors as a confounder between attention mechanisms and output. Specifically, by employing backdoor adjustment and counterfactual reasoning at both the visual and language attention levels, our method mitigates the negative effects of modality priors and enhances the alignment of MLLM's inputs and outputs, with a maximum score improvement of 65.3% on 6 VLind-Bench indicators and 164 points on MME Benchmark compared to conventional methods. Extensive experiments validate the effectiveness of our approach while being a plug-and-play solution. Our code is available at: https://github.com/The-Martyr/CausalMM.
Guanyu Zhou, Xin Zou 0001, Kun Wang 0056, Aiwei Liu, Xuming Hu
ICLR5
2025 GapDNER: A Gap-Aware Grid Tagging Model for Discontinuous Named Entity Recognition
abstract
In biomedical fields, one named entity may consist of a series of non-adjacent tokens and overlap with other entities. Previous methods recognize discontinuous entities by connecting entity fragments or internal tokens, which face challenges of error propagation and decoding ambiguity due to the wide variety of span or word combinations. To address these issues, we deeply explore discontinuous entity structures and propose an effective Gap-aware grid tagging model for Discontinuous Named Entity Recognition, named GapDNER. Our GapDNER innovatively applies representation learning on the context gaps between entity fragments to resolve decoding ambiguity and enhance discontinuous NER performance. Specifically, we treat the context gap as an additional type of span and convert span classification into a token-pair grid tagging task. Subsequently, we design two interactive components to comprehensively model token-pair grid features from both intra- and inter-span perspectives. The intra-span regularity extraction module employs the biaffine mechanism along with linear attention to capture the internal regularity of each span, while the inter-span relation enhancement module utilizes criss-cross attention to obtain semantic relations among different spans. At the inference stage of entity decoding, we assign a directed edge to each entity fragment and context gap, then use the BFS algorithm to search for all valid paths from the head to tail of grids with entity tags. Experimental results on three datasets demonstrate that our GapDNER achieves new state-of-the-art performance on discontinuous NER and exhibits remarkable advantages in recognizing complex entity structures.
Yawen Yang, Fukun Ma, Shiao Meng, Aiwei Liu, Lijie Wen 0001
IJCNN4
2025 GenCNER: A Generative Framework for Continual Named Entity Recognition
abstract
Traditional named entity recognition (NER) aims to identify text mentions into pre-defined entity types. Continual Named Entity Recognition (CNER) is introduced since entity categories are continuously increasing in various real-world scenarios. However, existing continual learning (CL) methods for NER face challenges of catastrophic forgetting and semantic shift of non-entity type. In this paper, we propose GenCNER, a simple but effective Generative framework for CNER to mitigate the above drawbacks. Specifically, we skillfully convert the CNER task into sustained entity triplet sequence generation problem and utilize a powerful pre-trained seq2seq model to solve it. Additionally, we design a type-specific confidence-based pseudo labeling strategy along with knowledge distillation (KD) to preserve learned knowledge and alleviate the impact of label noise at the triplet level. Experimental results on two benchmark datasets show that our framework outperforms previous state-of-the-art methods in multiple CNER settings, and achieves the smallest gap compared with non-CL results.
Yawen Yang, Fukun Ma, Shiao Meng, Aiwei Liu, Lijie Wen 0001
IJCNN4
2025 Entropy-Based Decoding for Retrieval-Augmented Large Language Models
abstract
Zexuan Qiu, Zijing Ou, Bin Wu, Jingjing Li, Aiwei Liu, Irwin King. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zexuan Qiu, Zijing Ou, Bin Wu 0025, Jingjing Li 0007, Aiwei Liu, Irwin King
NAACL (Long Papers)5
2024 Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models
abstract
Zhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, Rui Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhiwei He 0002, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang 0007, Zhaopeng Tu, Zhuosheng Zhang 0001, Rui Wang 0015
ACL (1)4
2024 Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
abstract
Aiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong, Xiaoming Wang, Jiulong Shan, Meng Cao, Lijie Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Aiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong, Jiulong Shan, Lijie Wen 0001
ACL (1)1
2024 An Entropy-based Text Watermarking Detection Method
abstract
Text watermarking algorithms for large language models (LLMs) can effectively identify machine-generated texts by embedding and detecting hidden features in the text.Although the current text watermarking algorithms perform well in most high-entropy scenarios, its performance in low-entropy scenarios still needs to be improved.In this work, we opine that the influence of token entropy should be fully considered in the watermark detection process, i.e., the weight of each token during watermark detection should be customized according to its entropy, rather than setting the weights of all tokens to the same value as in previous methods.Specifically, we propose Entropybased Text Watermarking Detection (EWD) that gives higher-entropy tokens higher influence weights during watermark detection, so as to better reflect the degree of watermarking.Furthermore, the proposed detection process is training-free and fully automated.From the experiments, we demonstrate that our EWD can achieve better detection performance in low-entropy scenarios, and our method is also general and can be applied to texts with different entropy distributions.Our code and data is available 1 .Additionally, our algorithm could be accessed through MarkLLM (Pan et al., 2024) 2 .
Yijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li 0007, Irwin King
ACL (1)2
2024 An Unforgeable Publicly Verifiable Watermark for Large Language Models
abstract
Recently, text watermarking algorithms for large language models (LLMs) have been proposed to mitigate the potential harms of text generated by LLMs, including fake news and copyright issues. However, current watermark detection algorithms require the secret key used in the watermark generation process, making them susceptible to security breaches and counterfeiting during public detection. To address this limitation, we propose an unforgeable publicly verifiable watermark algorithm named UPV that uses two different neural networks for watermark generation and detection, instead of using the same key at both stages. Meanwhile, the token embedding parameters are shared between the generation and detection networks, which makes the detection network achieve a high accuracy very efficiently. Experiments demonstrate that our algorithm attains high detection accuracy and computational efficiency through neural networks. Subsequent analysis confirms the high complexity involved in forging the watermark from the detection network. Our code is available at https://github.com/THU-BPM/unforgeable_watermark
Aiwei Liu, Leyi Pan, Xuming Hu, Shuang Li 0015, Lijie Wen 0001, Irwin King, Philip S. Yu
ICLR1
2024 A Semantic Invariant Robust Watermark for Large Language Models
abstract
Watermark algorithms for large language models (LLMs) have achieved extremely high accuracy in detecting text generated by LLMs. Such algorithms typically involve adding extra watermark logits to the LLM's logits at each generation step. However, prior algorithms face a trade-off between attack robustness and security robustness. This is because the watermark logits for a token are determined by a certain number of preceding tokens; a small number leads to low security robustness, while a large number results in insufficient attack robustness. In this work, we propose a semantic invariant watermarking method for LLMs that provides both attack robustness and security robustness. The watermark logits in our work are determined by the semantics of all preceding tokens. Specifically, we utilize another embedding LLM to generate semantic embeddings for all preceding tokens, and then these semantic embeddings are transformed into the watermark logits through our trained watermark model. Subsequent analyses and experiments demonstrated the attack robustness of our method in semantically invariant settings: synonym substitution and text paraphrasing settings. Finally, we also show that our watermark possesses adequate security robustness.
Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, Lijie Wen 0001
ICLR1
2024 Preventing and Detecting Misinformation Generated by Large Language Models
abstract
As large language models (LLMs) become increasingly capable and widely deployed, the risk of them generating misinformation poses a critical challenge. Misinformation from LLMs can take various forms, from factual errors due to hallucination to intentionally deceptive content, and can have severe consequences in high-stakes domains.This tutorial covers comprehensive strategies to prevent and detect misinformation generated by LLMs. We first introduce the types of misinformation LLMs can produce and their root causes. We then explore two broad categories: Preventing misinformation generation: a) AI alignment training techniques to reduce LLMs' propensity for misinformation and refuse malicious instructions during model training. b) Training-free mitigation methods like prompt guardrails, retrieval-augmented generation (RAG), and decoding strategies to curb misinformation at inference time. Detecting misinformation after generation, including a) using LLMs themselves to detect misinformation through embedded knowledge or retrieval-enhanced judgments, and b) distinguishing LLM-generated text from human-written text through black-box approaches (e.g., classifiers, probability analysis) and white-box approaches (e.g., watermarking). We also discuss the challenges and limitations of detecting LLM-generated misinformation.
Aiwei Liu, Qiang Sheng 0001, Xuming Hu
SIGIR1
2024 Reading Broadly to Open Your Mind: Improving Open Relation Extraction With Search Documents Under Self-Supervisions
abstract
Open relation extraction is the task of extracting open-domain relation facts from natural language sentences. Existing works either utilize distant-supervised annotations to train a supervised classifier over pre-defined relations, or adopt unsupervised methods with additional dependency on external assumptions. However, these works can only obtain information signals from limited existing knowledge bases or datasets. In this work, we propose a self-supervised framework namedWeb-SelfORE, which exploits self-supervised signals by requiring a large pretrained language model to extensively read real-world relevant documents from the web, and obtain contextualized relational features by mixing contextualized representations of entities from different documents. We perform adaptive clustering on contextualized relational features and bootstrap the self-supervised signals by improving contextualized features in relation classification. We additionally compare the effectiveness of self-supervisions brought by different document sources, and introduce relevance and redundancy evaluation metrics to obtain higher-quality self-supervisions. Experimental results on four public datasets show the effectiveness and robustness ofWeb-SelfOREon open-domain relation extraction task when comparing with competitive baselines.
Xuming Hu, Zhaochen Hong, Aiwei Liu, Shiao Meng, Lijie Wen 0001, Irwin King, Philip S. Yu
IEEE Trans. Knowl. Data Eng.4
2023 AMR-based Network for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment classification task.Many recent works have used dependency trees to extract the relation between aspects and contexts and have achieved significant improvements.However, further improvement is limited due to the potential mismatch between the dependency tree as a syntactic structure and the sentiment classification as a semantic task.To alleviate this gap, we replace the syntactic dependency tree with the semantic structure named Abstract Meaning Representation (AMR) and propose a model called AMR-based Path Aggregation Relational Network (APARN) to take full advantage of semantic structures.In particular, we design the path aggregator and the relation-enhanced selfattention mechanism that complement each other.The path aggregator extracts semantic features from AMRs under the guidance of sentence information, while the relationenhanced self-attention mechanism in turn improves sentence features with refined semantic information.Experimental results on four public datasets demonstrate 1.13% average F1 improvement of APARN in ABSA when compared with state-of-the-art baselines.1
Fukun Ma, Xuming Hu, Aiwei Liu, Yawen Yang, Shuang Li 0015, Philip S. Yu, Lijie Wen 0001
ACL (1)3
2023 RAPL: A Relation-Aware Prototype Learning Approach for Few-Shot Document-Level Relation Extraction
abstract
How to identify semantic relations among entities in a document when only a few labeled documents are available?Few-shot documentlevel relation extraction (FSDLRE) is crucial for addressing the pervasive data scarcity problem in real-world scenarios.Metric-based meta-learning is an effective framework widely adopted for FSDLRE, which constructs class prototypes for classification.However, existing works often struggle to obtain class prototypes with accurate relational semantics: 1) To build prototype for a target relation type, they aggregate the representations of all entity pairs holding that relation, while these entity pairs may also hold other relations, thus disturbing the prototype.2) They use a set of generic NOTA (none-of-the-above) prototypes across all tasks, neglecting that the NOTA semantics differs in tasks with different target relation types.In this paper, we propose a relation-aware prototype learning method for FSDLRE to strengthen the relational semantics of prototype representations.By judiciously leveraging the relation descriptions and realistic NOTA instances as guidance, our method effectively refines the relation prototypes and generates task-specific NOTA prototypes.Extensive experiments demonstrate that our method outperforms state-of-the-art approaches by average 2.61% F 1 across various settings of two FSDLRE benchmarks. 1
Shiao Meng, Xuming Hu, Aiwei Liu, Shuang Li 0015, Fukun Ma, Yawen Yang, Lijie Wen 0001
EMNLP3
2023 Gaussian Prior Reinforcement Learning for Nested Named Entity Recognition
abstract
Named Entity Recognition (NER) is a well and widely studied task in natural language processing. Recently, the nested NER has attracted more attention since its practicality and difficulty. Existing works for nested NER ignore the recognition order and boundary position relation of nested entities. To address these issues, we propose a novel seq2seq model named GPRL, which formulates the nested NER task as an entity triplet sequence generation process. GPRL adopts the reinforcement learning method to generate entity triplets de-coupling the entity order in gold labels and expects to learn a reasonable recognition order of entities via trial and error. Based on statistics of boundary distance for nested entities, GPRL designs a Gaussian prior to represent the boundary distance distribution between nested entities and adjust the out-put probability distribution of nested boundary tokens. Experiments on three nested NER datasets demonstrate that GPRL outperforms previous nested NER models.
Yawen Yang, Xuming Hu, Fukun Ma, Shuang Li 0015, Aiwei Liu, Lijie Wen 0001, Philip S. Yu
ICASSP5
2023 Prompt Me Up: Unleashing the Power of Alignments for Multimodal Entity and Relation Extraction
abstract
How can we better extract entities and relations from text? Using multimodal extraction with images and text obtains more signals for entities and relations, and aligns them through graphs or hierarchical fusion, aiding in extraction. Despite attempts at various fusions, previous works have overlooked many unlabeled image-caption pairs, such as NewsCLIPing. This paper proposes innovative pre-training objectives for entity-object and relation-image alignment, extracting objects from images and aligning them with entity and relation prompts for soft pseudo-labels. These labels are used as self-supervised signals for pre-training, enhancing the ability to extract entities and relations. Experiments on three datasets show an average 3.41% F1 improvement over prior SOTA. Additionally, our method is orthogonal to previous multimodal fusions, and using it on prior SOTA fusions further improves 5.47% F1.
Xuming Hu, Junzhe Chen 0001, Aiwei Liu, Shiao Meng, Lijie Wen 0001, Philip S. Yu
ACM Multimedia3
2023 A Multi-Level Supervised Contrastive Learning Framework for Low-Resource Natural Language Inference
abstract
Natural Language Inference (NLI) is a growingly essential task in natural language understanding, which requires inferring the relationship between the sentence pairs (premiseandhypothesis). Recently, low-resource natural language inference has gained increasing attention, due to significant savings in manual annotation costs and a better fit with real-world scenarios. Existing works fail to characterize discriminative representations between different classes with limited training data, which may cause faults in label prediction. Here we propose a multi-level supervised contrastive learning framework named MultiSCL for low-resource natural language inference. MultiSCL leverages a sentence-level and pair-level contrastive learning objective to discriminate between different classes of sentence pairs by bringing those in one class together and pushing away those in different classes. MultiSCL adopts a data augmentation module that generates different views for input samples to better learn the latent representation. The pair-level representation is obtained from a cross attention module. We conduct extensive experiments on two public NLI datasets in low-resource settings, and the accuracy of MultiSCL exceeds other models by 1.8%, 3.1% and 4.1% on SNLI, MNLI and Sick with 5 instances per label respectively. Moreover, our method outperforms the previous state-of-the-art method on cross-domain tasks of text classification.
Shuang Li 0015, Xuming Hu, Li Lin 0011, Aiwei Liu, Lijie Wen 0001, Philip S. Yu
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords Substitution
abstract
We propose the first character-level white-box adversarial attack method against transformer models.The intuition of our method comes from the observation that words are split into subtokens before being fed into the transformer models and the substitution between two close subtokens has a similar effect to the character modification.Our method mainly contains three steps.First, a gradient-based method is adopted to find the most vulnerable words in the sentence.Then we split the selected words into subtokens to replace the origin tokenization result from the transformer tokenizer.Finally, we utilize an adversarial loss to guide the substitution of attachable subtokens in which the Gumbel-softmax trick is introduced to ensure gradient propagation.Meanwhile, we introduce the visual and length constraint in the optimization process to achieve minimum character modifications.Extensive experiments on both sentence-level and token-level tasks demonstrate that our method could outperform the previous attack methods in terms of success rate and edit distance.Furthermore, human evaluation verifies our adversarial examples could preserve their origin labels.
Aiwei Liu, Honghai Yu, Xuming Hu, Shuang Li 0015, Li Lin 0011, Fukun Ma, Yawen Yang, Lijie Wen 0001
EMNLP1
2022 Semantic Enhanced Text-to-SQL Parsing via Iteratively Learning Schema Linking Graph
abstract
The generalizability to new databases is of vital importance to Text-to-SQL systems which aim to parse human utterances into SQL statements. Existing works achieve this goal by leveraging the exact matching method to identify the lexical matching between the question words and the schema items. However, these methods fail in other challenging scenarios, such as the synonym substitution in which the surface form differs between the corresponding question words and schema items. In this paper, we propose a framework named ISESL-SQL to iteratively build a semantic enhanced schema-linking graph between question tokens and database schemas. First, we extract a schema linking graph from PLMs through a probing procedure in an unsupervised manner. Then the schema linking graph is further optimized during the training process through a deep graph learning method. Meanwhile, we also design an auxiliary task called graph regularization to improve the schema information mentioned in the schema-linking graph. Extensive experiments on three benchmarks demonstrate that ISESL-SQL could consistently outperform the baselines and further investigations show its generalizability and robustness.
Aiwei Liu, Xuming Hu, Li Lin 0011, Lijie Wen 0001
KDD1
2022 CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking
abstract
Xuming Hu, Zhijiang Guo, GuanYu Wu, Aiwei Liu, Lijie Wen, Philip Yu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xuming Hu, Zhijiang Guo, Guanyu Wu, Aiwei Liu, Lijie Wen 0001, Philip S. Yu
NAACL-HLT4