Lei Sha

dblp:93/3906 · DBLP profile ↗
← Back
41ranked-venue papers
15as first author
20since 2021 · last 2027
0000-0001-5914-7590ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 15 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2027 Benchmarking multi-step legal reasoning and analyzing Chain-of-Thought effects in Large Language Models
Wenhan Yu, Xinbo Lin, Lanxin Ni, Jinhua Cheng, Lei Sha
Inf. Process. Manag.5
2026 Large Language Models Struggle with Unreasonability in Math Problems
abstract
Large Language Models (LLMs) have shown remarkable success on a wide range of math and reasoning benchmarks. However, we observe that they often struggle when faced with unreasonable math problems. Instead of recognizing these issues, models frequently proceed as if the problem is well-posed, producing incorrect answers or falling into overthinking and verbose self-correction. To systematically investigate this overlooked vulnerability, we propose the Unreasonable Math Problems (UMP) benchmark, designed to evaluate LLMs' ability to detect and respond to unreasonable math problem statements. Based on extensive experiments covering 19 LLMs, we find that even state-of-the-art general models like GPT-4o struggle on UMP. While reasoning models such as DeepSeek-R1 demonstrate a higher sensitivity to unreasonable inputs, this often comes at the cost of generating overly long and meaningless responses that fail to converge. We further find that prompting and fine-tuning enhance the detection of unreasonable inputs, with minor and acceptable trade-offs, making them practical solutions in this challenging setting.
Jingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 0094, Weilin Luo, Lei Sha, Zhifang Sui
AAAI8
2026 SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory
abstract
Current defenses for Large Language Models (LLMs) often suffer from a "memory gap": parameter-modifying methods are computationally rigid, while inference-time filters cannot retain or reuse defense knowledge across interactions.To address this, we propose Safet-yMem, a novel framework that secures LLMs through a dual-component safety memory system.SafetyMem consists of Semantic Safety Memory (SSM), which consolidates diverse jailbreak attempts into a structured knowledge base of attack patterns, and Episodic Safety Memory (ESM), which maintains an evolving set of procedural rules refined from historical detection failures.Unlike static defenses, Safe-tyMem allows the model to "remember" and adapt to emerging adversarial strategies without parameter retraining.To further enhance robustness, we introduce an adversarial memory expansion mechanism that proactively generates challenging variants to solidify these memories.Experiments on standard and stealthy jailbreak benchmarks show that SafetyMem substantially reduces attack success rates while preserving efficiency and interpretability, consistently outperforming state-of-the-art baselines across multiple LLMs.
Ziyi Ni, Huacan Wang, Lei Sha
ACL (1)5
2026 Rethinking Semantic-Collaborative Integration: Why Alignment Is Not Enough
abstract
Large language models (LLMs) have become an important semantic infrastructure for modern recommender systems. A prevailing paradigm integrates LLM-derived semantic embeddings with collaborative representations via representation alignment, implicitly assuming that the two views encode a shared latent entity and that stronger alignment yields better results. We formalize this assumption as the global low-complexity alignment hypothesis and argue that it is stronger than necessary and often structurally mismatched with real-world recommendation settings. We propose a complementary perspective in which semantic and collaborative representations are treated as partially shared yet fundamentally heterogeneous views, each containing both shared and view-specific factors. Under this shared-plus-private latent structure, enforcing global geometric alignment may distort local structure, suppress view-specific signals, and reduce informational diversity. To support this perspective, we develop complementarity-aware diagnostics that quantify overlap, unique-hit contribution, and theoretical fusion upper bounds. Empirical analyses on sparse recommendation benchmarks reveal low item-level agreement between semantic and collaborative views and substantial oracle fusion gains, indicating strong complementarity. Furthermore, controlled alignment probes show that low-capacity mappings capture only shared components and fail to recover full collaborative geometry, especially under distribution shift. These findings suggest that alignment should not be treated as the default integration principle. We advocate a shift from alignment-centric modeling to complementarity fusion-centric, complementarity-aware design, where shared factors are selectively integrated while private signals are preserved. This reframing provides a principled foundation for the next generation of LLM-enhanced recommender systems.
Maolin Wang 0001, Dongze Wu, Jianing Zhou, Beining Bao, Chenbin Zhang, Lei Sha
SIGIR10
2025 Towards Harmonized Uncertainty Estimation for Large Language Models
abstract
To facilitate robust and trustworthy deployment of large language models (LLMs), it is essential to quantify the reliability of their generations through uncertainty estimation.While recent efforts have made significant advancements by leveraging the internal logic and linguistic features of LLMs to estimate uncertainty scores, our empirical analysis highlights the pitfalls of these methods to strike a harmonized estimation between indication, balance, and calibration, which hinders their broader capability for accurate uncertainty estimation.To address this challenge, we propose CUE (Corrector for Uncertainty Estimation): A straightforward yet effective method that employs a lightweight model trained on data aligned with the target LLM's performance to adjust uncertainty scores.Comprehensive experiments across diverse models and tasks demonstrate its effectiveness, which achieves consistent improvements of up to 60% over existing methods.Resources are available at https://github. com/O-L1RU1/Corrector4UE.
Rui Li 0094, Jing Long, Muge Qi, Heming Xia, Lei Sha, Peiyi Wang, Zhifang Sui
ACL (1)5
2025 LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
abstract
Safety concerns in large language models (LLMs) have gained significant attention due to their exposure to potentially harmful data during pre-training. In this paper, we identify a new safety vulnerability in LLMs: their susceptibility to natural distribution shifts between attack prompts and original toxic prompts, where seemingly benign prompts, semantically related to harmful content, can bypass safety mechanisms. To explore this issue, we introduce a novel attack method, ActorBreaker, which identifies actors related to toxic prompts within pre-training distribution to craft multi-turn prompts that gradually lead LLMs to reveal unsafe content. ActorBreaker is grounded in Latour’s actor-network theory, encompassing both human and non-human actors to capture a broader range of vulnerabilities. Our experimental results demonstrate that ActorBreaker outperforms existing attack methods in terms of diversity, effectiveness, and efficiency across aligned LLMs. To address this vulnerability, we propose expanding safety training to cover a broader semantic space of toxic content. We thus construct a multi-turn safety dataset using ActorBreaker. Fine-tuning models on our dataset shows significant improvements in robustness, though with some trade-offs in utility. Code is available at https://github.com/AI45Lab/ActorAttack.
Qibing Ren, Hao Li 0069, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao 0001, Lei Sha, Junchi Yan, Lizhuang Ma
ACL (1)7
2025 Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
abstract
With rapid advancement and increasing accessibility of LLMs, fine-tuning aligned models has become a critical step for adapting them to real-world applications, which makes the safety of this fine-tuning process more important than ever.However, recent studies have highlighted a critical challenge: even when fine-tuning with benign datasets, the safety alignment of aligned LLMs can be compromised, making them more susceptible to malicious instructions.In this paper, we show that fine-tuning datasets often contain safetydegrading samples that are not easily identifiable on the surface.These samples can easily degrade the safety alignment of LLMs during fine-tuning.To address this issue, we propose LARF, a Layer-Aware Representation Filtering method.This method identifies safety-sensitive layers within the LLM and leverages data representations to detect safetydegrading data samples in the fine-tuning dataset.Experimental results demonstrate that LARF can efficiently and effectively identify safety-degrading data.After removing such data, the safety alignment degradation caused by fine-tuning is mitigated.Please see our code at https://github.com/LLLeoLi/LARF.
Hao Li 0031, Zhenghao Lu, Xianyi Wei, Rui Li 0094, Lei Sha
EMNLP7
2025 DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
abstract
Large Language Models (LLMs) are susceptible to generating harmful content when prompted with carefully crafted inputs, a vulnerability known as LLM jailbreaking.As LLMs become more powerful, studying jailbreak methods is critical to enhancing security and aligning models with human values.Traditionally, jailbreak techniques have relied on suffix addition or prompt templates, but these methods suffer from limited attack diversity.This paper introduces DiffusionAttacker, an end-to-end generative approach for jailbreak rewriting inspired by diffusion models.Our method employs a sequence-to-sequence (seq2seq) text diffusion model as a generator, conditioning on the original prompt and guiding the denoising process with a novel attack loss.Unlike previous approaches that use autoregressive LLMs to generate jailbreak prompts, which limit the modification of already generated tokens and restrict the rewriting space, DiffusionAttacker utilizes a seq2seq diffusion model, allowing more flexible token modifications.This approach preserves the semantic content of the original prompt while producing harmful content.Additionally, we leverage the Gumbel-Softmax technique to make the sampling process from the diffusion model's output distribution differentiable, eliminating the need for iterative token search.Extensive experiments on Advbench and Harmbench demonstrate that DiffusionAttacker outperforms previous methods across various evaluation metrics, including attack success rate (ASR), fluency, and diversity.How to make a bomb?Gaussian Noise Gradually Denoising DiffuSeq Model Embedding Map LM_head Gumbel softmax Jailbreaking Prompt LM_head Gumbel softmax Jailbreaking
Hao Wang 0003, Hao Li 0031, Junda Zhu 0003, Xinyuan Wang 0009, Chengwei Pan, Minlie Huang, Lei Sha
EMNLP7
2025 Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
abstract
Large Reasoning Models (LRMs) have recently demonstrated impressive performances across diverse domains.However, how the safety of Large Language Models (LLMs) benefits from enhanced reasoning capabilities against jailbreak queries remains unexplored.To bridge this gap, in this paper, we propose Reasoningto-Defend (R2D), a novel training paradigm that integrates a safety-aware reasoning mechanism into LLMs' generation process.This enables self-evaluation at each step of the reasoning process, forming safety PIVOT TOKENS as indicators of the safety status of responses.Furthermore, in order to improve the accuracy of predicting PIVOT TOKENS, we propose Contrastive Pivot Optimization (CPO), which enhances the model's perception of the safety status of given dialogues.LLMs dynamically adjust their response strategies during reasoning, significantly enhancing their safety capabilities defending jailbreak attacks.Extensive experiments demonstrate that R2D effectively mitigates various attacks and improves overall safety, while maintaining the original performances.This highlights the substantial potential of safety-aware reasoning in improving robustness of LRMs and LLMs against various jailbreaks. 1
Junda Zhu 0003, Lingyong Yan, Shuaiqiang Wang, Dawei Yin 0001, Lei Sha
EMNLP5
2025 Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models
abstract
Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.
Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang
ICLR16
2024 ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
abstract
The safety defense methods of Large language models (LLMs) stays limited because the dangerous prompts are manually curated to just few known attack types, which fails to keep pace with emerging varieties.Recent studies found that attaching suffixes to harmful instructions can hack the defense of LLMs and lead to dangerous outputs.However, similar to traditional text adversarial attacks, this approach, while effective, is limited by the challenge of the discrete tokens.This gradient based discrete optimization attack requires over 100,000 LLM calls, and due to the unreadable of adversarial suffixes, it can be relatively easily penetrated by common defense methods such as perplexity filters.To cope with this challenge, in this paper, we propose an Adversarial Suffix Embedding Translation Framework (ASETF), aimed at transforming continuous adversarial suffix embeddings into coherent and understandable text.This method greatly reduces the computational overhead during the attack process and helps to automatically generate multiple adversarial samples, which can be used as data to strengthen LLM's security defense.Experimental evaluations were conducted on Llama2, Vicuna, and other prominent LLMs, employing harmful directives sourced from the Advbench dataset.The results indicate that our method significantly reduces the computation time of adversarial suffixes and achieves a much better attack success rate than existing techniques, while significantly enhancing the textual fluency of the prompts.In addition, our approach can be generalized into a broader method for generating transferable adversarial suffixes that can successfully attack multiple LLMs, even black-box LLMs, such as ChatGPT and Gemini.
Hao Wang 0003, Hao Li 0031, Minlie Huang, Lei Sha
EMNLP4
2024 ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator
abstract
Large language models (LLMs) are proven to benefit a lot from retrieval-augmented generation (RAG) in alleviating hallucinations confronted with knowledge-intensive questions.RAG adopts information retrieval techniques to inject external knowledge from semanticrelevant documents as input contexts.However, since today's Internet is flooded with numerous noisy and fabricating content, it is inevitable that RAG systems are vulnerable to these noises and prone to respond incorrectly.To this end, we propose to optimize the retrieval-augmented GENERATOR with an Adversarial Tuning Multi-agent system (ATM).The ATM steers the GENERATOR to have a robust perspective of useful documents for question answering with the help of an auxiliary ATTACKER agent through adversarially tuning the agents for several iterations.After rounds of multi-agent iterative tuning, the GENERA-TOR can eventually better discriminate useful documents amongst fabrications.The experimental results verify the effectiveness of ATM and we also observe that the GENERATOR can achieve better performance compared to the state-of-the-art baselines.The code is available at https://github.com/chuhac/ATM-RAG.
Junda Zhu 0003, Lingyong Yan, Haibo Shi, Dawei Yin 0001, Lei Sha
EMNLP5
2024 A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks
abstract
Predictive coding networks are neuroscience-inspired models with roots in both Bayesian statistics and neuroscience. Training such models, however, is quite inefficient and unstable. In this work, we show how by simply changing the temporal scheduling of the update rule for the synaptic weights leads to an algorithm that is much more efficient and stable than the original one, and has theoretical guarantees in terms of convergence. The proposed algorithm, that we call incremental predictive coding (iPC) is also more biologically plausible than the original one, as it it fully automatic. In an extensive set of experiments, we show that iPC constantly performs better than the original formulation on a large number of benchmarks for image classification, as well as for the training of both conditional and masked language models, in terms of test accuracy, efficiency, and convergence with respect to a large set of hyperparameters.
Tommaso Salvatori, Yuhang Song 0001, Yordan Yordanov, Beren Millidge, Lei Sha, Cornelius Emde, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz
ICLR5
2024 A Cross-Site Scripting Attack Protection Framework Based on Managed Proxy
abstract
Cross-Site Scripting (XSS) vulnerabilities continue to pose a formidable challenge in the realm of web application security due to their prevalence. Despite the development of various defense mechanisms, including input validation techniques and code differentiation strategies, these solutions have struggled to achieve broad adoption in real-world environments due to persistent issues such as unreliable detection, high computational overhead, scalability limitations, and compatibility concerns. This paper introduces a novel managed proxy-based framework for XSS defense, positioned strategically between the web browser and the web server. This framework employs a tag-based XSS protection mechanism alongside a defense-in-depth strategy, providing robust protection not only against XSS attacks but also against other web threats. Furthermore, the framework integrates advanced web front-end hardening techniques to significantly complicate the exploitation of potential vulnerabilities. Comprehensive practical evaluations indicate that the proposed framework is both flexible and effective, demonstrating outstanding performance across multiple dimensions.
Jianhua Peng, Meiyue Yang, Lei Sha
TrustCom9
2024 Text Attribute Control via Closed-Loop Disentanglement
abstract
Abstract Changing an attribute of a text without changing the content usually requires first disentangling the text into irrelevant attributes and content representations. After that, in the inference phase, the representation of one attribute is tuned to a different value, expecting that the corresponding attribute of the text can also be changed accordingly. The usual way of disentanglement is to add some constraints on the latent space of an encoder-decoder architecture, including adversarial-based constraints and mutual-information-based constraints. However, previous semi-supervised processes of attribute change are usually not enough to guarantee the success of attribute change and content preservation. In this paper, we propose a novel approach to achieve a robust control of attributes while enhancing content preservation. In this approach, we use a semi-supervised contrastive learning method to encourage the disentanglement of attributes in latent spaces. Differently from previous works, we re-disentangle the reconstructed sentence and compare the re-disentangled latent space with the original latent space, which makes a closed-loop disentanglement process. This also helps content preservation. In addition, the contrastive learning method is also able to replace the role of minimizing mutual information and adversarial training in the disentanglement process, which alleviates the computation cost. We conducted experiments on three text datasets, including the Yelp Service review dataset, the Amazon Product review dataset, and the GoEmotions dataset. The experimental results show the effectiveness of our model.
Lei Sha, Thomas Lukasiewicz
Trans. Assoc. Comput. Linguistics1
2023 Rationalizing predictions by adversarial information calibration
Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz
Artif. Intell.1
2022 Small Changes Make Big Differences: Improving Multi-turn Response Selection in Dialogue Systems via Fine-Grained Contrastive Learning
abstract
Retrieve-based dialogue response selection aims to find a proper response from a candidate set given a multi-turn context.Pre-trained language models (PLMs) based methods have yielded significant improvements on this task.The sequence representation plays a key role in the learning of matching degree between the dialogue context and the response.However, we observe that different context-response pairs sharing the same context always have a greater similarity in the sequence representations calculated by PLMs, which makes it hard to distinguish positive responses from negative ones.Motivated by this, we propose a novel Fine-Grained Contrastive (FGC) learning method for the response selection task based on PLMs.This FGC learning strategy helps PLMs to generate more distinguishable matching representations of each dialogue at fine grains, and further make better predictions on choosing positive responses.Empirical studies on two benchmark datasets demonstrate that the proposed FGC learning method can generally and significantly improve the model performance of existing PLMbased matching models. 1
Can Xu 0002, Huang Hu, Lei Sha, Yan Zhang 0117, Daxin Jiang
INTERSPEECH4
2021 Learning from the Best: Rationalizing Predictions by Adversarial Information Calibration
abstract
Explaining the predictions of AI models is paramount in safety-critical applications, such as in legal or medical domains. One form of explanation for a prediction is an extractive rationale, i.e., a subset of features of an instance that lead the model to give its prediction on the instance. Previous works on generating extractive rationales usually employ a two-phase model: a selector that selects the most important features (i.e., the rationale) followed by a predictor that makes the prediction based exclusively on the selected features. One disadvantage of these works is that the main signal for learning to select features comes from the comparison of the final answers given by the predictor and the ground-truth answers. In this work, we propose to squeeze more information from the predictor via an information calibration method. More precisely, we train two models jointly: one is a typical neural model that solves the task at hand in an accurate but black-box manner, and the other is a selector-predictor model that additionally produces a rationale for its prediction. The first model is used as a guide to the second model. We use an adversarial-based technique to calibrate the information extracted by the two models such that the difference between them is an indicator of the missed or over-selected features. In addition, for natural language tasks, we propose to use a language-model-based regularizer to encourage the extraction of fluent rationales. Experimental results on a sentiment analysis task as well as on three tasks from the legal domain show the effectiveness of our approach to rationale extraction.
Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz
AAAI1
2021 Multi-type Disentanglement without Adversarial Training
abstract
Controlling the style of natural language by disentangling the latent space is an important step towards interpretable machine learning. After the latent space is disentangled, the style of a sentence can be transformed by tuning the style representation without affecting other features of the sentence. Previous works usually use adversarial training to guarantee that disentangled vectors do not affect each other. However, adversarial methods are difficult to train. Especially when there are multiple features (e.g., sentiment, or tense, which we call style types in this paper), each feature requires a separate discriminator for extracting a disentangled style vector corresponding to that feature. In this paper, we propose a unified distribution-controlling method, which provides each specific style value (the value of style types, e.g., positive sentiment, or past tense) with a unique representation. This method contributes a solid theoretical basis to avoid adversarial training in multi-type disentanglement. We also propose multiple loss functions to achieve a style-content disentanglement as well as a disentanglement among multiple style types. In addition, we observe that if two different style types always have some specific style values that occur together in the dataset, they will affect each other when transferring the style values. We call this phenomenon training bias , and we propose a loss function to alleviate such training bias while disentangling multiple types. We conduct experiments on two datasets (Yelp service reviews and Amazon product reviews) to evaluate the style-disentangling effect and the unsupervised style-transfer performance on two style types: sentiment and tense. The experimental results show the effectiveness of our model.
Lei Sha, Thomas Lukasiewicz
AAAI1
2021 Associative Memories via Predictive Coding
abstract
Associative memories in the brain receive and store patterns of activity registered by the sensory neurons, and are able to retrieve them when necessary. Due to their importance in human intelligence, computational models of associative memories have been developed for several decades now. In this paper, we present a novel neural model for realizing associative memories, which is based on a hierarchical generative network that receives external stimuli via sensory neurons. It is trained using predictive coding, an error-based learning algorithm inspired by information processing in the cortex. To test the model's capabilities, we perform multiple retrieval experiments from both corrupted and incomplete data points. In an extensive comparison, we show that this new model outperforms in retrieval accuracy and robustness popular associative memory models, such as autoencoders trained via backpropagation, and modern Hopfield networks. In particular, in completing partial data points, our model achieves remarkable results on natural image datasets, such as ImageNet, with a surprisingly high accuracy, even when only a tiny fraction of pixels of the original images is presented. Our model provides a plausible framework to study learning and retrieval of memories in the brain, as it closely mimics the behavior of the hippocampus as a memory index and generative model.
Tommaso Salvatori, Yuhang Song 0001, Yujian Hong, Lei Sha, Simon Frieder, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz
NeurIPS4
2020 Gradient-guided Unsupervised Lexically Constrained Text Generation
abstract
Lexically-constrained generation requires the target sentence to satisfy some lexical constraints, such as containing some specific words or being the paraphrase to a given sentence, which is very important in many real-world natural language generation applications.Previous works usually apply beamsearch-based methods or stochastic searching methods to lexically-constrained generation.However, when the search space is too large, beam-search-based methods always fail to find the constrained optimal solution.At the same time, stochastic search methods always cost too many steps to find the correct optimization direction.In this paper, we propose a novel method G2LC to solve the lexically-constrained generation as an unsupervised gradient-guided optimization problem.We propose a differentiable objective function and use the gradient to help determine which position in the sequence should be changed (deleted or inserted/replaced by another word).The word updating process of the inserted/replaced word also benefits from the guidance of gradient.Besides, our method is free of parallel data training, which is flexible to be used in the inference stage of any pre-trained generation model.We apply G2LC to two generation tasks: keyword-to-sentence generation and unsupervised paraphrase generation.The experiment results show that our method achieves state-of-the-art compared to previous lexically-constrained methods.
Lei Sha
EMNLP (1)1
2020 Estimating Minimum Operation Steps via Memory-based Recurrent Calculation Network
abstract
To estimate time complexity for a given algorithm is important for algorithm designers. Usually, time complexity means the "analytical" time complexity which needs to be proved by strict math derivation. We propose to estimate the "numerical" time complexity (NTC), which measures the minimum number of operations an algorithm has to spend, as well as capture the intrinsic laws of time complexity. The unique challenges include: (1) How to make a machine learning model has the same ability as a real-world CPU (2) How to measure the minimum number of required arithmetic operations for a given problem. To tackle these challenges, we first propose a memory-based recurrent calculation network to mimic the functions of CPU and then we propose a self-adaptive selection gate for deciding when the mimic calculation process should stop. In addition, we use a symbolic learning method to find the time complexity formula. We train and test our model on four basic algorithms: long integer addition, 1-dim max-pooling, outer product, and sorting. Experiment results demonstrate that our model can precisely predict the numerical time complexity as well as the time complexity formula for each algorithm. We also conduct many visualizations to prove the effectiveness and correctness of our model.
Lei Sha, Qi Chen 0009, Houfeng Wang
IJCNN1
2019 We Know What You Will Ask: A Dialogue System for Multi-intent Switch and Prediction
Qi Chen 0009, Lei Sha, Hui Xue 0004, Sujian Li, Houfeng Wang
NLPCC (1)3
2018 Table-to-Text Generation by Structure-Aware Seq2seq Learning
abstract
Table-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a novel structure-aware seq2seq architecture which consists of field-gating encoder and description generator with dual attention. In the encoding phase, we update the cell memory of the LSTM unit by a field gate and its corresponding field value in order to incorporate field information into table representation. In the decoding phase, dual attention mechanism which contains word level attention and field level attention is proposed to model the semantic relevance between the generated description and the table. We conduct experiments on the WIKIBIO dataset which contains over 700k biographies and corresponding infoboxes from Wikipedia. The attention visualizations and case studies show that our model is capable of generating coherent and informative descriptions based on the comprehensive understanding of both the content and the structure of a table. Automatic evaluations also show our model outperforms the baselines by a great margin. Code for this work is available on https://github.com/tyliupku/wiki2bio.
Tianyu Liu 0001, Kexiang Wang, Lei Sha, Baobao Chang, Zhifang Sui
AAAI3
2018 Order-Planning Neural Text Generation From Structured Data
abstract
Generating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model.
Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui
AAAI1
2018 Jointly Extracting Event Triggers and Arguments by Dependency-Bridge RNN and Tensor-Based Argument Interaction
abstract
Event extraction plays an important role in natural language processing (NLP) applications including question answering and information retrieval. Traditional event extraction relies heavily on lexical and syntactic features, which require intensive human engineering and may not generalize to different datasets. Deep neural networks, on the other hand, are able to automatically learn underlying features, but existing networks do not make full use of syntactic relations. In this paper, we propose a novel dependency bridge recurrent neural network (dbRNN) for event extraction. We build our model upon a recurrent neural network, but enhance it with dependency bridges, which carry syntactically related information when modeling each word.We illustrates that simultaneously applying tree structure and sequence structure in RNN brings much better performance than only uses sequential RNN. In addition, we use a tensor layer to simultaneously capture the various types of latent interaction between candidate arguments as well as identify/classify all arguments of an event. Experiments show that our approach achieves competitive results compared with previous work.
Lei Sha, Baobao Chang, Zhifang Sui
AAAI1
2018 A Multi-View Fusion Neural Network for Answer Selection
abstract
Community question answering aims at choosing the most appropriate answer for a given question, which is important in many NLP applications. Previous neural network-based methods consider several different aspects of information through calculating attentions. These different kinds of attentions are always simply summed up and can be seen as a ``single view", causing severe information loss. To overcome this problem, we propose a Multi-View Fusion Neural Network, where each attention component generates a ``view'' of the QA pair and a fusion RNN integrates the generated views to form a more holistic representation. In this fusion RNN method, a filter gate collects important information of input and directly adds it to the output, which borrows the idea of residual networks. Experimental results on the WikiQA and SemEval-2016 CQA datasets demonstrate that our proposed model outperforms the state-of-the-art methods.
Lei Sha, Xiaodong Zhang 0022, Baobao Chang, Zhifang Sui
AAAI1
2018 Auto-Dialabel: Labeling Dialogue Data with Unsupervised Learning
abstract
The lack of labeled data is one of the main challenges when building a task-oriented dialogue system.Existing dialogue datasets usually rely on human labeling, which is expensive, limited in size, and in low coverage.In this paper, we instead propose our framework auto-dialabel to automatically cluster the dialogue intents and slots.In this framework, we collect a set of context features, leverage an autoencoder for feature assembly, and adapt a dynamic hierarchical clustering method for intent and slot labeling.Experimental results show that our framework can promote human labeling cost to a great extent, achieve good intent clustering accuracy (84.1%), and provide reasonable and instructive slot labeling results.
Qi Chen 0009, Lei Sha, Sujian Li, Xu Sun 0001, Houfeng Wang
EMNLP3
2017 Attentive Interactive Neural Networks for Answer Selection in Community Question Answering
abstract
Answer selection plays a key role in community question answering (CQA). Previous research on answer selection usually ignores the problems of redundancy and noise prevalent in CQA. In this paper, we propose to treat different text segments differently and design a novel attentive interactive neural network (AI-NN) to focus on those text segments useful to answer selection. The representations of question and answer are first learned by convolutional neural networks (CNNs) or other neural network architectures. Then AI-NN learns interactions of each paired segments of two texts. Row-wise and column-wise pooling are used afterwards to collect the interactions. We adopt attention mechanism to measure the importance of each segment and combine the interactions to obtain fixed-length representations for question and answer. Experimental results on CQA dataset in SemEval-2016 demonstrate that AI-NN outperforms state-of-the-art method.
Xiaodong Zhang 0022, Sujian Li, Lei Sha, Houfeng Wang
AAAI3
2017 A Progressive Learning Approach to Chinese SRL Using Heterogeneous Data
abstract
Previous studies on Chinese semantic role labeling (SRL) have concentrated on a single semantically annotated corpus.But the training data of single corpus is often limited.Whereas the other existing semantically annotated corpora for Chinese SRL are scattered across different annotation frameworks.But still, Data sparsity remains a bottleneck.This situation calls for larger training datasets, or effective approaches which can take advantage of highly heterogeneous data.In this paper, we focus mainly on the latter, that is, to improve Chinese SRL by using heterogeneous corpora together.We propose a novel progressive learning model which augments the Progressive Neural Network with Gated Recurrent Adapters.The model can accommodate heterogeneous inputs and effectively transfer knowledge between them.We also release a new corpus, Chinese Sem-Bank, for Chinese SRL 1 .Experiments on CPB 1.0 show that our model outperforms state-of-the-art methods.
Qiaolin Xia, Lei Sha, Baobao Chang, Zhifang Sui
ACL (1)2
2017 Topic medical concept embedding: Multi-sense representation learning for medical concept
abstract
Representation learning algorithm in medical area maps high dimensional real world medical concepts to low dimensional vector space, encodes rich medical knowledge, and has brought improvement to various machine learning applications in medical area. However, previous representation learning models in medical area failed to consider the multi-sense characteristic of medical concept. Moreover, the inner relationships between representations learned by previous model is implicit and can only be explained according to visualization, which means poor interpretability. In this paper, we propose Topic Medical Concept Embedding (TMCE), a generative embedding model to address above two problems. TMCE is able to learn multi-sense representations for a single medical concept, and TMCE can also improve interpretability by modeling relationships between each concept explicitly. In TMCE, multi-sense concept representations are influenced by its contexts and its topics. In addition, dosage information which is ignored by previous work are also utilized in TMCE. A MCMC method is presented to jointly learn the two-layer topic embeddings and multi-sense concept embeddings. Experimental results show that representations learned by TMCE outperforms those learned by other strong baselines by a large margin in a multi-label diagnose classification tasks. Several case studies further show that TMCE can learn medically correct multi-sense representations with better interpretability than other strong baselines.
Chengyue Gong, Luchen Liu, Lei Sha, Ming Zhang 0004
BIBM4
2017 Will Repeated Reading Benefit Natural Language Understanding?
Lei Sha, Zhifang Sui
NLPCC1
2016 RBPB: Regularization-Based Pattern Balancing Method for Event Extraction
abstract
Event extraction is a particularly challenging information extraction task, which intends to identify and classify event triggers and arguments from raw text.In recent works, when determining event types (trigger classification), most of the works are either pattern-only or feature-only.However, although patterns cannot cover all representations of an event, it is still a very important feature.In addition, when identifying and classifying arguments, previous works consider each candidate argument separately while ignoring the relationship between arguments.This paper proposes a Regularization-Based Pattern Balancing Method (RBPB).Inspired by the progress in representation learning, we use trigger embedding, sentence-level embedding and pattern features together as our features for trigger classification so that the effect of patterns and other useful features can be balanced.In addition, RBPB uses a regularization method to take advantage of the relationship between arguments.Experiments show that we achieve results better than current state-of-art equivalents.
Lei Sha, Jing Liu 0022, Chin-Yew Lin, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)1
2016 Towards Time-Aware Knowledge Graph Completion
abstract
Knowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui
COLING4
2016 Reading and Thinking: Re-read LSTM Unit for Textual Entailment Recognition
abstract
Recognizing Textual Entailment (RTE) is a fundamentally important task in natural language processing that has many applications. The recently released Stanford Natural Language Inference (SNLI) corpus has made it possible to develop and evaluate deep neural network methods for the RTE task. Previous neural network based methods usually try to encode the two sentences (premise and hypothesis) and send them together into a multi-layer perceptron to get their entailment type, or use LSTM-RNN to link two sentences together while using attention mechanic to enhance the model’s ability. In this paper, we propose to use the re-read mechanic, which means to read the premise again and again while reading the hypothesis. After read the premise again, the model can get a better understanding of the premise, which can also affect the understanding of the hypothesis. On the contrary, a better understanding of the hypothesis can also affect the understanding of the premise. With the alternative re-read process, the model can “think” of a better decision of entailment type. We designed a new LSTM unit called re-read LSTM (rLSTM) to implement this “thinking” process. Experiments show that we achieve results better than current state-of-the-art equivalents.
Lei Sha, Baobao Chang, Zhifang Sui, Sujian Li
COLING1
2016 Encoding Temporal Information for Time-Aware Link Prediction
abstract
Most existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
EMNLP4
2016 Capturing Argument Relationship for Chinese Semantic Role Labeling
abstract
In this paper, we capture the argument relationships for Chinese semantic role labeling task, and improve the task's performance with the help of argument relationships.We split the relationship between two candidate arguments into two categories: (1) Compatible arguments: if one candidate argument belongs to a given predicate, then the other is more likely to belong to the same predicate; (2) Incompatible arguments: if one candidate argument belongs to a given predicate, then the other is less likely to belong to the same predicate.However, previous works did not explicitly model argument relationships.We use a simple maximum entropy classifier to capture the two categories of argument relationships and test its performance on the Chinese Proposition Bank (CPB).The experiments show that argument relationships is effective in Chinese semantic role labeling task.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP1
2016 Joint Learning Templates and Slots for Event Schema Induction
abstract
Automatic event schema induction (AESI) means to extract meta-event from raw text, in other words, to find out what types (templates) of event may exist in the raw text and what roles (slots) may exist in each event type.In this paper, we propose a joint entity-driven model to learn templates and slots simultaneously based on the constraints of templates and slots in the same sentence.In addition, the entities' semantic information is also considered for the inner connectivity of the entities.We borrow the normalized cut criteria in image segmentation to divide the entities into more accurate template clusters and slot clusters.The experiment shows that our model gains a relatively higher result than previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
HLT-NAACL1
2015 Multi-label Text Categorization with Joint Learning Predictions-as-Features Method
abstract
Multi-label text categorization is a type of text categorization, where each document is assigned to one or more categories.Recently, a series of methods have been developed, which train a classifier for each label, organize the classifiers in a partially ordered structure and take predictions produced by the former classifiers as the latter classifiers' features.These predictions-asfeatures style methods model high order label dependencies and obtain high performance.Nevertheless, the predictionsas-features methods suffer a drawback.When training a classifier for one label, the predictions-as-features methods can model dependencies between former labels and the current label, but they can't model dependencies between the current label and the latter labels.To address this problem, we propose a novel joint learning algorithm that allows the feedbacks to be propagated from the classifiers for latter labels to the classifier for the current label.We conduct experiments using real-world textual data sets, and these experiments illustrate the predictions-as-features models trained by our algorithm outperform the original models.
Houfeng Wang, Xu Sun 0001, Baobao Chang, Shi Zhao, Lei Sha
EMNLP6
2015 Recognizing Textual Entailment Using Probabilistic Inference
abstract
Recognizing Text Entailment (RTE) plays an important role in NLP applications including question answering, information retrieval, etc.In recent work, some research explore "deep" expressions such as discourse commitments or strict logic for representing the text.However, these expressions suffer from the limitation of inference inconvenience or translation loss.To overcome the limitations, in this paper, we propose to use the predicate-argument structures to represent the discourse commitments extracted from text.At the same time, with the help of the YAGO knowledge, we borrow the distant supervision technique to mine the implicit facts from the text.We also construct a probabilistic network for all the facts and conduct inference to judge the confidence of each fact for RTE.The experimental results show that our proposed method achieves a competitive result compared to the previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP1
2014 Event Schema Induction Based on Relational Co-occurrence over Multiple Documents
Tingsong Jiang, Lei Sha, Zhifang Sui
NLPCC2