EDBT 2026 Demo / reviewers in the wild / expert
Wei Lu 0011
dblp:98/6613-11
· DBLP profile ↗
81ranked-venue papers
10as first author
22since 2021 · last 2025
0000-0003-0827-0382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 9 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network ServicesabstractAs interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce PET-BENCH, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, PET-BENCH emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. PET-BENCH serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions. Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang 0006, Weiting Liu 0001, Zheyu Ye, Zhoujun Li 0001, Zuozhu Liu, Wei Lu 0011 |
CIKM | 9 |
| 2025 | Self-Harmonized Chain of ThoughtabstractChain-of-thought (CoT) prompting has demonstrated the capacity of large language models to perform complex reasoning through intermediate steps.While effective, current CoT methods face challenges: Zero-shot-CoT can lead to reasoning errors, and Few-shot-CoT requires labor-intensive manual demonstrations.Auto-CoT attempts to address these issues by automatically generating diverse demonstrations, but this diversity can lead to inconsistent reasoning patterns.We propose ECHO (Self-Harmonized Chain of Thought), a novel method that unifies diverse solution paths into a consistent and effective reasoning pattern.ECHO employs an iterative process to refine and harmonize automatically generated demonstrations, mitigating the limitations of existing approaches.Our comprehensive experiments across arithmetic, commonsense, and symbolic reasoning tasks demonstrate that ECHO outperforms Auto-CoT by an average of 2.8%.These findings suggest that ECHO represents a significant step towards more robust and generalizable automated reasoning in large language models. 1 Q: Mary found 18 seashells and Jessica found 41 seashells on the beach.How many seashells did they find together?A: Let's think step by step.First, we need to add the number of seashells Mary found to the number of seashells Jessica found.18 (seashells found by Mary) + 41 (seashells found by Jessica) = 59 So, they found a total of 59 seashells together.Q: Jason grew 23 watermelons and 18 turnips.Nancy grew 28 watermelons.How many watermelons did they grow in total?A: Let's think step by step.First, we need to add the number of watermelons Jason grew to the number of watermelons Nancy grew.23 (watermelons grown by Jason) + 28 (watermelons grown by Nancy) = 51 So, they grew a total of 51 watermelons.Q: Last week Tim had 12 dollars and Keith had 36 dollars . Tim washed cars over the weekend and now has 75 dollars . How much money did Tim make washing cars ?A: Let's think step by step.... Demo Regeneration Demonstrations in random order Demo i to be regenerated First, we need to find out how much money Tim made by subtracting his initial amount from his final amount.75 -12 = 63 This means that Tim made 63 dollars by washing cars over the weekend. Ziqi Jin, Wei Lu 0011 |
NAACL (Long Papers) | 2 |
| 2025 | Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security ReviewabstractLanguage models (LMs) are becoming increasingly popular in real-world applications. Outsourcing model training and data hosting to third-party platforms has become a standard method for reducing costs. In such a situation, the attacker can manipulate the training process or data to inject a backdoor into models. Backdoor attacks are a serious threat where malicious behavior is activated when triggers are present; otherwise, the model operates normally. However, there is still no systematic and comprehensive review of LMs from the attacker's capabilities and purposes on different backdoor attack surfaces. Moreover, there is a shortage of analysis and comparison of the diverse emerging backdoor countermeasures. Therefore, this work aims to provide the natural language processing (NLP) community with a timely review of backdoor attacks and countermeasures. According to the attackers' capability and affected stage of the LMs, the attack surfaces are formalized into four categorizations: attacking the pretrained model with fine-tuning (APMF) or parameter-efficient fine-tuning (PEFT), attacking the final model with training (AFMT), and attacking large language model (ALLM). Thus, attacks under each categorization are combed. The countermeasures are categorized into two general classes: sample inspection and model inspection. Thus, we review countermeasures and analyze their advantages and disadvantages. Also, we summarize the benchmark datasets and provide comparable evaluations for representative attacks and defenses. Drawing the insights from the review, we point out the crucial areas for future research on the backdoor, especially soliciting more efficient and practical countermeasures. Pengzhou Cheng, Zongru Wu, Haodong Zhao, Wei Lu 0011, Gongshen Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Constrained Layout Generation with Factor GraphsabstractThis paper addresses the challenge of object-centric lay-out generation under spatial constraints, seen in multi-ple domains including floorplan design process. The de-sign process typically involves specifying a set of spa-tial constraints that include object attributes like size and inter-object relations such as relative positioning. Existing works, which typically represent objects as single nodes, lack the granularity to accurately model complex interactions between objects. For instance, often only certain parts of an object, like a room's right wall, interact with adjacent objects. To address this gap, we introduce a factor graph based approach with four latent variable nodes for each room, and a factor node for each constraint. The factor nodes represent dependencies among the variables to which they are connected, effectively capturing constraints that are potentially of a higher order. We then develop message-passing on the bipartite graph, forming a factor graph neu-ral network that is trained to produce a floorplan that aligns with the desired requirements. Our approach is simple and generates layouts faithful to the user requirements, demon-strated by a large improvement in IOU scores over existing methods. Additionally, our approach, being inferential and accurate, is well-suited to the practical human-in-the-loop design process where specifications evolve iteratively, offering a practical and powerful tool for AI-guided design. Mohammed Haroon Dupty, Yanfei Dong, Sicong Leng, Guoji Fu, Yong Liang Goh, Wei Lu 0011, Wee Sun Lee |
CVPR | 6 |
| 2024 | OptScaler: A Collaborative Framework for Robust Autoscaling in the CloudabstractAutoscaling is a critical mechanism in cloud computing, enabling the autonomous adjustment of computing resources in response to dynamic workloads. This is particularly valuable for co-located, long-running applications with diverse workload patterns. The primary objective of autoscaling is to regulate resource utilization at a desired level, effectively balancing the need for resource optimization with the fulfillment of Service Level Objectives (SLOs). Many existing proactive autoscaling frameworks may encounter prediction deviations arising from the frequent fluctuations of cloud workloads. Reactive frameworks, on the other hand, rely on realtime system feedback, but their hysteretic nature could lead to violations of stringent SLOs. Hybrid frameworks, while prevalent, often feature independently functioning proactive and reactive modules, potentially leading to incompatibility and undermining the overall decision-making efficacy. In addressing these challenges, we propose OptScaler, a collaborative autoscaling framework that integrates proactive and reactive modules through an optimization module. The proactive module delivers reliable future workload predictions to the optimization module, while the reactive module offers a self-tuning estimator for real-time updates. By embedding a Model Predictive Control (MPC) mechanism and chance constraints into the optimization module, we further enhance its robustness. Numerical results have demonstrated the superiority of our workload prediction model and the collaborative framework, leading to over a 36% reduction in SLO violations compared to prevalent reactive, proactive, or hybrid autoscalers. Notably, OptScaler has been successfully deployed at Alipay, providing autoscaling support for the world-leading payment platform. Aaron Zou, Wei Lu 0011, Zhibo Zhu, Xingyu Lu 0004, Jun Zhou 0011, Xiaojin Wang, Kangyu Liu, Kefan Wang, Renen Sun |
Proc. VLDB Endow. | 2 |
| 2023 | Tell2Design: A Dataset for Language-Guided Floor Plan GenerationabstractWe consider the task of generating designs directly from natural language descriptions, and consider floor plan generation as the initial research area.Language conditional generative models have recently been very successful in generating high-quality artistic images.However, designs must satisfy different constraints that are not present in generating artistic images, particularly spatial and relational constraints.We make multiple contributions to initiate research on this task.First, we introduce a novel dataset, Tell2Design (T2D), which contains more than 80k floor plan designs associated with natural language instructions.Second, we propose a Sequence-to-Sequence model that can serve as a strong baseline for future research.Third, we benchmark this task with several text-conditional image generation models.We conclude by conducting human evaluations on the generated samples and providing an analysis of human performance.We hope our contributions will propel the research on language-guided design generation forward 1 . Sicong Leng, Yang Zhou 0017, Mohammed Haroon Dupty, Wee Sun Lee, Sam Joyce, Wei Lu 0011 |
ACL (1) | 6 |
| 2023 | Contextual Distortion Reveals Constituency: Masked Language Models are Implicit ParsersabstractRecent advancements in pre-trained language models (PLMs) have demonstrated that these models possess some degree of syntactic awareness.To leverage this knowledge, we propose a novel chart-based method for extracting parse trees from masked language models (LMs) without the need to train separate parsers.Our method computes a score for each span based on the distortion of contextual representations resulting from linguistic perturbations.We design a set of perturbations motivated by the linguistic concept of constituency tests, and use these to score each span by aggregating the distortion scores.To produce a parse tree, we use chart parsing to find the tree with the minimum score.Our method consistently outperforms previous state-of-the-art methods on English with masked LMs, and also demonstrates superior performance in a multilingual setting, outperforming the state of the art in 6 out of 8 languages.Notably, although our method does not involve parameter updates or extensive hyperparameter search, its performance can even surpass some unsupervised parsing methods that require fine-tuning.Our analysis highlights that the distortion of contextual representation resulting from syntactic perturbation can serve as an effective indicator of constituency across languages.1 Jiaxi Li 0001, Wei Lu 0011 |
ACL (1) | 2 |
| 2023 | One Network, Many Masks: Towards More Parameter-Efficient Transfer LearningabstractFine-tuning pre-trained language models for multiple tasks tends to be expensive in terms of storage.To mitigate this, parameter-efficient transfer learning (PETL) methods have been proposed to address this issue, but they still require a significant number of parameters and storage when being applied to broader ranges of tasks.To achieve even greater storage reduction, we propose PROPETL, a novel method that enables efficient sharing of a single PETL module which we call prototype network (e.g., adapter, LoRA, and prefix-tuning) across layers and tasks.We then learn binary masks to select different sub-networks from the shared prototype network and apply them as PETL modules into different layers.We find that the binary masks can determine crucial information from the network, which is often ignored in previous studies.Our work can also be seen as a type of pruning method, where we find that overparameterization also exists in the seemingly small PETL modules.We evaluate PROPETL on various downstream tasks and show that it can outperform other PETL methods with approximately 10% of the parameter storage required by the latter. 1 Guangtao Zeng, Peiyuan Zhang, Wei Lu 0011 |
ACL (1) | 3 |
| 2023 | Global-Aware Model-Free Self-distillation for Recommendation System
Ang Li 0043, Jian Hu 0002, Wei Lu 0011, Ke Ding 0001, Jun Zhou 0011, Yong He 0009, Liang Zhang 0045, Lihong Gu |
DASFAA (4) | 3 |
| 2023 | Unraveling Feature Extraction Mechanisms in Neural NetworksabstractThe underlying mechanism of neural networks in capturing precise knowledge has been the subject of consistent research efforts.In this work, we propose a theoretical approach based on Neural Tangent Kernels (NTKs) to investigate such mechanisms.Specifically, considering the infinite network width, we hypothesize the learning dynamics of target models may intuitively unravel the features they acquire from training data, deepening our insights into their internal mechanisms.We apply our approach to several fundamental models and reveal how these models leverage statistical features during gradient descent and how they are integrated into final decisions.We also discovered that the choice of activation function can affect feature extraction.For instance, the use of the ReLU activation function could potentially introduce a bias in features, providing a plausible explanation for its replacement with alternative functions in recent pre-trained language models.Additionally, we find that while self-attention and CNN models may exhibit limitations in learning n-grams, multiplication-based models seem to excel in this area.We verify these theoretical findings through experiments and find that they can be applied to analyze language modeling tasks, which can be regarded as a special variant of classification.Our contributions offer insights into the roles and capacities of fundamental components within large language models, thereby aiding the broader understanding of these complex systems. Xiaobing Sun 0002, Jiaxi Li 0001, Wei Lu 0011 |
EMNLP | 3 |
| 2023 | Towards Hard Few-Shot Relation ClassificationabstractFew-shot relation classification (FSRC) focuses on recognizing novel relations by learning with merely a handful of annotated instances. Meta-learning has been widely adopted for such a task, which trains on randomly generated few-shot tasks to learn generic data representations. Despite impressive results achieved, existing models still perform suboptimally when handling hard FSRC tasks with similar categories that confuse the model to distinguish correctly. We argue this is largely due to two reasons, 1) ignoring pivotal and discriminate information that is crucial to distinguish confusing classes, and 2) training indiscriminately via randomly sampled tasks of varying difficulty. In this article, we introduce a novel prototypical network approach with contrastive learning that learns more informative and discriminative representations by exploiting relation label information. We further design two strategies that increase the difficulty of training tasks and allow the model to adaptively learn to focus on hard tasks. By doing so, our model can better represent subtle inter-relation variance and grow up through task difficulty. Extensive experiments on three standard benchmarks demonstrate the effectiveness of our method. Jiale Han 0001, Bo Cheng 0001, Zhiguo Wan, Wei Lu 0011 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Learning to Reason Deductively: Math Word Problem Solving as Complex Relation ExtractionabstractSolving math word problems requires deductive reasoning over the quantities in the text.Various recent research efforts mostly relied on sequence-to-sequence or sequence-to-tree models to generate mathematical expressions without explicitly performing relational reasoning between quantities in the given context.While empirically effective, such approaches typically do not provide explanations for the generated expressions.In this work, we view the task as a complex relation extraction problem, proposing a novel approach that presents explainable deductive reasoning steps to iteratively construct target expressions, where each step involves a primitive operation over two quantities defining their relation.Through extensive experiments on four benchmark datasets, we show that the proposed model significantly outperforms existing strong baselines.We further demonstrate that the deductive procedure not only presents more explainable steps but also enables us to make more accurate predictions on questions that require more complex reasoning.1 Zhanming Jie, Jierui Li, Wei Lu 0011 |
ACL (1) | 3 |
| 2022 | Differentiable Data Augmentation for Contrastive Sentence Representation LearningabstractFine-tuning a pre-trained language model via the contrastive learning framework with a large amount of unlabeled sentences or labeled sentence pairs is a common way to obtain highquality sentence representations.Although the contrastive learning framework has shown its superiority on sentence representation learning over previous methods, the potential of such a framework is under-explored so far due to the simple method it used to construct positive pairs.Motivated by this, we propose a method that makes hard positives from the original training examples.A pivotal ingredient of our approach is the use of prefix that is attached to a pre-trained language model, which allows for differentiable data augmentation during contrastive learning.Our method can be summarized in two steps: supervised prefixtuning followed by joint contrastive fine-tuning with unlabeled or labeled examples.Our experiments confirm the effectiveness of our data augmentation approach.The proposed method yields significant improvements over existing methods under both semi-supervised and supervised settings.Our experiments under a low labeled data setting also show that our method is more label-efficient than the state-of-the-art contrastive learning methods. 1 Tianduo Wang, Wei Lu 0011 |
EMNLP | 2 |
| 2022 | Unsupervised Non-transferable Text ClassificationabstractTraining a good deep learning model requires substantial data and computing resources, which makes the resulting neural model a valuable intellectual property.To prevent the neural network from being undesirably exploited, non-transferable learning has been proposed to reduce the model generalization ability in specific target domains.However, existing approaches require labeled data for the target domain which can be difficult to obtain.Furthermore, they do not have the mechanism to still recover the model's ability to access the target domain.In this paper, we propose a novel unsupervised non-transferable learning method for the text classification task that does not require annotated target domain data.We further introduce a secret key component in our approach for recovering the access to the target domain, where we design both an explicit and an implicit method for doing so.Extensive experiments demonstrate the effectiveness of our approach. Guangtao Zeng, Wei Lu 0011 |
EMNLP | 2 |
| 2022 | Better Few-Shot Relation Extraction with Label Prompt DropoutabstractFew-shot relation extraction aims to learn to identify the relation between two entities based on very limited training examples.Recent efforts found that textual labels (i.e., relation names and relation descriptions) could be extremely useful for learning class representations, which will benefit the few-shot learning task.However, what is the best way to leverage such label information in the learning process is an important research question.Existing works largely assume such textual labels are always present during both learning and prediction.In this work, we argue that such approaches may not always lead to optimal results.Instead, we present a novel approach called label prompt dropout, which randomly removes label descriptions in the learning process.Our experiments show that our approach is able to lead to improved class representations, yielding significantly better results on the few-shot relation extraction task. 1 Peiyuan Zhang, Wei Lu 0011 |
EMNLP | 2 |
| 2022 | Implicit n-grams Induced by RecurrenceabstractAlthough self-attention based models such as Transformers have achieved remarkable successes on natural language processing (NLP) tasks, recent studies reveal that they have limitations on modeling sequential transformations (Hahn, 2020), which may prompt re-examinations of recurrent neural networks (RNNs) that demonstrated impressive results on handling sequential data.Despite many prior attempts to interpret RNNs, their internal mechanisms have not been fully understood, and the question on how exactly they capture sequential features remains largely unclear.In this work, we present a study that shows there actually exist some explainable components that reside within the hidden states, which are reminiscent of the classical n-grams features.We evaluated such extracted explainable features from trained RNNs on downstream sentiment analysis tasks and found they could be used to model interesting linguistic phenomena such as negation and intensification.Furthermore, we examined the efficacy of using such n-gram components alone as encoders on tasks such as sentiment analysis and language modeling, revealing they could be playing important roles in contributing to the overall performance of RNNs.We hope our findings could add interpretability to RNN architectures, and also provide inspirations for proposing new architectures for sequential data. Xiaobing Sun 0002, Wei Lu 0011 |
NAACL-HLT | 2 |
| 2021 | Interventional Video Grounding With Dual Contrastive LearningabstractVideo grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies, i.e., P(Y |X). Consequently, these models may suffer from spurious correlations between the language and video features due to the selection bias of the dataset. 1) To uncover the causality behind the model and data, we first propose a novel paradigm from the perspective of the causal inference, i.e., interventional video grounding (IVG) that leverages backdoor adjustment to deconfound the selection bias based on structured causal model (SCM) and do-calculus P(Y |do(X)). Then, we present a simple yet effective method to approximate the unobserved confounder as it cannot be directly sampled from the dataset. 2) Meanwhile, we introduce a dual contrastive learning approach (DCL) to better align the text and video by maximizing the mutual information (MI) between query and video clips, and the MI between start/end frames of a target moment and the others within a video to learn more informative visual representations. Experiments on three standard benchmarks show the effectiveness of our approaches. Guoshun Nan, Rui Qiao 0006, Jun Liu 0036, Sicong Leng, Hao Zhang 0048, Wei Lu 0011 |
CVPR | 7 |
| 2021 | Exploring Task Difficulty for Few-Shot Relation ExtractionabstractFew-shot relation extraction (FSRE) focuses on recognizing novel relations by learning with merely a handful of annotated instances.Meta-learning has been widely adopted for such a task, which trains on randomly generated few-shot tasks to learn generic data representations.Despite impressive results achieved, existing models still perform suboptimally when handling hard FSRE tasks, where the relations are fine-grained and similar to each other.We argue this is largely because existing models do not distinguish hard tasks from easy ones in the learning process.In this paper, we introduce a novel approach based on contrastive learning that learns better representations by exploiting relation label information.We further design a method that allows the model to adaptively learn how to focus on hard tasks.Experiments on two standard datasets demonstrate the effectiveness of our method. Jiale Han 0001, Bo Cheng 0001, Wei Lu 0011 |
EMNLP (1) | 3 |
| 2021 | Uncovering Main Causalities for Long-tailed Information ExtractionabstractInformation Extraction (IE) aims to extract structural information from unstructured texts.In practice, long-tailed distributions caused by the selection bias of a dataset, may lead to incorrect correlations, also known as spurious correlations, between entities and labels in the conventional likelihood models.This motivates us to propose counterfactual IE (CFIE), a novel framework that aims to uncover the main causalities behind data in the view of causal inference.Specifically, 1) we first introduce a unified structural causal model (SCM) for various IE tasks, describing the relationships among variables; 2) with our SCM, we then generate counterfactuals based on an explicit language structure to better calculate the direct causal effect during the inference stage; 3) we further propose a novel debiasing approach to yield more robust predictions.Experiments on three IE tasks across five public datasets show the effectiveness of our CFIE model in mitigating the spurious correlation issues. Guoshun Nan, Jiaqi Zeng, Rui Qiao 0006, Zhijiang Guo, Wei Lu 0011 |
EMNLP (1) | 5 |
| 2021 | To be Closer: Learning to Link up Aspects with OpinionsabstractDependency parse trees are helpful for discovering the opinion words in aspect-based sentiment analysis (ABSA) (Huang and Carley, 2019).However, the trees obtained from offthe-shelf dependency parsers are static, and could be sub-optimal in ABSA.This is because the syntactic trees are not designed for capturing the interactions between opinion words and aspect words.In this work, we aim to shorten the distance between aspects and corresponding opinion words by learning an aspect-centric tree structure.The aspect and opinion words are expected to be closer along such tree structure compared to the standard dependency parse tree.The learning process allows the tree structure to adaptively correlate the aspect and opinion words, enabling us to better identify the polarity in the ABSA task.We conduct experiments on five aspectbased sentiment datasets, and the proposed model significantly outperforms recent strong baselines.Furthermore, our thorough analysis demonstrates the average distance between aspect and opinion words are shortened by at least 19% on the standard SemEval Restau-rant14 (Pontiki et al., 2014) dataset 1 . Lejian Liao, Yang Gao 0016, Zhanming Jie, Wei Lu 0011 |
EMNLP (1) | 5 |
| 2021 | Mixed Cross Entropy Loss for Neural Machine TranslationabstractIn neural machine translation, Cross Entropy loss (CE) is the standard loss function in two training methods of auto-regressive models, i.e., teacher forcing and scheduled sampling. In this paper, we propose mixed Cross Entropy loss (mixed CE) as a substitute for CE in both training approaches. In teacher forcing, the model trained with CE regards the translation problem as a one-to-one mapping process, while in mixed CE this process can be relaxed to one-to-many. In scheduled sampling, we show that mixed CE has the potential to encourage the training and testing behaviours to be similar to each other, more effectively mitigating the exposure bias problem. We demonstrate the superiority of mixed CE over CE on several machine translation datasets, WMT’16 Ro-En, WMT’16 Ru-En, and WMT’14 En-De in both teacher forcing and scheduled sampling setups. Furthermore, in WMT’14 En-De, we also find mixed CE consistently outperforms CE on a multi-reference set as well as a challenging paraphrased reference set. We also found the model trained with mixed CE is able to provide a better probability distribution defined over the translation output space. Our code is available at https://github.com/haorannlp/mix. Wei Lu 0011 |
ICML | 2 |
| 2021 | Better Feature Integration for Named Entity RecognitionabstractIt has been shown that named entity recognition (NER) could benefit from incorporating the long-distance structured information captured by dependency trees.We believe this is because both types of features -the contextual information captured by the linear sequences and the structured information captured by the dependency trees may complement each other.However, existing approaches largely focused on stacking the LSTM and graph neural networks such as graph convolutional networks (GCNs) for building improved NER models, where the exact interaction mechanism between the two different types of features is not very clear, and the performance gain does not appear to be significant.In this work, we propose a simple and robust solution to incorporate both types of features with our Synergized-LSTM (Syn-LSTM), which clearly captures how the two types of features interact.We conduct extensive experiments on several standard datasets across four languages.The results demonstrate that the proposed model achieves better performance than previous approaches while requiring fewer parameters.Our further analysis demonstrates that our model can capture longer dependencies compared with strong baselines.1 Lu Xu 0007, Zhanming Jie, Wei Lu 0011, Lidong Bing |
NAACL-HLT | 3 |
| 2020 | Knowing What, How and Why: A Near Complete Solution for Aspect-Based Sentiment AnalysisabstractTarget-based sentiment analysis or aspect-based sentiment analysis (ABSA) refers to addressing various sentiment analysis tasks at a fine-grained level, which includes but is not limited to aspect extraction, aspect sentiment classification, and opinion extraction. There exist many solvers of the above individual subtasks or a combination of two subtasks, and they can work together to tell a complete story, i.e. the discussed aspect, the sentiment on it, and the cause of the sentiment. However, no previous ABSA research tried to provide a complete solution in one shot. In this paper, we introduce a new subtask under ABSA, named aspect sentiment triplet extraction (ASTE). Particularly, a solver of this task needs to extract triplets (What, How, Why) from the inputs, which show WHAT the targeted aspects are, HOW their sentiment polarities are and WHY they have such polarities (i.e. opinion reasons). For instance, one triplet from “Waiters are very friendly and the pasta is simply average” could be (‘Waiters’, positive, ‘friendly’). We propose a two-stage framework to address this task. The first stage predicts what, how and why in a unified model, and then the second stage pairs up the predicted what (how) and why from the first stage to output triplets. In the experiments, our framework has set a benchmark performance in this novel triplet extraction task. Meanwhile, it outperforms a few strong baselines adapted from state-of-the-art related methods. Haiyun Peng, Lu Xu 0007, Lidong Bing, Fei Huang 0002, Wei Lu 0011, Luo Si |
AAAI | 5 |
| 2020 | Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionabstractDocument-level relation extraction requires integrating information within and across multiple sentences of a document and capturing complex interactions between inter-sentence entities.However, effective aggregation of relevant information in the document remains a challenging research question.Existing approaches construct static document-level graphs based on syntactic trees, co-references or heuristics from the unstructured text to model the dependencies.Unlike previous methods that may not be able to capture rich non-local interactions for inference, we propose a novel model that empowers the relational reasoning across sentences by automatically inducing the latent document-level graph.We further develop a refinement strategy, which enables the model to incrementally aggregate relevant information for multi-hop reasoning.Specifically, our model achieves an F 1 score of 59.05 on a large-scale documentlevel dataset (DocRED), significantly improving over the previous results, and also yields new state-of-the-art results on the CDR and GDA dataset.Furthermore, extensive analyses show that the model is able to discover more accurate inter-sentence relations. Guoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei Lu 0011 |
ACL | 4 |
| 2020 | Understanding Attention for Text ClassificationabstractAttention has been proven successful in many natural language processing (NLP) tasks.Recently, many researchers started to investigate the interpretability of attention on NLP tasks.Many existing approaches focused on examining whether the local attention weights could reflect the importance of input representations.In this work, we present a study on understanding the internal mechanism of attention by looking into the gradient update process, checking its behavior when approaching a local minimum during training.We propose to analyze for each word token the following two quantities: its polarity score and its attention score, where the latter is a global assessment on the token's significance.We discuss conditions under which the attention mechanism may become more (or less) interpretable, and show how the interplay between the two quantities may impact the model performance.1 Xiaobing Sun 0002, Wei Lu 0011 |
ACL | 2 |
| 2020 | Attention-Based Context Aware Reasoning for Situation RecognitionabstractSituation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action. Predicting semantic roles is very challenging: a vast variety of possibilities can be the match for a semantic role. Existing work has focused on dependency modelling architectures to solve this issue. Inspired by the success achieved by query-based visual reasoning (e.g., Visual Question Answering), we propose to address semantic role prediction as a query-based visual reasoning problem. However, existing query-based reasoning methods have not considered handling of inter-dependent queries which is a unique requirement of semantic role prediction in SR. Therefore, to the best of our knowledge, we propose the first set of methods to address inter-dependent queries in query-based visual reasoning. Extensive experiments demonstrate the effectiveness of our proposed method which achieves outstanding performance on Situation Recognition task. Furthermore, leveraging query inter-dependency, our methods improve upon a state-of-the-art method that answers queries separately. Our code: https://github.com/thilinicooray/context-aware-reasoning-for-sr. Thilini Cooray, Ngai-Man Cheung, Wei Lu 0011 |
CVPR | 3 |
| 2020 | APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task LearningabstractPeer review and rebuttal, with rich interactions and argumentative discussions in between, are naturally a good resource to mine arguments.However, few works study both of them simultaneously.In this paper, we introduce a new argument pair extraction (APE) task on peer review and rebuttal in order to study the contents, the structure and the connections between them.We prepare a challenging dataset that contains 4,764 fully annotated review-rebuttal passage pairs from an open review platform to facilitate the study of this task.To automatically detect argumentative propositions and extract argument pairs from this corpus, we cast it as the combination of a sequence labeling task and a text relation classification task.Thus, we propose a multitask learning framework based on hierarchical LSTM networks.Extensive experiments and analysis demonstrate the effectiveness of our multi-task framework, and also show the challenges of the new task as well as motivate future research directions. 1 Liying Cheng, Lidong Bing, Wei Lu 0011, Luo Si |
EMNLP (1) | 4 |
| 2020 | ENT-DESC: Entity Description Generation by Exploring Knowledge GraphabstractPrevious works on knowledge-to-text generation take as input a few RDF triples or keyvalue pairs conveying the knowledge of some entities to generate a natural language description.Existing datasets, such as WIKIBIO, WebNLG, and E2E, basically have a good alignment between an input triple/pair set and its output text.However, in practice, the input knowledge could be more than enough, since the output description may only cover the most significant knowledge.In this paper, we introduce a large-scale and challenging dataset to facilitate the study of such a practical scenario in KG-to-text.Our dataset involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter explosion while generating the descriptions.We address these challenges by proposing a multi-graph structure that is able to represent the original graph information more comprehensively.Furthermore, we also incorporate aggregation methods that learn to extract the rich graph information.Extensive experiments demonstrate the effectiveness of our model architecture.1 Liying Cheng, Dekun Wu, Lidong Bing, Yan Zhang 0004, Zhanming Jie, Wei Lu 0011, Luo Si |
EMNLP (1) | 6 |
| 2020 | Re-examining the Role of Schema Linking in Text-to-SQLabstractIn existing sophisticated text-to-SQL models, schema linking is often considered as a simple, minor component, belying its importance.By providing a schema linking corpus based on the Spider text-to-SQL dataset, we systematically study the role of schema linking.We also build a simple BERT-based baseline, called Schema-Linking SQL (SLSQL) to perform a data-driven study.We find when schema linking is done well, SLSQL demonstrates good performance on Spider despite its structural simplicity.Many remaining errors are attributable to corpus noise.This suggests schema linking is the crux for the current textto-SQL task.Our analytic studies provide insights on the characteristics of schema linking for future developments of text-to-SQL tasks. 1 Wenqiang Lei, Zhixin Ma 0001, Tian Gan 0002, Wei Lu 0011, Min-Yen Kan, Tat-Seng Chua |
EMNLP (1) | 5 |
| 2020 | Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence EncodersabstractNamed entity recognition and relation extraction are two important fundamental problems.Joint learning algorithms have been proposed to solve both tasks simultaneously, and many of them cast the joint task as a table-filling problem.However, they typically focused on learning a single encoder (usually learning representation in the form of a table) to capture information required for both tasks within the same space.We argue that it can be beneficial to design two distinct encoders to capture such two different types of information in the learning process.In this work, we propose the novel table-sequence encoders where two different encoders -a table encoder and a sequence encoder are designed to help each other in the representation learning process.Our experiments confirm the advantages of having two encoders over one encoder.On several standard datasets, our model shows significant improvements over existing approaches. 1 Jue Wang 0019, Wei Lu 0011 |
EMNLP (1) | 2 |
| 2020 | Aspect Sentiment Classification with Aspect-Specific Opinion SpansabstractAspect sentiment classification, predicting the sentiment polarity of given aspects, has drawn extensive attention.Previous attention-based models emphasize using aspect semantics to help extract opinion features for classification.However, these works are either not able to capture opinion spans as a whole or capture variable-length opinion spans.In this paper, we present a neat and effective multiple CRFs based structured attention model that is capable of extracting aspect-specific opinion spans.The sentiment polarity of the target is then classified based on the extracted opinion features and contextual information.The experimental results on four datasets demonstrate the effectiveness of the proposed model, and our further analysis shows that our model can capture aspect-specific opinion spans. 1 Lu Xu 0007, Lidong Bing, Wei Lu 0011, Fei Huang 0002 |
EMNLP (1) | 3 |
| 2020 | Position-Aware Tagging for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE)is the task of extracting the triplets of target entities, their associated sentiment, and opinion spans explaining the reason for the sentiment.Existing research efforts mostly solve this problem using pipeline approaches, which break the triplet extraction process into several stages.Our observation is that the three elements within a triplet are highly related to each other, and this motivates us to build a joint model to extract such triplets using a sequence tagging approach.However, how to effectively design a tagging approach to extract the triplets that can capture the rich interactions among the elements is a challenging research question.In this work, we propose the first end-to-end model with a novel positionaware tagging scheme that is capable of jointly extracting the triplets.Our experimental results on several existing datasets show that jointly capturing elements in the triplet using our approach leads to improved performance over the existing approaches.We also conducted extensive experiments to investigate the model effectiveness and robustness 1 . Lu Xu 0007, Wei Lu 0011, Lidong Bing |
EMNLP (1) | 3 |
| 2020 | Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text GenerationabstractAMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters. Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing |
EMNLP (1) | 4 |
| 2020 | Pre-training for Abstractive Document Summarization by Reinstating Source TextabstractAbstractive document summarization is usually modeled as a sequence-to-sequence (SEQ2SEQ) learning problem.Unfortunately, training large SEQ2SEQ based summarization models on limited supervised summarization data is challenging.This paper presents three sequence-to-sequence pre-training (in shorthand, STEP) objectives which allow us to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text.The main idea is that, given an input text artificially constructed from a document, a model is pre-trained to reinstate the original document.These objectives include sentence reordering, next sentence generation and masked document generation, which have close relations with the abstractive document summarization task.Experiments on two benchmark summarization datasets (i.e., CNN/DailyMail and New York Times) show that all three objectives can improve performance upon baselines.Compared to models pre-trained on large-scale data (≥160GB), our method, with only 19GB text for pre-training, achieves comparable results, which demonstrates its effectiveness.Code and models are public available at https://github.com/ zoezou2015/abs_pretraining. Xingxing Zhang 0002, Wei Lu 0011, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 3 |
| 2020 | Learning Latent Forests for Medical Relation ExtractionabstractThe goal of medical relation extraction is to detect relations among entities, such as genes, mutations and drugs in medical texts. Dependency tree structures have been proven useful for this task. Existing approaches to such relation extraction leverage off-the-shelf dependency parsers to obtain a syntactic tree or forest for the text. However, for the medical domain, low parsing accuracy may lead to error propagation downstream the relation extraction pipeline. In this work, we propose a novel model which treats the dependency structure as a latent variable and induces it from the unstructured text in an end-to-end fashion. Our model can be understood as composing task-specific dependency forests that capture non-local interactions for better relation extraction. Extensive results on four datasets show that our model is able to significantly outperform state-of-the-art systems without relying on any direct tree supervision or pre-training. Zhijiang Guo, Guoshun Nan, Wei Lu 0011, Shay B. Cohen |
IJCAI | 3 |
| 2019 | A Neural Multi-digraph Model for Chinese NER with GazetteersabstractGazetteers were shown to be useful resources for named entity recognition (NER) (Ratinov and Roth, 2009).Many existing approaches to incorporating gazetteers into machine learning based NER systems rely on manually defined selection strategies or handcrafted templates, which may not always lead to optimal effectiveness, especially when multiple gazetteers are involved.This is especially the case for the task of Chinese NER, where the words are not naturally tokenized, leading to additional ambiguities.To automatically learn how to incorporate multiple gazetteers into an NER system, we propose a novel approach based on graph neural networks with a multidigraph structure that captures the information that the gazetteers offer.Experiments on various datasets show that our model is effective in incorporating rich gazetteer information while resolving ambiguities, outperforming previous approaches. Ruixue Ding, Pengjun Xie, Wei Lu 0011, Linlin Li 0001, Luo Si |
ACL (1) | 4 |
| 2019 | Attention Guided Graph Convolutional Networks for Relation ExtractionabstractDependency trees convey rich structural information that is proven useful for extracting relations among entities in text.However, how to effectively make use of relevant information while ignoring irrelevant information from the dependency trees remains a challenging research question.Existing approaches employing rule based hard-pruning strategies for selecting relevant partial dependency structures may not always yield optimal results.In this work, we propose Attention Guided Graph Convolutional Networks (AGGCNs), a novel model which directly takes full dependency trees as inputs.Our model can be understood as a soft-pruning approach that automatically learns how to selectively attend to the relevant sub-structures useful for the relation extraction task.Extensive results on various tasks including cross-sentence n-ary relation extraction and large-scale sentence-level relation extraction show that our model is able to better leverage the structural information of the full dependency trees, giving significantly better results than previous approaches. Zhijiang Guo, Yan Zhang 0004, Wei Lu 0011 |
ACL (1) | 3 |
| 2019 | Twitter Homophily: Network Based Prediction of User's OccupationabstractIn this paper, we investigate the importance of social network information compared to content information in the prediction of a Twitter user's occupational class.We show that the content information of a user's tweets, the profile descriptions of a user's follower/following community, and the user's social network provide useful information for classifying a user's occupational group.In our study, we extend an existing dataset for this problem, and we achieve significantly better performance by using social network homophily that has not been fully exploited in previous work.In our analysis, we found that by using the graph convolutional network to exploit social homophily, we can achieve competitive performance on this dataset with just a small fraction of the training data. Rishabh Bhardwaj, Wei Lu 0011, Hai Leong Chieu, Xinghao Pan, Ni Yi Puay |
ACL (1) | 3 |
| 2019 | Quantity Tagger: A Latent-Variable Sequence Labeling Approach to Solving Addition-Subtraction Word ProblemsabstractAn arithmetic word problem typically includes a textual description containing several constant quantities.The key to solving the problem is to reveal the underlying mathematical relations (such as addition and subtraction) among quantities, and then generate equations to find solutions.This work presents a novel approach, Quantity Tagger, that automatically discovers such hidden relations by tagging each quantity with a sign corresponding to one type of mathematical operation.For each quantity, we assume there exists a latent, variable-sized quantity span surrounding the quantity token in the text, which conveys information useful for determining its sign.Empirical results show that our method achieves 5 and 8 points of accuracy gains on two datasets respectively, compared to prior approaches. Wei Lu 0011 |
ACL (1) | 2 |
| 2019 | Dependency-Guided LSTM-CRF for Named Entity RecognitionabstractZhanming Jie, Wei Lu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhanming Jie, Wei Lu 0011 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Learning Explicit and Implicit Structures for Targeted Sentiment AnalysisabstractTargeted sentiment analysis is the task of jointly predicting target entities and their associated sentiment information.Existing research efforts mostly regard this joint task as a sequence labeling problem, building models that can capture explicit structures in the output space.However, the importance of capturing implicit global structural information that resides in the input space is largely unexplored.In this work, we argue that both types of information (implicit and explicit structural information) are crucial for building a successful targeted sentiment analysis model.Our experimental results show that properly capturing both information is able to lead to better performance than competitive existing approaches.We also conduct extensive experiments to investigate our model's effectiveness and robustness 1 . Wei Lu 0011 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Combining Spans into Entities: A Neural Two-Stage Approach for Recognizing Discontiguous EntitiesabstractBailin Wang, Wei Lu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bailin Wang, Wei Lu 0011 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Aligning Cross-Lingual Entities with Multi-Aspect InformationabstractHsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hsiu-Wei Yang, Peng Shi 0010, Wei Lu 0011, Jimmy Lin, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Text2Math: End-to-end Parsing Text into Math ExpressionsabstractYanyan Zou, Wei Lu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wei Lu 0011 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Spatio-Temporal Event Detection from Multiple Data Sources
Aman Ahuja, Ashish Baghudana, Wei Lu 0011, Edward A. Fox, Chandan K. Reddy |
PAKDD (1) | 3 |
| 2019 | Densely Connected Graph Convolutional Networks for Graph-to-Sequence LearningabstractWe focus on graph-to-sequence learning, which can be framed as transducing graph structures to sequences for text generation. To capture structural information associated with graphs, we investigate the problem of encoding graphs using graph convolutional networks (GCNs). Unlike various existing approaches where shallow architectures were used for capturing local structural information only, we introduce a dense connection strategy, proposing a novel Densely Connected Graph Convolutional Network (DCGCN). Such a deep architecture is able to integrate both local and non-local features to learn a better structural representation of a graph. Our model outperforms the state-of-the-art neural models significantly on AMR-to-text generation and syntax-based neural machine translation. Zhijiang Guo, Yan Zhang 0004, Zhiyang Teng, Wei Lu 0011 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2018 | cw2vec: Learning Chinese Word Embeddings with Stroke n-gram InformationabstractWe propose cw2vec, a novel method for learning Chinese word embeddings. It is based on our observation that exploiting stroke-level information is crucial for improving the learning of Chinese word embeddings. Specifically, we design a minimalist approach to exploit such features, by using stroke n-grams, which capture semantic and morphological level information of Chinese words. Through qualitative analysis, we demonstrate that our model is able to extract semantic information that cannot be captured by existing methods. Empirical results on the word similarity, word analogy, text classification and named entity recognition tasks show that the proposed approach consistently outperforms state-of-the-art approaches such as word-based word2vec and GloVe, character-based CWE, component-based JWE and pixel-based GWE. Shaosheng Cao, Wei Lu 0011, Jun Zhou 0011, Xiaolong Li 0005 |
AAAI | 2 |
| 2018 | Learning Latent Opinions for Aspect-level Sentiment ClassificationabstractAspect-level sentiment classification aims at detecting the sentiment expressed towards a particular target in a sentence. Based on the observation that the sentiment polarity is often related to specific spans in the given sentence, it is possible to make use of such information for better classification. On the other hand, such information can also serve as justifications associated with the predictions.We propose a segmentation attention based LSTM model which can effectively capture the structural dependencies between the target and the sentiment expressions with a linear-chain conditional random field (CRF) layer. The model simulates human's process of inferring sentiment information when reading: when given a target, humans tend to search for surrounding relevant text spans in the sentence before making an informed decision on the underlying sentiment information.We perform sentiment classification tasks on publicly available datasets on online reviews across different languages from SemEval tasks and social comments from Twitter. Extensive experiments show that our model achieves the state-of-the-art performance while extracting interpretable sentiment expressions. Bailin Wang, Wei Lu 0011 |
AAAI | 2 |
| 2018 | Better Transition-Based AMR Parsing with Refined Search SpaceabstractThis paper introduces a simple yet effective transition-based system for Abstract Meaning Representation (AMR) parsing.We argue that a well-defined search space for a transition system is crucial for building an effective parser.We propose to conduct the search in a refined search space based on a new compact AMR graph and an improved oracle.Our end-to-end parser achieves the state-of-the-art performance on various datasets with minimal additional information.1 Zhijiang Guo, Wei Lu 0011 |
EMNLP | 2 |
| 2018 | Dependency-based Hybrid Trees for Semantic ParsingabstractWe propose a novel dependency-based hybrid tree model for semantic parsing, which converts natural language utterance into machine interpretable meaning representations.Unlike previous state-of-the-art models, the semantic information is interpreted as the latent dependency between the natural language words in our joint representation.Such dependency information can capture the interactions between the semantics and natural language words.We integrate a neural component into our model and propose an efficient dynamicprogramming algorithm to perform tractable inference.Through extensive experiments on the standard multilingual GeoQuery dataset with eight languages, we demonstrate that our proposed approach is able to achieve state-ofthe-art performance across several languages.Analysis also justifies the effectiveness of using our new dependency-based representation. 1 Zhanming Jie, Wei Lu 0011 |
EMNLP | 2 |
| 2018 | Neural Adaptation Layers for Cross-domain Named Entity RecognitionabstractRecent research efforts have shown that neural architectures can be effective in conventional information extraction tasks such as named entity recognition, yielding state-of-the-art results on standard newswire datasets.However, despite significant resources required for training such models, the performance of a model trained on one domain typically degrades dramatically when applied to a different domain, yet extracting entities from new emerging domains such as social media can be of significant interest.In this paper, we empirically investigate effective methods for conveniently adapting an existing, well-trained neural NER model for a new domain.Unlike existing approaches, we propose lightweight yet effective methods for performing domain adaptation for neural models.Specifically, we introduce adaptation layers on top of existing neural architectures, where no re-training using the source domain data is required.We conduct extensive empirical studies and show that our approach significantly outperforms stateof-the-art methods. Bill Y. Lin, Wei Lu 0011 |
EMNLP | 2 |
| 2018 | Neural Segmental Hypergraphs for Overlapping Mention RecognitionabstractIn this work, we propose a novel segmental hypergraph representation to model overlapping entity mentions that are prevalent in many practical datasets.We show that our model built on top of such a new representation is able to capture features and interactions that cannot be captured by previous models while maintaining a low time complexity for inference.We also present a theoretical analysis to formally assess how our representation is better than alternative representations reported in the literature in terms of representational power.Coupled with neural networks for feature learning, our model achieves the state-of-the-art performance in three benchmark datasets annotated with overlapping mentions.1 1 We make our system and code available at: http Bailin Wang, Wei Lu 0011 |
EMNLP | 2 |
| 2018 | A Neural Transition-based Model for Nested Mention RecognitionabstractIt is common that entity mentions can contain other mentions recursively.This paper introduces a scalable transition-based method to model the nested structure of mentions.We first map a sentence with nested mentions to a designated forest where each mention corresponds to a constituent of the forest.Our shiftreduce based system then learns to construct the forest structure in a bottom-up manner through an action sequence whose maximal length is guaranteed to be three times of the sentence length.Based on Stack-LSTM which is employed to efficiently and effectively represent the states of the system in a continuous space, our system is further incorporated with a character-based component to capture letterlevel patterns.Our model achieves the stateof-the-art results on ACE datasets, showing its effectiveness in detecting nested mentions.1 Bailin Wang, Wei Lu 0011, Yu Wang 0091, Hongxia Jin |
EMNLP | 2 |
| 2017 | Improving Word Embeddings with Convolutional Feature Learning and Subword InformationabstractWe present a novel approach to learning word embeddings by exploring subword information (character n-gram, root/affix and inflections) and capturing the structural information of their context with convolutional feature learning. Specifically, we introduce a convolutional neural network architecture that allows us to measure structural information of context words and incorporate subword features conveying semantic, syntactic and morphological information related to the words. To assess the effectiveness of our model, we conduct extensive experiments on the standard word similarity and word analogy tasks. We showed improvements over existing state-of-the-art methods for learning word embeddings, including skipgram, GloVe, char n-gram and DSSM. Shaosheng Cao, Wei Lu 0011 |
AAAI | 2 |
| 2017 | Efficient Dependency-Guided Named Entity RecognitionabstractNamed entity recognition (NER), which focuses on the extraction of semantically meaningful named entities and their semantic classes from text, serves as an indispensable component for several down-stream natural language processing (NLP) tasks such as relation extraction and event extraction. Dependency trees, on the other hand, also convey crucial semantic-level information. It has been shown previously that such information can be used to improve the performance of NER. In this work, we investigate on how to better utilize the structured information conveyed by dependency trees to improve the performance of NER. Specifically, unlike existing approaches which only exploit dependency information for designing local features, we show that certain global structured information of the dependency trees can be exploited when building NER models where such information can provide guided learning and inference. Through extensive experiments, we show that our proposed novel dependency-guided NER model performs competitively with models based on conventional semi-Markov conditional random fields, while requiring significantly less running time. Zhanming Jie, Aldrian Obaja Muis, Wei Lu 0011 |
AAAI | 3 |
| 2017 | Learning Latent Sentiment Scopes for Entity-Level Sentiment AnalysisabstractIn this paper, we focus on the task of extracting named entities together with their associated sentiment information in a joint manner. Our key observation in such an entity-level sentiment analysis (a.k.a. targeted sentiment analysis) task is that there exists a sentiment scope within which each named entity is embedded, which largely decides the sentiment information associated with the entity. However, such sentiment scopes are typically not explicitly annotated in the data, and their lengths can be unbounded. Motivated by this, unlike traditional approaches that cast this problem as a simple sequence labeling task, we propose a novel approach that can explicitly model the latent sentiment scopes. Our experiments on the standard datasets demonstrate that our approach is able to achieve better results compared to existing approaches based on conventional conditional random fields (CRFs) and a more recent work based on neural networks. Wei Lu 0011 |
AAAI | 2 |
| 2017 | Semantic Parsing with Neural Hybrid TreesabstractWe propose a neural graphical model for parsing natural language sentences into their logical representations. The graphical model is based on hybrid tree structures that jointly represent both sentences and semantics. Learning and decoding are done using efficient dynamic programming algorithms. The model is trained under a discriminative setting, which allows us to incorporate a rich set of features. Hybrid tree structures have shown to achieve state-of-the-art results on standard semantic parsing datasets. In this work, we propose a novel model that incorporates a rich, nonlinear featurization by a feedforward neural network. The error signals are computed with respect to the conditional random fields (CRFs) objective using an inside-outside algorithm, which are then backpropagated to the neural network. We demonstrate that by combining the strengths of the exact global inference in the hybrid tree models and the power of neural networks to extract high level features, our model is able to achieve new state-of-the-art results on standard benchmark datasets across different languages. Raymond Hendy Susanto, Wei Lu 0011 |
AAAI | 2 |
| 2017 | Topical Coherence in LDA-based Models through Induced SegmentationabstractHesam Amoualian, Wei Lu, Eric Gaussier, Georgios Balikas, Massih R. Amini, Marianne Clausel. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Hesam Amoualian, Wei Lu 0011, Éric Gaussier, Georgios Balikas, Massih-Reza Amini, Marianne Clausel |
ACL (1) | 2 |
| 2017 | MalwareTextDB: A Database for Annotated Malware ArticlesabstractCybersecurity risks and malware threats are becoming increasingly dangerous and common.Despite the severity of the problem, there has been few NLP efforts focused on tackling cybersecurity.In this paper, we discuss the construction of a new database for annotated malware texts.An annotation framework is introduced based around the MAEC vocabulary for defining malware characteristics, along with a database consisting of 39 annotated APT reports with a total of 6,819 sentences.We also use the database to construct models that can potentially help cybersecurity researchers in their data collection and analytics efforts. Swee Kiat Lim, Aldrian Obaja Muis, Wei Lu 0011, Ong Chen Hui |
ACL (1) | 3 |
| 2017 | Labeling Gaps Between Words: Recognizing Overlapping Mentions with Mention SeparatorsabstractIn this paper, we propose a new model that is capable of recognizing overlapping mentions.We introduce a novel notion of mention separators that can be effectively used to capture how mentions overlap with one another.On top of a novel multigraph representation that we introduce, we show that efficient and exact inference can still be performed.We present some theoretical analysis on the differences between our model and a recently proposed model for recognizing overlapping mentions, and discuss the possible implications of the differences.Through extensive empirical analysis on standard datasets, we demonstrate the effectiveness of our approach. Aldrian Obaja Muis, Wei Lu 0011 |
EMNLP | 2 |
| 2017 | A Simple Regularization-based Algorithm for Learning Cross-Domain Word EmbeddingsabstractLearning word embeddings has received a significant amount of attention recently.Often, word embeddings are learned in an unsupervised manner from a large collection of text.The genre of the text typically plays an important role in the effectiveness of the resulting embeddings.How to effectively train word embedding models using data from different domains remains a problem that is underexplored.In this paper, we present a simple yet effective method for learning word embeddings based on text from different domains.We demonstrate the effectiveness of our approach through extensive experiments on various down-stream NLP tasks. Wei Yang 0017, Wei Lu 0011, Vincent Wenchen Zheng |
EMNLP | 2 |
| 2017 | A Probabilistic Geographical Aspect-Opinion Model for Geo-Tagged MicroblogsabstractDue to the rapid increase in the number of users owning location-based devices, there is a considerable amount of geo-tagged data available on social media websites, such as Twitter and Facebook. This geo-tagged data can be useful in a variety of ways to extract location-specific information, as well as to comprehend the variation of information across different geographical regions. A lot of techniques have been proposed for extracting location-based information from social media, but none of these techniques aim to utilize an important characteristic of this data, which is the presence of aspects and their opinions, expressed by the users on these platforms. In this paper, we propose Geographic Aspect Opinion model (GASPOP), a probabilistic model that jointly discovers the variation of aspect and opinion, that correspond to different topics across various geographical regions from geo-tagged social media data. It incorporates the syntactic features of text in the generative process to differentiate aspect and opinion words from general background words. The user-based modeling of topics, also enables it to determine the interest distribution of various users. Furthermore, our model can be used to predict the location of different tweets based on their text. We evaluated our model on Twitter data, and our experimental results show that GASPOP can jointly discover latent aspect and opinion words for different topics across latent geographical regions. Moreover, a quantitative analysis of GASPOP using widely used evaluation metrics shows that it outperforms the state-of-the-art methods. Aman Ahuja, Wei Wei 0019, Wei Lu 0011, Kathleen M. Carley, Chandan K. Reddy |
ICDM | 3 |
| 2016 | Deep Neural Networks for Learning Graph RepresentationsabstractIn this paper, we propose a novel model for learning graph representations, which generates a low-dimensional vector representation for each vertex by capturing the graph structural information. Different from other previous research efforts, we adopt a random surfing model to capture graph structural information directly, instead of using the sampling-based method for generating linear sequences proposed by Perozzi et al. (2014). The advantages of our approach will be illustrated from both theorical and empirical perspectives. We also give a new perspective for the matrix factorization method proposed by Levy and Goldberg (2014), in which the pointwise mutual information (PMI) matrix is considered as an analytical solution to the objective function of the skip-gram model with negative sampling proposed by Mikolov et al. (2013). Unlike their approach which involves the use of the SVD for finding the low-dimensitonal projections from the PMI matrix, however, the stacked denoising autoencoder is introduced in our model to extract complex features and model non-linearities. To demonstrate the effectiveness of our model, we conduct experiments on clustering and visualization tasks, employing the learned vertex representations as features. Empirical results on datasets of varying sizes show that our model outperforms other stat-of-the-art models in such tasks. Shaosheng Cao, Wei Lu 0011, Qiongkai Xu |
AAAI | 2 |
| 2016 | A General Regularization Framework for Domain AdaptationabstractWe propose a domain adaptation framework, and formally prove that it generalizes the feature augmentation technique in (Daumé III, 2007) and the multi-task regularization framework in (Evgeniou and Pontil, 2004).We show that our framework is strictly more general than these approaches and allows practitioners to tune hyper-parameters to encourage transfer between close domains and avoid negative transfer between distant ones. Wei Lu 0011, Hai Leong Chieu, Jonathan Löfgren |
EMNLP | 1 |
| 2016 | Learning to Recognize Discontiguous EntitiesabstractThis paper focuses on the study of recognizing discontiguous entities.Motivated by a previous work, we propose to use a novel hypergraph representation to jointly encode discontiguous entities of unbounded length, which can overlap with one another.To compare with existing approaches, we first formally introduce the notion of model ambiguity, which defines the difficulty level of interpreting the outputs of a model, and then formally analyze the theoretical advantages of our model over previous existing approaches based on linearchain CRFs.Our empirical results also show that our model is able to achieve significantly better results when evaluated on standard data with many discontiguous entities. Aldrian Obaja Muis, Wei Lu 0011 |
EMNLP | 2 |
| 2016 | Learning to Capitalize with Character-Level Recurrent Neural Networks: An Empirical StudyabstractIn this paper, we investigate case restoration for text without case information.Previous such work operates at the word level.We propose an approach using character-level recurrent neural networks (RNN), which performs competitively compared to language modeling and conditional random fields (CRF) approaches.We further provide quantitative and qualitative analysis on how RNN helps improve truecasing. Raymond Hendy Susanto, Hai Leong Chieu, Wei Lu 0011 |
EMNLP | 3 |
| 2016 | Weak Semi-Markov CRFs for Noun Phrase Chunking in Informal TextabstractThis paper introduces a new annotated corpus based on an existing informal text corpus: the NUS SMS Corpus (Chen and Kan, 2013). The new corpus includes 76,490 noun phrases from 26,500 SMS messages, annotated by university students. We then explored several graphical models, including a novel variant of the semi-Markov conditional random fields (semi-CRF) for the task of noun phrase chunking. We demonstrated through empirical evaluations on the new dataset that the new variant yielded similar accuracy but ran in significantly lower running time compared to the conventional semi-CRF. Aldrian Obaja Muis, Wei Lu 0011 |
HLT-NAACL | 2 |
| 2016 | Improving Semantic Parsing with Enriched Synchronous Context-Free Grammars in Statistical Machine TranslationabstractSemantic parsing maps a sentence in natural language into a structured meaning representation. Previous studies show that semantic parsing with synchronous context-free grammars (SCFGs) achieves favorable performance over most other alternatives. Motivated by the observation that the performance of semantic parsing with SCFGs is closely tied to the translation rules, this article explores to extend translation rules with high quality and increased coverage in three ways. First, we examine the difference between word alignments for semantic parsing and statistical machine translation (SMT) to better adapt word alignment in SMT to semantic parsing. Second, we introduce both structure and syntax informed nonterminals, better guiding the parsing in favor of well-formed structure, instead of using a uninformed nonterminal in SCFGs. Third, we address the unknown word translation issue via synthetic translation rules. Last but not least, we use a filtering approach to improve performance via predicting answer type. Evaluation on the standard GeoQuery benchmark dataset shows that our approach greatly outperforms the state of the art across various languages, including English, Chinese, Thai, German, and Greek. Junhui Li 0001, Muhua Zhu, Wei Lu 0011, Guodong Zhou 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | GraRep: Learning Graph Representations with Global Structural InformationabstractIn this paper, we present {GraRep}, a novel model for learning vertex representations of weighted graphs. This model learns low dimensional vectors to represent vertices appearing in a graph and, unlike existing work, integrates global structural information of the graph into the learning process. We also formally analyze the connections between our work and several previous research efforts, including the DeepWalk model of Perozzi et al. as well as the skip-gram model with negative sampling of Mikolov et al. Shaosheng Cao, Wei Lu 0011, Qiongkai Xu |
CIKM | 2 |
| 2015 | Improving Semantic Parsing with Enriched Synchronous Context-Free GrammarabstractSemantic parsing maps a sentence in natural language into a structured meaning representation.Previous studies show that semantic parsing with synchronous contextfree grammars (SCFGs) achieves favorable performance over most other alternatives.Motivated by the observation that the performance of semantic parsing with SCFGs is closely tied to the translation rules, this paper explores extending translation rules with high quality and increased coverage in three ways.First, we introduce structure informed non-terminals, better guiding the parsing in favor of well formed structure, instead of using a uninformed non-terminal in SCFGs.Second, we examine the difference between word alignments for semantic parsing and statistical machine translation (SMT) to better adapt word alignment in SMT to semantic parsing.Finally, we address the unknown word translation issue via synthetic translation rules.Evaluation on the standard GeoQuery benchmark dataset shows that our approach achieves the state-of-the-art across various languages, including English, German and Greek. Junhui Li 0001, Muhua Zhu, Wei Lu 0011, Guodong Zhou 0001 |
EMNLP | 3 |
| 2015 | Joint Mention Extraction and Classification with Mention HypergraphsabstractWe present a novel model for the task of joint mention extraction and classifi-cation. Unlike existing approaches, our model is able to effectively capture over-lapping mentions with unbounded lengths. The model is highly scalable, with a time complexity that is linear in the number of words in the input sentence and linear in the number of possible mention classes. Our model can be extended to additionally capture mention heads explicitly in a joint manner under the same time complexity. We demonstrate the effectiveness of our model through extensive experiments on standard datasets. 1 Wei Lu 0011, Dan Roth 0001 |
EMNLP | 1 |
| 2015 | Language Processing with Perl and Prolog: Theories, Implemetation, and Application Pierre M. Nugues (Lund University, Sweden) Springer (Cognitive technologies series, edited by A. Bundy et. al), 2014, Second Edition, xxv+662 pp; hardcover, ISBN 978-3-642-41463-3 $89.99; ebook, ISBN 978-3-642-41464-0, $69.99; doi 10.1007/978-3-642-41464-0
Wei Lu 0011 |
Comput. Linguistics | 1 |
| 2014 | Multilingual Semantic Parsing : Parsing Multiple Languages into Semantic Representations
Zhanming Jie, Wei Lu 0011 |
COLING | 2 |
| 2014 | Semantic Parsing with Relaxed Hybrid TreesabstractWe propose a novel model for parsing natural language sentences into their for-mal semantic representations. The model is able to perform integrated lexicon ac-quisition and semantic parsing, mapping each atomic element in a complete seman-tic representation to a contiguous word sequence in the input sentence in a re-cursive manner, where certain overlap-pings amongst such word sequences are allowed. It defines distributions over the novel relaxed hybrid tree structures which jointly represent both sentences and se-mantics. Such structures allow tractable dynamic programming algorithms to be developed for efficient learning and decod-ing. Trained under a discriminative set-ting, our model is able to incorporate a rich set of features where certain unbounded long-distance dependencies can be cap-tured in a principled manner. We demon-strate through experiments that by exploit-ing a large collection of simple features, our model is shown to be competitive to previous works and achieves state-of-the-art performance on standard benchmark data across four different languages. The system and code can be downloaded from Wei Lu 0011 |
EMNLP | 1 |
| 2012 | Automatic Event Extraction with Structured Preference Modeling
Wei Lu 0011, Dan Roth 0001 |
ACL (1) | 1 |
| 2012 | Joint Inference for Event Timeline Construction
Quang Do, Wei Lu 0011, Dan Roth 0001 |
EMNLP-CoNLL | 2 |
| 2011 | A Probabilistic Forest-to-String Model for Language Generation from Typed Lambda Calculus Expressions
Wei Lu 0011, Hwee Tou Ng |
EMNLP | 1 |
| 2010 | Better Punctuation Prediction with Dynamic Conditional Random Fields
Wei Lu 0011, Hwee Tou Ng |
EMNLP | 1 |
| 2009 | Natural Language Generation with Tree Conditional Random Fields
Wei Lu 0011, Hwee Tou Ng, Wee Sun Lee |
EMNLP | 1 |
| 2008 | A Generative Model for Parsing Natural Language to Meaning Representations
Wei Lu 0011, Hwee Tou Ng, Wee Sun Lee, Luke Zettlemoyer |
EMNLP | 1 |
| 2007 | Supervised categorization of JavaScriptTM using program analysis features
Wei Lu 0011, Min-Yen Kan |
Inf. Process. Manag. | 1 |