Lidong Bing

dblp:53/6625 · DBLP profile ↗
← Back
149ranked-venue papers
17as first author
75since 2021 · last 2026
0000-0003-4565-6313ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 137 · 13 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 22 · 7 first-author · 2 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
abstract
Universal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and lack diversity in negative samples. Moreover, the embeddings exhibit limited discriminative ability in distinguishing false and hard negatives. In this paper, we leverage the advanced understanding capabilities of MLLMs to enhance representation learning, and present a novel Universal Multimodal Embedding(UniME-V2) model. Our approach first constructs a potential hard negative set through global retrieval. We then introduce the MLLM-as-a-Judge mechanism, which utilizes MLLMs to assess the semantic alignment of query-candidate pairs and generate soft semantic matching scores. These scores serve as a foundation for hard negative mining, mitigating the impact of false negatives and enabling the identification of diverse, high-quality hard negatives. Furthermore, the semantic matching scores are used as soft labels to mitigate the rigid one-to-one mapping constraint. By aligning the similarity matrix with the soft semantic matching score matrix, the model learns semantic distinctions among candidates, significantly enhancing its discriminative capacity. To further improve performance, we propose UniME-V2, a reranking model trained on our mined hard negatives through a joint pairwise and listwise optimization approach. We conduct comprehensive experiments on the MMEB benchmark and multiple retrieval tasks, demonstrating that our method achieves state-of-the-art performance across all tasks.
Tiancheng Gu, Kaicheng Yang 0002, Kaichen Zhang, Xiang An, Ziyong Feng, Tom Weidong Cai, Jiankang Deng, Lidong Bing
AAAI9
2026 EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
abstract
Chuanrui Hu, Xingze Gao, Zuyi Zhou, Dannong Xu, Yi Bai, Xintong Li, Hui Zhang, Tong Li, Chong Zhang, Lidong Bing, Yafeng Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chuanrui Hu, Xingze Gao, Zuyi Zhou, Dannong Xu, Hui Zhang 0093, Lidong Bing, Yafeng Deng
ACL (1)10
2025 FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
abstract
Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guizhen Chen, Weiwen Xu, Hao Zhang 0048, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong 0001
ACL (1)6
2025 Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
abstract
Large language models excel at problemsolving but often struggle with complex reasoning and factual accuracy.While chainof-thought and retrieval-augmented generation help break down problems and retrieve knowledge, they still falter on challenging tasks like competitive programming due to frequent reasoning errors and irrelevant retrieval.To address this, we introduce Critic-guided planning with Retrieval-augmentation, CR-Planner, a novel framework that leverages fine-tuned critic models to guide both reasoning and retrieval processes through planning.CR-Planner iteratively selects and executes sub-goals, guided by critic models.A sub-goal critic identifies promising sub-goals from reasoning, query generation, and retrieval, while an execution critic evaluates outputs of sub-goal executions.We employ Monte Carlo Tree Search to collect data for critic training, allowing systematic exploration of action sequences and effective navigation toward the final answer.We evaluate CR-Planner on challenging domain-knowledgeintensive and reasoning-heavy tasks, including competitive programming, theorem-driven math reasoning, and complex domain retrieval problems.It significantly outperforms baselines, demonstrating effectiveness in both reasoning and retrieval.Our code is available at https://github.com/xingxuanli/CR-Planner.
Xingxuan Li, Weiwen Xu, Fangkai Jiao, Shafiq R. Joty, Lidong Bing
ACL (1)6
2025 Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
abstract
While understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research on the knowledge boundaries of LLMs has predominantly focused on English. In this work, we present the first study to analyze how LLMs recognize knowledge boundaries across different languages by probing their internal representations when processing known and unknown questions in multiple languages. Our empirical studies reveal three key findings: 1) LLMs' perceptions of knowledge boundaries are encoded in the middle to middle-upper layers across different languages. 2) Language differences in knowledge boundary perception follow a linear structure, which motivates our proposal of a training-free alignment method that effectively transfers knowledge boundary perception ability across languages, thereby helping reduce hallucination risk in low-resource languages; 3) Fine-tuning on bilingual question pair translation further enhances LLMs' recognition of knowledge boundaries across languages. Given the absence of standard testbeds for cross-lingual knowledge boundary analysis, we construct a multilingual evaluation suite comprising three representative types of knowledge boundary data. Our code and datasets are publicly available at https://github.com/DAMO-NLP-SG/ LLM-Multilingual-Knowledge-Boundaries.
Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Mahani Aljunied, Lidong Bing, Noura Al Moubayed, Yu Rong 0001
ACL (1)5
2025 Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization
abstract
Yao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Xiaoli Li, Roy Ka-Wei Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Roy Ka-Wei Lee
ACL (1)5
2025 Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
abstract
As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly.Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting platforms like Chatbot Arena.However, human evaluations require significant manual effort.Therefore, we propose Auto-Arena, an innovative framework that automates the entire evaluation process using LLM-powered agents.Firstly, an LLM examiner generates questions.Then, two LLM candidates engage in a multi-round peer battle based on the questions, aiming at revealing their true performance differences.Finally, a committee of LLM judges collaboratively discusses and decides the winner, reducing bias and enhancing fairness.During the peer battles, we observe intriguing scenarios where the LLM candidates display competitive behaviors and learn from the opponents.In our extensive experiments involving 15 recent LLMs, Auto-Arena shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts.Auto-Arena offers a promising alternative to current human evaluation platforms for evaluating LLMs automatically. 1
Wenxuan Zhang 0001, Yew Ken Chia, Weiwen Xu, Deli Zhao, Lidong Bing
ACL (1)6
2025 Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
abstract
Large Language Models (LLMs) have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. However, this paradigm is limited by the availability of gold labels, while in certain scenarios, LLMs may need to perform tasks that are too complex for humans to provide such labels. To tackle this challenge, this study explores whether solely utilizing unlabeled data can elicit strong model capabilities. We propose a new paradigm termed zero-to-strong generalization. We iteratively prompt LLMs to annotate unlabeled data and retain high-quality labels by filtering. Surprisingly, we obverse that this iterative process gradually unlocks LLMs’ potential on downstream tasks. Our experiments on extensive classification and reasoning tasks confirm the effectiveness of our proposed framework. Our analysis indicates that this paradigm is effective for both in-context learning and fine-tuning, and for various model sizes.
Chaoqun Liu, Qin Chao, Wenxuan Zhang 0001, Xiaobao Wu, Boyang Li 0001, Anh Tuan Luu, Lidong Bing
COLING7
2025 Breaking the Memory Barrier of Contrastive Loss via Tile-Based Strategy
abstract
Contrastive loss is a powerful approach for representation learning, where larger batch sizes enhance performance by providing more negative samples to better distinguish between similar and dissimilar data. However, the full instantiation of the similarity matrix demands substantial GPU memory, making large batch training highly resource-intensive. To address this, we propose a tile-based computation strategy that partitions the contrastive loss calculation into small blocks, avoiding full materialization of the similarity matrix. Additionally, we introduce a multi-level tiling implementation to leverage the hierarchical structure of distributed systems, using ring-based communication at the GPU level to optimize synchronization and fused kernels at the CUDA core level to reduce I/O overhead. Experimental results show that the proposed method significantly reduces GPU memory usage in contrastive loss. For instance, it enables contrastive training of a CLIP-ViT-L/14 model with a batch size of 4M using only 8 A800 80GB GPUs, without sacrificing accuracy. Compared to state-of-the-art memory-efficient solutions, it achieves a two-order-of-magnitude reduction in memory while maintaining comparable speed. The code will be made publicly available.1
Zesen Cheng, Sicong Leng, Deli Zhao, Xin Li 0056, Lidong Bing
CVPR9
2025 ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
abstract
The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code is available at https://github.com/RhDang/ECBench.
Ronghao Dang, Yuqian Yuan, Wenqi Zhang 0001, Yifei Xin, Boqiang Zhang, Liuyi Wang, Qinyang Zeng, Xin Li 0056, Lidong Bing
CVPR10
2025 VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
abstract
Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a largescale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities.
Yuqian Yuan, Wentong Li 0001, Zesen Cheng, Boqiang Zhang, Xin Li 0056, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing
CVPR12
2025 M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
abstract
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing
EMNLP8
2025 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
abstract
Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook.
Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing
ICCV9
2025 LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities through pretraining and alignment. However, superior short-context LLMs may underperform in long-context scenarios due to insufficient long-context alignment. This alignment process remains challenging due to the impracticality of human annotation for extended contexts and the difficulty in balancing short- and long-context performance. To address these challenges, we introduce LongPO, that enables short-context LLMs to self-evolve to excel on long-context tasks by internally transferring short-context capabilities. LongPO harnesses LLMs to learn from self-generated short-to-long preference data, comprising paired responses generated for identical instructions with long-context inputs and their compressed short-context counterparts, respectively. This preference reveals capabilities and potentials of LLMs cultivated during short-context alignment that may be diminished in under-aligned long-context scenarios. Additionally, LongPO incorporates a short-to-long KL constraint to mitigate short-context performance decline during long-context alignment. When applied to Mistral-7B-Instruct-v0.2 from 128K to 512K context lengths, LongPO fully retains short-context performance and largely outperforms naive SFT and DPO in both long- and short-context tasks. Specifically, LongPO-trained models can achieve results on long-context benchmarks comparable to, or even surpassing, those of superior LLMs (e.g., GPT-4-128K) that involve extensive long-context annotation and larger parameter scales. Our code is available at https://github.com/DAMO-NLP-SG/LongPO.
Guanzheng Chen, Xin Li 0056, Michael Shieh, Lidong Bing
ICLR4
2025 Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
abstract
We propose a novel prompt design paradigm that challenges conventional wisdom in large language model (LLM) prompting. While conventional wisdom prioritizes well-crafted instructions and demonstrations for in-context learning (ICL), we show that pruning random demonstrations into seemingly incoherent ''gibberish'' can remarkably improve performance across diverse tasks. Notably, the ''gibberish'' always matches or surpasses state-of-the-art automatic prompt optimization techniques, achieving substantial gains regardless of LLM alignment. Nevertheless, discovering an effective pruning strategy is non-trivial, as existing attribution methods and prompt compression algorithms fail to deliver robust results, let alone human intuition. In terms of this, we propose a self-discover prompt optimization framework, PromptQuine, an evolutionary search framework that automatically searches for the pruning strategy by itself using only low-data regimes. Much like the emergent complexity in nature—such as symbiosis and self-organization—arising in response to resource constraints, our framework evolves and refines unconventional yet highly effective prompts by leveraging only the tokens present within the context. We demonstrate its effectiveness across classification, multi-choice question answering, generation and math reasoning tasks across LLMs, while achieving decent runtime efficiency. We hope our findings can guide mechanistic studies on in-context learning, and provide a call to action, to pave the way for more open-ended search algorithms for more effective LLM prompting.
Lidong Bing
ICML3
2025 ParaICL: Towards Parallel In-Context Learning
abstract
Xingxuan Li, Xuan-Phi Nguyen, Shafiq Joty, Lidong Bing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xingxuan Li, Xuan-Phi Nguyen, Shafiq R. Joty, Lidong Bing
NAACL (Long Papers)4
2025 Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
abstract
Chaoqun Liu, Wenxuan Zhang, Yiran Zhao, Anh Tuan Luu, Lidong Bing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Chaoqun Liu, Wenxuan Zhang 0001, Yiran Zhao 0006, Anh Tuan Luu, Lidong Bing
NAACL (Long Papers)5
2025 AdaMergeX: Cross-Lingual Transfer with Large Language Models via Adaptive Adapter Merging
abstract
Yiran Zhao, Wenxuan Zhang, Huiming Wang, Kenji Kawaguchi, Lidong Bing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yiran Zhao 0006, Wenxuan Zhang 0001, Kenji Kawaguchi, Lidong Bing
NAACL (Long Papers)5
2025 The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
abstract
Recent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most existing LMMs remain vulnerable to hallucinations, the discrepancy between the factual multimodal input and the generated textual output, which has limited their applicability in various real-world scenarios. This paper presents the first systematic investigation of hallucinations in LMMs involving the three most common modalities: language, visual, and audio. Our study reveals two key contributors to hallucinations: overreliance on unimodal priors and spurious inter-modality correlations. To address these challenges, we introduce the benchmark The Curse of Multi-Modalities (CMM), which comprehensively evaluates hallucinations in LMMs, providing a detailed analysis of their underlying issues. Our findings highlight key vulnerabilities, including imbalances in modality integration and biases from training data, underscoring the need for balanced cross-modal learning and enhanced hallucination mitigation strategies. Based on our observations and findings, we suggest potential research directions that could enhance the reliability of LMMs.
Sicong Leng, Zesen Cheng, Xin Li 0056, Deli Zhao, Shijian Lu, Chunyan Miao, Lidong Bing
NeurIPS10
2025 MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
abstract
Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothesis discovery, which entails generating detailed, experimentally actionable hypotheses from coarse initial research directions. We frame this as a combinatorial optimization problem and investigate the upper limits of LLMs' capacity to solve it when maximally leveraged. Specifically, we explore four foundational questions: (1) how to best harness an LLM's internal heuristics to formulate the fine-grained hypothesis it itself would judge as the most promising among all the possible hypotheses it might generate, based on its own internal scoring-thus defining a latent reward landscape over the hypothesis space; (2) whether such LLM-judged better hypotheses exhibit stronger alignment with ground-truth hypotheses; (3) whether shaping the reward landscape using an ensemble of diverse LLMs of similar capacity yields better outcomes than defining it with repeated instances of the strongest LLM among them; and (4) whether an ensemble of identical LLMs provides a more reliable reward landscape than a single LLM. To address these questions, we propose a hierarchical search method that incrementally proposes and integrates details into the hypothesis, progressing from general concepts to specific experimental configurations. We show that this hierarchical process smooths the reward landscape and enables more effective optimization. Empirical evaluations on a new benchmark of expert-annotated fine-grained hypotheses from recent literature show that our method consistently outperforms strong baselines.
Zonglin Yang 0001, Wanhao Liu, Ben Gao, Wei Li 0076, Tong Xie, Lidong Bing, Wanli Ouyang, Erik Cambria, Dongzhan Zhou
NeurIPS7
2024 Exploring the Potential of Large Language Models in Computational Argumentation
abstract
Computational argumentation has become an essential tool in various domains, including law, public policy, and artificial intelligence.It is an emerging research field in natural language processing that attracts increasing attention.Research on computational argumentation mainly involves two types of tasks: argument mining and argument generation.As large language models (LLMs) have demonstrated impressive capabilities in understanding context and generating natural language, it is worthwhile to evaluate the performance of LLMs on diverse computational argumentation tasks.This work aims to embark on an assessment of LLMs, such as ChatGPT, Flan models, and LLaMA2 models, in both zero-shot and few-shot settings.We organize existing tasks into six main categories and standardize the format of fourteen openly available datasets.In addition, we present a new benchmark dataset on counter speech generation that aims to holistically evaluate the end-to-end performance of LLMs on argument mining and argument generation.Extensive experiments show that LLMs exhibit commendable performance across most of the datasets, demonstrating their capabilities in the field of argumentation.Our analysis offers valuable suggestions for evaluating computational argumentation and its integration with LLMs in future research endeavors.1
Guizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong Bing
ACL (1)4
2024 Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
abstract
Large language models (LLMs) are known to perform tasks by simply observing few exemplars.Moreover, competent generative capabilities of LLMs are observed mostly in highresource languages, while their performances among under-represented languages fall behind due to pre-training data imbalance.To elicit LLMs' ability onto low-resource languages without any supervised data, we propose to assemble synthetic exemplars from a diverse set of high-resource languages.These prompts can directly induce generative capabilities in lowresource languages and serve as intra-lingual exemplars to even improve tasks in these languages.Our unsupervised prompting method performs on par with supervised few-shot learning in LLMs of different sizes for translations between English and 34 Indic and African languages, and surpasses supervised prompting in non-English tasks.The method also significantly improves low-resource performances in many other intra-lingual tasks like summarization (XLSum), question answering (XQUAD & TydiQA) and conversational instruction following (Sea-Bench).
Xuan-Phi Nguyen, Mahani Aljunied, Shafiq R. Joty, Lidong Bing
ACL (1)4
2024 Order-Agnostic Data Augmentation for Few-Shot Named Entity Recognition
abstract
Data augmentation (DA) methods have been proven to be effective for pre-trained language models (PLMs) in low-resource settings, including few-shot named entity recognition (NER).However, existing NER DA techniques either perform rule-based manipulations on words that break the semantic coherence of the sentence, or exploit generative models for entity or context substitution, which requires a substantial amount of labeled data and contradicts the objective of operating in low-resource settings.In this work, we propose orderagnostic data augmentation (OADA), an alternative solution that exploits the often overlooked order-agnostic property in the training data construction phase of sequence-tosequence NER methods for data augmentation.To effectively utilize the augmented data without suffering from the one-to-many issue, where multiple augmented target sequences exist for one single sentence, we further propose the use of ordering instructions and an innovative OADA-XE loss.Specifically, by treating each permutation of entity types as an ordering instruction, we rearrange the entity set accordingly, ensuring a distinct input-output pair, while OADA-XE assigns loss based on the best match between the target sequence and model predictions.We conduct comprehensive experiments and analyses across three major NER benchmarks and can significantly enhance the few-shot capabilities of PLMs with OADA.Our code is available at https://github.com/Circle-Ming/OADA-NER.
Liying Cheng, Wenxuan Zhang 0001, De Wen Soh, Lidong Bing
ACL (1)5
2024 Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
abstract
Large Vision-Language Models (LVLMs) have advanced considerably, intertwining visual recognition and language understanding to generate content that is not only coherent but also contextually attuned. Despite their success, LVLMs still suffer from the issue of object hallucinations, where models generate plausible yet incorrect outputs that include objects that do not exist in the images. To mitigate this issue, we introduce Visual Contrastive Decoding (VCD), a simple and training-free method that contrasts output distributions derived from original and distorted visual inputs. The proposed VCD effectively reduces the over-reliance on statistical bias and unimodal priors, two essential causes of object hallucinations. This adjustment ensures the generated content is closely grounded to visual inputs, resulting in contextually accurate outputs. Our experiments show that VCD, without either additional training or the usage of external tools, significantly mitigates the object hallucination issue across different LVLM families. Beyond mitigating object hallucinations, VCD also excels in general LVLM benchmarks, highlighting its wide-ranging applicability.
Sicong Leng, Guanzheng Chen, Xin Li 0056, Shijian Lu, Chunyan Miao, Lidong Bing
CVPR7
2024 Evaluating Psychological Safety of Large Language Models
abstract
In this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs).First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI).All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern.Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3.Then, we evaluated the LLMs in the GPT series by using well-being tests to study the impact of fine-tuning with more training data.We observed a continuous increase in the well-being scores of GPT models.Following these observations, we showed that finetuning Llama-2-chat-7B with responses from BFI using direct preference optimization could effectively reduce the psychological toxicity of the model.Based on the findings, we recommended the application of systematic and comprehensive psychological metrics to further evaluate and improve the safety of LLMs.Our code is available at https://github.com/DAMO- NLP-SG/PsychSafety.Warning: This paper contains examples with potentially harmful content.
Xingxuan Li, Shafiq R. Joty, Lidong Bing
EMNLP5
2024 AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code Generation
abstract
The impressive performance of proprietary LLMs like GPT4 in code generation has led to a trend to replicate these capabilities in open-source models through knowledge distillation (e.g.Code Evol-Instruct).However, these efforts often neglect the crucial aspect of response quality, relying heavily on teacher models for direct response distillation.This paradigm, especially for complex instructions, can degrade the quality of synthesized data, compromising the knowledge distillation process.To this end, our study introduces the Adaptive Modular Response Evolution (AMR-Evol) framework, which employs a two-stage process to refine response distillation.The first stage, modular decomposition, breaks down the direct response into more manageable sub-modules.The second stage, adaptive response evolution, automatically evolves the response with the related function modules.Our experiments with three popular code benchmarks-HumanEval, MBPP, and EvalPlus-attests to the superiority of the AMR-Evol framework over baseline response distillation methods.By comparing with the open-source Code LLMs trained on a similar scale of data, we observed performance enhancements: more than +3.0 points on HumanEval-Plus and +1.0 points on MBPP-Plus, which underscores the effectiveness of our framework.Our codes are available at https://github.com/ChiYeungLaw/ AMR-Evol. Module 3def validate_matrix(matrix: list) -> None:""" Description:Validates if the input matrix is a 3x3 matrix.Parameters:-matrix (list): The input matrix to be validated.Raises:-ValueError: If the matrix is not a 3x3 matrix.""" Coding Tasks Design a function that implements the Laplace expansion theorem to calculate the determinant of a 3x3 matrix, but now with the added condition of utilizing nested for loops and ensuring that the matrix elements are limited to integers only.Teacher Model Response def calculate_determinant(matrix): if len(matrix) != 3 or len(matrix[0]) != 3: raise ValueError('Matrix must be a 3x3 matrix') determinant = 0 for i in range(3): submatrix = [[matrix[j][k] for k in range(3) if k != i] for j in range(1, 3)] sub_determinant = calculate_determinant(submatrix) sign = (-1) ** i determinant += sign * matrix[0][i] * sub_determinant return determinant if len(matrix) == 2 else round(determinant) Direct Distillation Module 2def calculate_minor_matrix(matrix: list, row: int, col: int) -> list:""" Description: Calculates the minor matrix by removing the specified row Parameters:-matrix (list): The input matrix.-row (int): The row index to be removed.-col (int): The column index to be removed.
Xin Li 0056, Hongzhan Lin 0001, Jing Ma 0004, Lidong Bing
EMNLP5
2024 Multilingual Jailbreak Challenges in Large Language Models
abstract
While large language models (LLMs) exhibit remarkable capabilities across a wide range of tasks, they pose potential safety concerns, such as the ``jailbreak'' problem, wherein malicious instructions can manipulate LLMs to exhibit undesirable behavior. Although several preventive measures have been developed to mitigate the potential risks associated with LLMs, they have primarily focused on English. In this study, we reveal the presence of multilingual jailbreak challenges within LLMs and consider two potential risky scenarios: unintentional and intentional. The unintentional scenario involves users querying LLMs using non-English prompts and inadvertently bypassing the safety mechanisms, while the intentional scenario concerns malicious users combining malicious instructions with multilingual prompts to deliberately attack LLMs. The experimental results reveal that in the unintentional scenario, the rate of unsafe content increases as the availability of languages decreases. Specifically, low-resource languages exhibit about three times the likelihood of encountering harmful content compared to high-resource languages, with both ChatGPT and GPT-4. In the intentional scenario, multilingual prompts can exacerbate the negative impact of malicious instructions, with astonishingly high rates of unsafe output: 80.92\% for ChatGPT and 40.71\% for GPT-4. To handle such a challenge in the multilingual context, we propose a novel \textsc{Self-Defense} framework that automatically generates multilingual training data for safety fine-tuning. Experimental results show that ChatGPT fine-tuned with such data can achieve a substantial reduction in unsafe content generation. Data is available at \url{https://github.com/DAMO-NLP-SG/multilingual-safety-for-LLMs}.
Yue Deng 0010, Wenxuan Zhang 0001, Sinno Jialin Pan, Lidong Bing
ICLR4
2024 CLEX: Continuous Length Extrapolation for Large Language Models
abstract
Transformer-based Large Language Models (LLMs) are pioneering advances in many natural language processing tasks, however, their exceptional capabilities are restricted within the preset context window of Transformer. Position Embedding (PE) scaling methods, while effective in extending the context window to a specific length, demonstrate either notable limitations in their extrapolation abilities or sacrificing partial performance within the context window. Length extrapolation methods, although theoretically capable of extending the context window beyond the training sequence length, often underperform in practical long-context applications. To address these challenges, we propose Continuous Length EXtrapolation (CLEX) for LLMs. We generalise the PE scaling approaches to model the continuous dynamics by ordinary differential equations over the length scaling factor, thereby overcoming the constraints of current PE scaling methods designed for specific lengths. Moreover, by extending the dynamics to desired context lengths beyond the training sequence length, CLEX facilitates the length extrapolation with impressive performance in practical tasks. We demonstrate that CLEX can be seamlessly incorporated into LLMs equipped with Rotary Position Embedding, such as LLaMA and GPT-NeoX, with negligible impact on training and inference latency. Experimental results reveal that CLEX can effectively extend the context window to over 4× or almost 8× training length, with no deterioration in performance. Furthermore, when evaluated on the practical LongBench benchmark, our model trained on a 4k length exhibits competitive performance against state-of-the-art open-source models trained on context lengths up to 32k. Our code is available at https://github.com/DAMO-NLP-SG/CLEX.
Guanzheng Chen, Xin Li 0056, Zaiqiao Meng, Shangsong Liang, Lidong Bing
ICLR5
2024 Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
abstract
We present chain-of-knowledge (CoK), a novel framework that augments large language models (LLMs) by dynamically incorporating grounding information from heterogeneous sources. It results in more factual rationales and reduced hallucination in generation. Specifically, CoK consists of three stages: reasoning preparation, dynamic knowledge adapting, and answer consolidation. Given a knowledge-intensive question, CoK first prepares several preliminary rationales and answers while identifying the relevant knowledge domains. If there is no majority consensus among the answers from samples, CoK corrects the rationales step by step by adapting knowledge from the identified domains. These corrected rationales can plausibly serve as a better foundation for the final answer consolidation. Unlike prior studies that primarily use unstructured data, CoK also leverages structured knowledge sources such as Wikidata and tables that provide more reliable factual information. To access both unstructured and structured knowledge sources in the dynamic knowledge adapting stage, we propose an adaptive query generator that allows the generation of queries for various types of query languages, including SPARQL, SQL, and natural sentences. Moreover, to minimize error propagation between rationales, CoK corrects the rationales progressively using preceding corrected rationales to generate and correct subsequent rationales. Extensive experiments show that CoK consistently improves the performance of LLMs on knowledge-intensive tasks across different domains.
Xingxuan Li, Yew Ken Chia, Bosheng Ding, Shafiq R. Joty, Soujanya Poria, Lidong Bing
ICLR7
2024 Large Language Models can Contrastively Refine their Generation for Better Sentence Representation Learning
abstract
Huiming Wang, Zhaodonghui Li, Liying Cheng, De Wen Soh, Lidong Bing. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhaodonghui Li, Liying Cheng, De Wen Soh, Lidong Bing
NAACL-HLT5
2024 Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective
abstract
Latent-based image generative models, such as Latent Diffusion Models (LDMs) and Mask Image Models (MIMs), have achieved notable success in image generation tasks. These models typically leverage reconstructive autoencoders like VQGAN or VAE to encode pixels into a more compact latent space and learn the data distribution in the latent space instead of directly from pixels. However, this practice raises a pertinent question: Is it truly the optimal choice? In response, we begin with an intriguing observation: despite sharing the same latent space, autoregressive models significantly lag behind LDMs and MIMs in image generation. This finding contrasts sharply with the field of NLP, where the autoregressive model GPT has established a commanding presence. To address this discrepancy, we introduce a unified perspective on the relationship between latent space and generative models, emphasizing the stability of latent space in image generative modeling. Furthermore, we propose a simple but effective discrete image tokenizer to stabilize the latent space for image generative modeling by applying K-Means on the latent features of self-supervised learning models. Experimental results show that image autoregressive modeling with our tokenizer (DiGIT) benefits both image understanding and image generation with the next token prediction principle, which is inherently straightforward for GPT models but challenging for other generative models. Remarkably, for the first time, a GPT-style autoregressive model for images outperforms LDMs, which also exhibits substantial improvement akin to GPT when scaling up model size. Our findings underscore the potential of an optimized latent space and the integration of discrete tokenization in advancing the capabilities of image generative models. The code is available at \url{https://github.com/DAMO-NLP-SG/DiGIT}.
Yongxin Zhu 0003, Bocheng Li, Xin Li 0056, Linli Xu 0002, Lidong Bing
NeurIPS6
2024 How do Large Language Models Handle Multilingualism?
abstract
Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow ($\texttt{MWork}$): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify $\texttt{MWork}$, we introduce Parallel Language-specific Neuron Detection ($\texttt{PLND}$) to identify activated neurons for inputs in different languages without any labeled data. Using $\texttt{PLND}$, we validate $\texttt{MWork}$ through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, $\texttt{MWork}$ allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of $3.6\%$ for high-resource languages and $2.3\%$ for low-resource languages across all tasks with just $400$ documents.
Yiran Zhao 0006, Wenxuan Zhang 0001, Guizhen Chen, Kenji Kawaguchi, Lidong Bing
NeurIPS5
2024 LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency
abstract
Query rewrite, which aims to improve query efficiency by altering an SQL query's structure without changing its result, has been an important research problem. In order to maintain equivalence between the rewritten query and the original one during rewriting, traditional query rewrite methods always rewrite the queries following certain rewrite rules. However, some problems still remain. First, existing methods of finding the optimal choice or sequence of rewrite rules are still limited and the process always costs a lot of resources. Methods involving discovering new rewrite rules typically require complicated proofs of structural logic or extensive user interactions. Second, current query rewrite methods usually rely highly on DBMS cost estimators which are often not accurate. In this paper, we address these problems by proposing a novel query rewrite method named LLM-R 2 , which leverages a large language model (LLM) to recommend rewrite rules for a database rewrite system. To further enhance the inference ability of the LLM in recommending rewrite rules, we train a contrastive model using a curriculum-based approach to learn query representations and select effective query demonstrations for the LLM. Experimental results show that our method significantly improves the query execution efficiency and outperforms the baseline methods. In addition, our method exhibits high robustness across different datasets.
Zhaodonghui Li, Haitao Yuan 0002, Gao Cong, Lidong Bing
Proc. VLDB Endow.5
2023 On the Effectiveness of Parameter-Efficient Fine-Tuning
abstract
Fine-tuning pre-trained models has been ubiquitously proven to be effective in a wide range of NLP tasks. However, fine-tuning the whole model is parameter inefficient as it always yields an entirely new model for each task. Currently, many research works propose to only fine-tune a small portion of the parameters while keeping most of the parameters shared across different tasks. These methods achieve surprisingly good performance and are shown to be more stable than their corresponding fully fine-tuned counterparts. However, such kind of methods is still not well understood. Some natural questions arise: How does the parameter sparsity lead to promising performance? Why is the model more stable than the fully fine-tuned models? How to choose the tunable parameters? In this paper, we first categorize the existing methods into random approaches, rule-based approaches, and projection-based approaches based on how they choose which parameters to tune. Then, we show that all of the methods are actually sparse fine-tuned models and conduct a novel theoretical analysis of them. We indicate that the sparsity is actually imposing a regularization on the original model by controlling the upper bound of the stability. Such stability leads to better generalization capability which has been empirically observed in a lot of recent research works. Despite the effectiveness of sparsity grounded by our theory, it still remains an open problem of how to choose the tunable parameters. Currently, the random and rule-based methods do not utilize task-specific data information while the projection-based approaches suffer from the projection discontinuity problem. To better choose the tunable parameters, we propose a novel Second-order Approximation Method (SAM) which approximates the original problem with an analytically solvable optimization function. The tunable parameters are determined by directly optimizing the approximation function. We conduct extensive experiments on several tasks. The experimental results show that our proposed SAM model outperforms many strong baseline models and it also verifies our theoretical analysis. The source code of this paper can be obtained from https://github.com/fuzihaofzh/AnalyzeParameterEff\/icientFinetune .
Anthony Man-Cho So, Wai Lam, Lidong Bing, Nigel Collier
AAAI5
2023 Improving Self-training for Cross-lingual Named Entity Recognition with Contrastive and Prototype Learning
abstract
In cross-lingual named entity recognition (NER), self-training is commonly used to bridge the linguistic gap by training on pseudolabeled target-language data.However, due to sub-optimal performance on target languages, the pseudo labels are often noisy and limit the overall performance.In this work, we aim to improve self-training for cross-lingual NER by combining representation learning and pseudo label refinement in one coherent framework.Our proposed method, namely ContProto mainly comprises two components: (1) contrastive self-training and (2) prototype-based pseudo-labeling.Our contrastive self-training facilitates span classification by separating clusters of different classes, and enhances crosslingual transferability by producing closelyaligned representations between the source and target language.Meanwhile, prototype-based pseudo-labeling effectively improves the accuracy of pseudo labels during training.We evaluate ContProto on multiple transfer pairs, and experimental results show our method brings in substantial improvements over current stateof-the-art methods. 1
Ran Zhou 0004, Xin Li 0056, Lidong Bing, Erik Cambria, Chunyan Miao
ACL (1)3
2023 Bidirectional Generative Framework for Cross-domain Aspect-based Sentiment Analysis
abstract
Cross-domain aspect-based sentiment analysis (ABSA) aims to perform various fine-grained sentiment analysis tasks on a target domain by transferring knowledge from a source domain.Since labeled data only exists in the source domain, a model is expected to bridge the domain gap for tackling cross-domain ABSA.Though domain adaptation methods have proven to be effective, most of them are based on a discriminative model, which needs to be specifically designed for different ABSA tasks.To offer a more general solution, we propose a unified bidirectional generative framework to tackle various cross-domain ABSA tasks.Specifically, our framework trains a generative model in both text-to-label and label-to-text directions.The former transforms each task into a unified format to learn domain-agnostic features, and the latter generates natural sentences from noisy labels for data augmentation, with which a more accurate model can be trained.To investigate the effectiveness and generality of our framework, we conduct extensive experiments on four cross-domain ABSA tasks and present new state-of-the-art results on all tasks.Our data and code are publicly available at https://github.com
Yue Deng 0010, Wenxuan Zhang 0001, Sinno Jialin Pan, Lidong Bing
ACL (1)4
2023 Is GPT-3 a Good Data Annotator?
abstract
Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Boyang Li, Shafiq Joty, Lidong Bing. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Bosheng Ding, Chengwei Qin, Yew Ken Chia, Boyang Li 0001, Shafiq R. Joty, Lidong Bing
ACL (1)7
2023 Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed Representations
abstract
Due to the huge amount of parameters, finetuning of pretrained language models (PLMs) is prone to overfitting in the low resource scenarios.In this work, we present a novel method that operates on the hidden representations of a PLM to reduce overfitting.During fine-tuning, our method inserts random autoencoders between the hidden layers of a PLM, which transform activations from the previous layers into multi-view compressed representations before feeding them into the upper layers.The autoencoders are plugged out after fine-tuning, so our method does not add extra parameters or increase computation cost during inference.Our method demonstrates promising performance improvement across a wide range of sequenceand token-level low-resource NLP tasks.Our code is available at https://github.com/DAMO- NLP-SG/MVCR.
Xingxuan Li, Megh Thakkar, Xin Li 0056, Shafiq R. Joty, Luo Si, Lidong Bing
ACL (1)7
2023 Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
abstract
Reasoning about time is of fundamental importance.Many facts are time-dependent.For example, athletes change teams from time to time, and different government officials are elected periodically.Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types.In this paper, we introduce a comprehensive probing dataset TEMPREASON to evaluate the temporal reasoning capability of large language models.Our dataset includes questions of three temporal reasoning levels.In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning.We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach 1 .
Hwee Tou Ng, Lidong Bing
ACL (1)3
2023 Information Screening whilst Exploiting! Multimodal Relation Extraction with Feature Denoising and Multimodal Topic Modeling
abstract
Existing research on multimodal relation extraction (MRE) faces two co-existing challenges, internal-information over-utilization and external-information under-exploitation.To combat that, we propose a novel framework that simultaneously implements the idea of internal-information screening and externalinformation exploiting.First, we represent the fine-grained semantic structures of the input image and text with the visual and textual scene graphs, which are further fused into a unified cross-modal graph (CMG).Based on CMG, we perform structure refinement with the guidance of the graph information bottleneck principle, actively denoising the less-informative features.Next, we perform topic modeling over the input image and text, incorporating latent multimodal topic features to enrich the contexts.On the benchmark MRE dataset, our system outperforms the current best model significantly.With further in-depth analyses, we reveal the great potential of our method for the MRE task.Our codes are open at https://github.com/ChocoWu/MRE-ISE.
Shengqiong Wu, Hao Fei 0001, Yixin Cao 0002, Lidong Bing, Tat-Seng Chua
ACL (1)4
2023 PeerDA: Data Augmentation via Modeling Peer Relation for Span Identification Tasks
abstract
Span identification aims at identifying specific text spans from text input and classifying them into pre-defined categories.Different from previous works that merely leverage the Subordinate (SUB) relation (i.e. if a span is an instance of a certain category) to train models, this paper for the first time explores the Peer (PR) relation, which indicates that two spans are instances of the same category and share similar features.Specifically, a novel Peer Data Augmentation (PeerDA) approach is proposed which employs span pairs with the PR relation as the augmentation data for training.PeerDA has two unique advantages: (1) There are a large number of PR span pairs for augmenting the training data.(2) The augmented data can prevent the trained model from over-fitting the superficial span-category mapping by pushing the model to leverage the span semantics.Experimental results on ten datasets over four diverse tasks across seven domains demonstrate the effectiveness of PeerDA.Notably, PeerDA achieves state-of-the-art results on six of them. 1
Weiwen Xu, Xin Li 0056, Yang Deng 0002, Wai Lam, Lidong Bing
ACL (1)5
2023 Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework
abstract
As large language models (LLMs) have become the norm in NLP, demonstrating good performance in generation and reasoning tasks, one of its most fatal disadvantages is the lack of factual correctness.Generating unfactual texts not only leads to lower performances but also degrades the trust and validity of their applications.Chain-of-Thought (CoT) prompting improves trust and model performance on complex reasoning tasks by generating interpretable reasoning chains, but still suffers from factuality concerns in knowledge-intensive tasks.In this paper, we propose the Verify-and-Edit framework for CoT prompting, which seeks to increase prediction factuality by post-editing reasoning chains according to external knowledge.Building on top of GPT-3, our framework lead to accuracy improvements in multiple open-domain question-answering tasks.For reproducing our results and extending the framework further, we make our codebase available at https://github.com/RuochenZhao/Verify- and-Edit * Equal contribution.
Xingxuan Li, Shafiq R. Joty, Chengwei Qin, Lidong Bing
ACL (1)5
2023 Towards Integration of Discriminability and Robustness for Document-Level Relation Extraction
abstract
Document-level relation extraction (DocRE)predicts relations for entity pairs that rely on long-range context-dependent reasoning in a document.As a typical multi-label classification problem, DocRE faces the challenge of effectively distinguishing a small set of positive relations from the majority of negative ones.This challenge becomes even more difficult to overcome when there exists a significant number of annotation errors in the dataset.In this work, we aim to achieve better integration of both the discriminability and robustness for the DocRE problem.Specifically, we first design an effective loss function to endow high discriminability to both probabilistic outputs and internal representations.We innovatively customize entropy minimization and supervised contrastive learning for the challenging multi-label and long-tailed learning problems.To ameliorate the impact of label errors, we equipped our method with a novel negative label sampling strategy to strengthen the model robustness.In addition, we introduce two new data regimes to mimic more realistic scenarios with annotation errors and evaluate our sampling strategy.Experimental results verify the effectiveness of each component and show that our method achieves new state-ofthe-art results on the DocRED dataset, its recently cleaned version, Re-DocRED, and the proposed data regimes. 1
Stanley Kok, Lidong Bing
EACL3
2023 SOUL: Towards Sentiment and Opinion Understanding of Language
abstract
Sentiment analysis is a well-established natural language processing task, with sentiment polarity classification being one of its most popular and representative tasks.However, despite the success of pre-trained language models in this area, they often fall short of capturing the broader complexities of sentiment analysis.To address this issue, we propose a new task called Sentiment and Opinion Understanding of Language (SOUL).SOUL aims to evaluate sentiment understanding through two subtasks: Review Comprehension (RC) and Justification Generation (JG).RC seeks to validate statements that focus on subjective information based on a review text, while JG requires models to provide explanations for their sentiment predictions.To enable comprehensive evaluation, we annotate a new dataset comprising 15,028 statements from 3,638 reviews.Experimental results indicate that SOUL is a challenging task for both small and large language models, with a performance gap of up to 27% when compared to human performance.Furthermore, evaluations conducted with both human experts and GPT-4 highlight the limitations of the small language model in generating reasoning-based justifications.These findings underscore the challenging nature of the SOUL task for existing models, emphasizing the need for further advancements in sentiment analysis to address its complexities.The new dataset and code are available at https://github.com
Yue Deng 0010, Wenxuan Zhang 0001, Sinno Jialin Pan, Lidong Bing
EMNLP4
2023 LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
abstract
The success of large language models (LLMs), like GPT-4 and ChatGPT, has led to the development of numerous cost-effective and accessible alternatives that are created by finetuning open-access LLMs with task-specific data (e.g., ChatDoctor) or instruction data (e.g., Alpaca).Among the various fine-tuning methods, adapter-based parameter-efficient fine-tuning (PEFT) is undoubtedly one of the most attractive topics, as it only requires fine-tuning a few external parameters instead of the entire LLMs while achieving comparable or even better performance.To enable further research on PEFT methods of LLMs, this paper presents LLM-Adapters, an easy-to-use framework that integrates various adapters into LLMs and can execute these adapter-based PEFT methods of LLMs for different tasks.The framework includes state-of-the-art open-access LLMs such as LLaMA, BLOOM, and GPT-J, as well as widely used adapters such as Series adapters, Parallel adapter, Prompt-based learning and Reparametrization-based methods.Moreover, we conduct extensive empirical studies on the impact of adapter types, placement locations, and hyper-parameters to the best design for each adapter-based methods.We evaluate the effectiveness of the adapters on fourteen datasets from two different reasoning tasks, Arithmetic Reasoning and Commonsense Reasoning.The results demonstrate that using adapter-based PEFT in smaller-scale LLMs (7B) with few extra trainable parameters yields comparable, and in some cases superior, performance to powerful LLMs (175B) in zero-shot inference on both reasoning tasks.The code and datasets can be found in https://github. com/AGI-Edgerunners/LLM-Adapters.
Lei Wang 0185, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu 0001, Soujanya Poria, Roy Ka-Wei Lee
EMNLP6
2023 Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal Reasoning
abstract
Our physical world is constantly evolving over time, rendering challenges for pre-trained language models to understand and reason over the temporal contexts of texts.Existing work focuses on strengthening the direct association between a piece of text and its time-stamp.However, the knowledge-time association is usually insufficient for the downstream tasks that require reasoning over temporal dependencies between knowledge.In this work, we make use of the underlying nature of time, all temporally-scoped sentences are strung together through a one-dimensional time axis, and suggest creating a graph structure based on the relative placements of events along the time axis.Inspired by the graph view, we propose REMEMO (Relative Time Modeling), which explicitly connects all temporally-scoped facts by modeling the time relations between any two sentences.Experimental results show that REMEMO outperforms the baseline T5 on multiple temporal question answering datasets under various settings.Further analysis suggests that REMEMO is especially good at modeling long-range complex temporal dependencies.We release our code and pretrained checkpoints at https://github.com/ DAMO-NLP-SG/RemeMo.
Sen Yang 0005, Xin Li 0056, Lidong Bing, Wai Lam
EMNLP3
2023 From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine Reader
abstract
We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy between model pre-training and downstream fine-tuning of existing MLMs. To build the proposed PMR, we constructed a large volume of general-purpose and high-quality MRC-style training data by using Wikipedia hyperlinks and designed a Wiki Anchor Extraction task to guide the MRC-style pre-training. Apart from its simplicity, PMR effectively solves extraction tasks, such as Extractive Question Answering and Named Entity Recognition. PMR shows tremendous improvements over existing approaches, especially in low-resource scenarios. When applied to the sequence classification task in the MRC formulation, PMR enables the extraction of high-quality rationales to explain the classification process, thereby providing greater prediction explainability. PMR also has the potential to serve as a unified model for tackling various extraction and classification tasks in the MRC formulation.
Weiwen Xu, Xin Li 0056, Wenxuan Zhang 0001, Wai Lam, Luo Si, Lidong Bing
NeurIPS7
2023 M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
abstract
Despite the existence of various benchmarks for evaluating natural language processing models, we argue that human exams are a more suitable means of evaluating general intelligence for large language models (LLMs), as they inherently demand a much wider range of abilities such as language understanding, domain knowledge, and problem-solving skills. To this end, we introduce M3Exam, a novel benchmark sourced from real and official human exam questions for evaluating LLMs in a multilingual, multimodal, and multilevel context. M3Exam exhibits three unique characteristics: (1) multilingualism, encompassing questions from multiple countries that require strong multilingual proficiency and cultural knowledge; (2) multimodality, accounting for the multimodal nature of many exam questions to test the model's multimodal understanding capability; and (3) multilevel structure, featuring exams from three critical educational periods to comprehensively assess a model's proficiency at different levels. In total, M3Exam contains 12,317 questions in 9 diverse languages with three educational levels, where about 23\% of the questions require processing images for successful solving. We assess the performance of top-performing LLMs on M3Exam and find that current models, including GPT-4, still struggle with multilingual text, particularly in low-resource and non-Latin script languages. Multimodal LLMs also perform poorly with complex multimodal questions. We believe that M3Exam can be a valuable resource for comprehensively evaluating LLMs by examining their multilingual and multimodal abilities and tracking their development. Data and evaluation code is available at \url{https://github.com/DAMO-NLP-SG/M3Exam}.
Wenxuan Zhang 0001, Mahani Aljunied, Yew Ken Chia, Lidong Bing
NeurIPS5
2023 A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges
abstract
As an important fine-grained sentiment analysis problem, aspect-based sentiment analysis (ABSA), aiming to analyze and understand people's opinions at the aspect level, has been attracting considerable interest in the last decade. To handle ABSA in different scenarios, various tasks are introduced for analyzing different sentiment elements and their relations, including the aspect term, aspect category, opinion term, and sentiment polarity. Unlike early ABSA works focusing on a single sentiment element, many compound ABSA tasks involving multiple elements have been studied in recent years for capturing more complete aspect-level sentiment information. However, a systematic review of various ABSA tasks and their corresponding solutions is still lacking, which we aim to fill in this survey. More specifically, we provide a new taxonomy for ABSA which organizes existing studies from the axes of concerned sentiment elements, with an emphasis on recent advances of compound ABSA tasks. From the perspective of solutions, we summarize the utilization of pre-trained language models for ABSA, which improved the performance of ABSA to a new stage. Besides, techniques for building more practical ABSA systems in cross-domain/lingual scenarios are discussed. Finally, we review some emerging topics and discuss some open challenges to outlook potential future directions of ABSA.
Wenxuan Zhang 0001, Xin Li 0056, Yang Deng 0002, Lidong Bing, Wai Lam
IEEE Trans. Knowl. Data Eng.4
2022 IAM: A Comprehensive and Large-Scale Dataset for Integrated Argument Mining Tasks
abstract
Traditionally, a debate usually requires a manual preparation process, including reading plenty of articles, selecting the claims, identifying the stances of the claims, seeking the evidence for the claims, etc.As the AI debate attracts more attention these years, it is worth exploring the methods to automate the tedious process involved in the debating system.In this work, we introduce a comprehensive and large dataset named IAM, which can be applied to a series of argument mining tasks, including claim extraction, stance classification, evidence extraction, etc.Our dataset is collected from over 1k articles related to 123 topics.Near 70k sentences in the dataset are fully annotated based on their argument properties (e.g., claims, stances, evidence, etc.).We further propose two new integrated argument mining tasks associated with the debate preparation process: (1) claim extraction with stance classification (CESC) and (2) claim-evidence pair extraction (CEPE).We adopt a pipeline approach and an end-to-end method for each integrated task separately.Promising experimental results are reported to show the values and challenges of our proposed tasks, and motivate future research on argument mining. 1
Liying Cheng, Lidong Bing, Ruidan He, Yan Zhang 0004, Luo Si
ACL (1)2
2022 GlobalWoZ: Globalizing MultiWoZ to Develop Multilingual Task-Oriented Dialogue Systems
abstract
Bosheng Ding, Junjie Hu, Lidong Bing, Mahani Aljunied, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Bosheng Ding, Junjie Hu 0001, Lidong Bing, Sharifah Mahani Aljunied, Shafiq R. Joty, Luo Si, Chunyan Miao
ACL (1)3
2022 MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER
abstract
Data augmentation is an effective solution to data scarcity in low-resource scenarios.However, when applied to token-level tasks such as NER, data augmentation methods often suffer from token-label misalignment, which leads to unsatsifactory performance.In this work, we propose Masked Entity Language Modeling (MELM) as a novel data augmentation framework for low-resource NER.To alleviate the token-label misalignment issue, we explicitly inject NER labels into sentence context, and thus the fine-tuned MELM is able to predict masked entity tokens by explicitly conditioning on their labels.Thereby, MELM generates high-quality augmented data with novel entities, which provides rich entity regularity knowledge and boosts NER performance.When training data from multiple languages are available, we also integrate MELM with codemixing for further improvement.We demonstrate the effectiveness of MELM on monolingual, cross-lingual and multilingual NER across various low-resource levels.Experimental results show that our MELM presents substantial improvement over the baseline methods. 1
Ran Zhou 0004, Xin Li 0056, Ruidan He, Lidong Bing, Erik Cambria, Luo Si, Chunyan Miao
ACL (1)4
2022 SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning
abstract
With the boom of e-commerce, Multimodal Review Helpfulness Prediction (MRHP) that identifies the helpfulness score of multimodal product reviews has become a research hotspot. Previous work on this task focuses on attention-based modality fusion, information integration, and relation modeling, which primarily exposes the following drawbacks: 1) the model may fail to capture the really essential information due to its indiscriminate attention formulation; 2) lack appropriate modeling methods that takes full advantage of correlation among provided data. In this paper, we propose SANCL: Selective Attention and Natural Contrastive Learning for MRHP. SANCL adopts a probe-based strategy to enforce high attention weights on the regions of greater significance. It also constructs a contrastive learning framework based on natural matching properties in the dataset. Experimental results on two benchmark datasets with three categories show that SANCL achieves state-of-the-art baseline performance with lower memory consumption.
Wei Han 0002, Hui Chen 0023, Zhen Hai, Soujanya Poria, Lidong Bing
COLING5
2022 Towards Multi-Sense Cross-Lingual Alignment of Contextual Embeddings
abstract
Cross-lingual word embeddings (CLWE) have been proven useful in many cross-lingual tasks. However, most existing approaches to learn CLWE including the ones with contextual embeddings are sense agnostic. In this work, we propose a novel framework to align contextual embeddings at the sense level by leveraging cross-lingual signal from bilingual dictionaries only. We operationalize our framework by first proposing a novel sense-aware cross entropy loss to model word senses explicitly. The monolingual ELMo and BERT models pretrained with our sense-aware cross entropy loss demonstrate significant performance improvement for word sense disambiguation tasks. We then propose a sense alignment objective on top of the sense-aware cross entropy loss for cross-lingual model pretraining, and pretrain cross-lingual models for several language pairs (English to German/Spanish/Japanese/Chinese). Compared with the best baseline results, our cross-lingual models achieve 0.52%, 2.09% and 1.29% average performance improvements on zero-shot cross-lingual NER, sentiment classification and XNLI tasks, respectively.
Thien Hai Nguyen, Shafiq R. Joty, Lidong Bing, Luo Si
COLING4
2022 Domain Generalization for Text Classification with Memory-Based Supervised Contrastive Learning
abstract
While there is much research on cross-domain text classification, most existing approaches focus on one-to-one or many-to-one domain adaptation. In this paper, we tackle the more challenging task of domain generalization, in which domain-invariant representations are learned from multiple source domains, without access to any data from the target domains, and classification decisions are then made on test documents in unseen target domains. We propose a novel framework based on supervised contrastive learning with a memory-saving queue. In this way, we explicitly encourage examples of the same class to be closer and examples of different classes to be further apart in the embedding space. We have conducted extensive experiments on two Amazon review sentiment datasets, and one rumour detection dataset. Experimental results show that our domain generalization method consistently outperforms state-of-the-art domain adaptation methods.
Ruidan He, Lidong Bing, Hwee Tou Ng
COLING3
2022 Retrofitting Multilingual Sentence Embeddings with Abstract Meaning Representation
abstract
We introduce a new method to improve existing multilingual sentence embeddings with Abstract Meaning Representation (AMR).Compared with the original textual input, AMR is a structured semantic representation that presents the core concepts and relations in a sentence explicitly and unambiguously.It also helps reduce surface variations across different expressions and languages.Unlike most prior work that only evaluates the ability to measure semantic similarity, we present a thorough evaluation of existing multilingual sentence embeddings and our improved versions, which include a collection of five transfer tasks in different downstream applications.Experiment results show that retrofitting multilingual sentence embeddings with AMR leads to better state-of-the-art performance on both semantic textual similarity and transfer tasks.Our codebase and evaluation scripts
Deng Cai 0002, Xin Li 0056, Jackie C. S. Ho, Lidong Bing, Wai Lam
EMNLP4
2022 A Dataset for Hyper-Relational Extraction and a Cube-Filling Approach
abstract
Relation extraction has the potential for largescale knowledge graph construction, but current methods do not consider the qualifier attributes for each relation triplet, such as time, quantity or location.The qualifiers form hyperrelational facts which better capture the rich and complex knowledge graph structure.For example, the relation triplet (Leonard Parker, Educated At, Harvard University) can be factually enriched by including the qualifier (End Time, 1967).Hence, we propose the task of hyper-relational extraction to extract more specific and complete facts from text.To support the task, we construct HyperRED, a large-scale and general-purpose dataset.Existing models cannot perform hyper-relational extraction as it requires a model to consider the interaction between three entities.Hence, we propose Cu-beRE, a cube-filling model inspired by tablefilling approaches and explicitly considers the interaction between relation triplets and qualifiers.To improve model scalability and reduce negative class imbalance, we further propose a cube-pruning method.Our experiments show that CubeRE outperforms strong baselines and reveal possible directions for future research.Our code and data are available at github.com/declare-lab/HyperRED.
Yew Ken Chia, Lidong Bing, Sharifah Mahani Aljunied, Luo Si, Soujanya Poria
EMNLP2
2022 Enhancing Multilingual Language Model with Massive Multilingual Knowledge Triples
abstract
Knowledge-enhanced language representation learning has shown promising results across various knowledge-intensive NLP tasks.However, prior methods are limited in efficient utilization of multilingual knowledge graph (KG) data for language model (LM) pretraining.They often train LMs with KGs in indirect ways, relying on extra entity/relation embeddings to facilitate knowledge injection.In this work, we explore methods to make better use of the multilingual annotation and language agnostic property of KG triples, and present novel knowledge based multilingual language models (KMLMs) trained directly on the knowledge triples.We first generate a large amount of multilingual synthetic sentences using the Wikidata KG triples.Then based on the intra-and inter-sentence structures of the generated data, we design pretraining tasks to enable the LMs to not only memorize the factual knowledge but also learn useful logical patterns.Our pretrained KMLMs demonstrate significant performance improvements on a wide range of knowledge-intensive crosslingual tasks, including named entity recognition (NER), factual knowledge retrieval, relation classification, and a newly designed logical reasoning task. 1
Xin Li 0056, Ruidan He, Lidong Bing, Shafiq R. Joty, Luo Si
EMNLP4
2022 Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Prediction
abstract
Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images.Unfortunately, those contemporary approaches pay scarce attention to polish representations of cross-modal relations and tend to suffer from inferior optimization.This might cause harm to model's predictions in numerous cases.To overcome the aforementioned issues, we propose Multimodal Contrastive Learning for Multimodal Review Helpfulness Prediction (MRHP) problem, concentrating on mutual information between input modalities to explicitly elaborate cross-modal relations.In addition, we introduce Adaptive Weighting scheme for our contrastive learning approach in order to increase flexibility in optimization.Lastly, we propose Multimodal Interaction module to address the unalignment nature of multimodal data, thereby assisting the model in producing more reasonable multimodal representations.Experimental results show that our method outperforms prior baselines and achieves state-of-the-art results on two publicly available benchmark datasets for MRHP problem.
Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, Lidong Bing
EMNLP5
2022 SentBS: Sentence-level Beam Search for Controllable Summarization
abstract
A wide range of control perspectives have been explored in controllable text generation.Structure-controlled summarization is recently proposed as a useful and interesting research direction.However, current structure-controlling methods have limited effectiveness in enforcing the desired structure.To address this limitation, we propose a sentence-level beam search generation method (SentBS), where evaluation is conducted throughout the generation process to select suitable sentences for subsequent generations.We experiment with different combinations of decoding methods to be used as subcomponents by SentBS and evaluate results on the structure-controlled dataset MReD.Experiments show that all explored combinations for SentBS can improve the agreement between the generated text and the desired structure, with the best method significantly reducing the structural discrepancies suffered by the existing model, by approximately 68%. 1
Chenhui Shen, Liying Cheng, Lidong Bing, Luo Si
EMNLP3
2022 Revisiting DocRED - Addressing the False Negative Problem in Relation Extraction
abstract
The DocRED dataset is one of the most popular and widely used benchmarks for documentlevel relation extraction (RE).It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset.However, we find that the annotation of DocRED is incomplete, i.e., false negative samples are prevalent.We analyze the causes and effects of the overwhelming false negative problem in the DocRED dataset.To address the shortcoming, we re-annotate 4,053 documents in the DocRED dataset by adding the missed relation triples back to the original DocRED.We name our revised DocRED dataset Re-DocRED.We conduct extensive experiments with state-ofthe-art neural models on both datasets, and the experimental results show that the models trained and evaluated on our Re-DocRED achieve performance improvements of around 13 F1 points.Moreover, we conduct a comprehensive analysis to identify the potential areas for further improvement.1
Lu Xu 0007, Lidong Bing, Hwee Tou Ng, Sharifah Mahani Aljunied
EMNLP3
2022 Interventional Training for Out-Of-Distribution Natural Language Understanding
abstract
Out-of-distribution (OOD) settings are used to measure a model's performance when the distribution of the test data is different from that of the training data.NLU models are known to suffer in OOD settings (Utama et al., 2020b).We study this issue from the perspective of causality, which sees confounding bias as the reason for models to learn spurious correlations.While a common solution is to perform intervention, existing methods handle only known and single confounder (Pearl and Mackenzie, 2018), but in many NLU tasks the confounders can be both unknown and multifactorial.In this paper, we propose a novel interventional training method called Bottom-up Automatic Intervention (BAI) that performs multi-granular intervention with identified multifactorial confounders.Our experiments on three NLU tasks, namely, natural language inference, fact verification and paraphrase identification, show the effectiveness of BAI for tackling different OOD settings.1
Sicheng Yu, Jing Jiang 0001, Hao Zhang 0048, Yulei Niu, Qianru Sun, Lidong Bing
EMNLP6
2022 ConNER: Consistency Training for Cross-lingual Named Entity Recognition
abstract
Cross-lingual named entity recognition (NER) suffers from data scarcity in the target languages, especially under zero-shot settings.Existing translate-train or knowledge distillation methods attempt to bridge the language gap, but often introduce a high level of noise.To solve this problem, consistency training methods regularize the model to be robust towards perturbations on data or hidden states.However, such methods are likely to violate the consistency hypothesis, or mainly focus on coarse-grain consistency.We propose ConNER as a novel consistency training framework for cross-lingual NER, which comprises of: (1) translation-based consistency training on unlabeled target-language data, and (2) dropoutbased consistency training on labeled sourcelanguage data.ConNER effectively leverages unlabeled target-language data and alleviates overfitting on the source language to enhance the cross-lingual adaptability.Experimental results show our ConNER achieves consistent improvement over various baseline methods. 1
Ran Zhou 0004, Xin Li 0056, Lidong Bing, Erik Cambria, Luo Si, Chunyan Miao
EMNLP3
2021 Exploring Auxiliary Reasoning Tasks for Task-oriented Dialog Systems with Meta Cooperative Learning
abstract
In this paper, we propose a Meta Cooperative Learning (MCL) framework for task-oriented dialog systems (TDSs). Our model consists of an auxiliary KB reasoning task for learning meta KB knowledge, an auxiliary dialogue reasoning task for learning dialogue patterns, and a TDS task (primary task) that aims at not only retrieving accurate entities from KB but also generating natural responses, which are coordinated to achieve collective success in both retrieving accurate KB entities and generating human-like responses via meta learning. Concretely, the dialog generation model amalgamates complementary meta KB and dialog knowledge from two novel auxiliary reasoning tasks that together provide integrated guidance to build a high-quality TDS by adding regularization terms to force primary network to produce similar results to auxiliary networks. While MCL automatically learns appropriate labels for the two auxiliary reasoning tasks from the primary task, without requiring access to any further data. The key idea behind MCL is to use the performance of the primary task, which is trained alongside the auxiliary tasks in one iteration, to improve the auxiliary labels for the next iteration with meta learning. Experimental results on three benchmark datasets show that MCL can generate higher quality responses compared to several strong baselines in terms of both automatic and human evaluations. Code to reproduce the results in this paper is available at: https://github.com/siat-nlp/MCL.
Bowen Qin, Min Yang 0007, Lidong Bing, Qingshan Jiang, Chengming Li 0004, Ruifeng Xu 0001
AAAI3
2021 Bootstrapped Unsupervised Sentence Representation Learning
abstract
Yan Zhang, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yan Zhang 0004, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li 0001
ACL/IJCNLP (1)4
2021 Argument Pair Extraction via Attention-guided Multi-Layer Multi-Cross Encoding
abstract
Liying Cheng, Tianyu Wu, Lidong Bing, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Liying Cheng, Lidong Bing, Luo Si
ACL/IJCNLP (1)3
2021 On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation
abstract
Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jiawei Low, Lidong Bing, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruidan He, Hai Ye, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, Luo Si
ACL/IJCNLP (1)8
2021 MulDA: A Multilingual Data Augmentation Framework for Low-Resource Cross-Lingual NER
abstract
Linlin Liu, Bosheng Ding, Lidong Bing, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bosheng Ding, Lidong Bing, Shafiq R. Joty, Luo Si, Chunyan Miao
ACL/IJCNLP (1)3
2021 Multi-perspective Coherent Reasoning for Helpfulness Prediction of Multimodal Reviews
abstract
Junhao Liu, Zhen Hai, Min Yang, Lidong Bing. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Junhao Liu 0001, Zhen Hai, Min Yang 0007, Lidong Bing
ACL/IJCNLP (1)4
2021 Learning Span-Level Interactions for Aspect Sentiment Triplet Extraction
abstract
Lu Xu, Yew Ken Chia, Lidong Bing. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lu Xu 0007, Yew Ken Chia, Lidong Bing
ACL/IJCNLP (1)3
2021 Aspect Sentiment Quad Prediction as Paraphrase Generation
abstract
Aspect-based sentiment analysis (ABSA) has been extensively studied in recent years, which typically involves four fundamental sentiment elements, including the aspect category, aspect term, opinion term, and sentiment polarity.Existing studies usually consider the detection of partial sentiment elements, instead of predicting the four elements in one shot.In this work, we introduce the Aspect Sentiment Quad Prediction (ASQP) task, aiming to jointly detect all sentiment elements in quads for a given opinionated sentence, which can reveal a more comprehensive and complete aspect-level sentiment structure.We further propose a novel PARAPHRASE modeling paradigm to cast the ASQP task to a paraphrase generation process.On one hand, the generation formulation allows solving ASQP in an end-to-end manner, alleviating the potential error propagation in the pipeline solution.On the other hand, the semantics of the sentiment elements can be fully exploited by learning to generate them in the natural language form.Extensive experiments on benchmark datasets show the superiority of our proposed method and the capacity of crosstask transfer with the proposed unified PARA-PHRASE modeling framework.
Wenxuan Zhang 0001, Yang Deng 0002, Xin Li 0056, Yifei Yuan 0002, Lidong Bing, Wai Lam
EMNLP (1)5
2021 Cross-lingual Aspect-based Sentiment Analysis with Aspect Term Code-Switching
abstract
Many efforts have been made in solving the Aspect-based sentiment analysis (ABSA) task.While most existing studies focus on English texts, handling ABSA in resource-poor languages remains a challenging problem.In this paper, we consider the unsupervised crosslingual transfer for the ABSA task, where only labeled data in the source language is available and we aim at transferring its knowledge to the target language having no labeled data.To this end, we propose an alignment-free label projection method to obtain high-quality pseudolabeled data of the target language with the help of the translation system, which could preserve more accurate task-specific knowledge in the target language.For better utilizing the source and translated data, as well as enhancing the cross-lingual alignment, we design an aspect code-switching mechanism to augment the training data with code-switched bilingual sentences.To further investigate the importance of language-specific knowledge in solving the ABSA problem, we distill the above model on the unlabeled target language data which improves the performance to the same level of the supervised method.
Wenxuan Zhang 0001, Ruidan He, Haiyun Peng, Lidong Bing, Wai Lam
EMNLP (1)4
2021 Better Feature Integration for Named Entity Recognition
abstract
It has been shown that named entity recognition (NER) could benefit from incorporating the long-distance structured information captured by dependency trees.We believe this is because both types of features -the contextual information captured by the linear sequences and the structured information captured by the dependency trees may complement each other.However, existing approaches largely focused on stacking the LSTM and graph neural networks such as graph convolutional networks (GCNs) for building improved NER models, where the exact interaction mechanism between the two different types of features is not very clear, and the performance gain does not appear to be significant.In this work, we propose a simple and robust solution to incorporate both types of features with our Synergized-LSTM (Syn-LSTM), which clearly captures how the two types of features interact.We conduct extensive experiments on several standard datasets across four languages.The results demonstrate that the proposed model achieves better performance than previous approaches while requiring fewer parameters.Our further analysis demonstrates that our model can capture longer dependencies compared with strong baselines.1
Lu Xu 0007, Zhanming Jie, Wei Lu 0011, Lidong Bing
NAACL-HLT4
2021 Overview of Argumentative Text Understanding for AI Debater Challenge
Liying Cheng, Ruidan He, Yinzi Li, Lidong Bing, Zhongyu Wei, Qin Liu 0010, Chenhui Shen, Shuonan Zhang, Changlong Sun, Luo Si, Changjian Jiang, Xuanjing Huang 0001
NLPCC (2)5
2021 Harvest shopping advice: Neural Question Generation from multiple information sources in E-commerce
Yongzhen Wang 0002, Kaisong Song, Lidong Bing, Xiaozhong Liu 0001
Neurocomputing3
2020 Open Domain Event Text Generation
abstract
Text generation tasks aim at generating human-readable text from different kinds of data. Normally, the generated text only contains the information included in the data and its application is thus restricted to some limited scenarios. In this paper, we extend the task to an open domain event text generation scenario with an entity chain as its skeleton. Specifically, given an entity chain containing several related event entities, the model should retrieve from a trustworthy repository (e.g. Wikipedia) the detailed information of these entities and generate a description text based on the retrieved sentences. We build a new dataset called WikiEvent1 that provides 34K pairs of entity chain and its corresponding description sentences. To solve the problem, we propose a wiki augmented generator framework that contains an encoder, a retriever, and a decoder. The encoder encodes the entity chain into a hidden space while the decoder decodes from the hidden space and generates description text. The retriever retrieves relevant text from a trustworthy repository which provides more information for generation. To alleviate the overfitting problem, we propose a novel random drop component that randomly deletes words from the retrieved sentences making our model more robust for handling long input sentences. We apply the proposed model on the WikiEvent dataset and compare it with a few baselines. The experimental results show that our carefully-designed architecture does help generate better event text, and extensive analysis further uncovers the characteristics of the proposed task.
Lidong Bing, Wai Lam
AAAI2
2020 Cross-Lingual Low-Resource Set-to-Description Retrieval for Global E-Commerce
abstract
With the prosperous of cross-border e-commerce, there is an urgent demand for designing intelligent approaches for assisting e-commerce sellers to offer local products for consumers from all over the world. In this paper, we explore a new task of cross-lingual information retrieval, i.e., cross-lingual set-to-description retrieval in cross-border e-commerce, which involves matching product attribute sets in the source language with persuasive product descriptions in the target language. We manually collect a new and high-quality paired dataset, where each pair contains an unordered product attribute set in the source language and an informative product description in the target language. As the dataset construction process is both time-consuming and costly, the new dataset only comprises of 13.5k pairs, which is a low-resource setting and can be viewed as a challenging testbed for model development and evaluation in cross-border e-commerce. To tackle this cross-lingual set-to-description retrieval task, we propose a novel cross-lingual matching network (CLMN) with the enhancement of context-dependent cross-lingual mapping upon the pre-trained monolingual BERT representations. Experimental results indicate that our proposed CLMN yields impressive results on the challenging task and the context-dependent cross-lingual mapping on BERT yields noticeable improvement over the pre-trained multi-lingual BERT model.
Juntao Li 0005, Chang Liu 0076, Lidong Bing, Hongsong Li, Xiaozhong Liu 0001, Dongyan Zhao 0001, Rui Yan 0001
AAAI4
2020 Knowing What, How and Why: A Near Complete Solution for Aspect-Based Sentiment Analysis
abstract
Target-based sentiment analysis or aspect-based sentiment analysis (ABSA) refers to addressing various sentiment analysis tasks at a fine-grained level, which includes but is not limited to aspect extraction, aspect sentiment classification, and opinion extraction. There exist many solvers of the above individual subtasks or a combination of two subtasks, and they can work together to tell a complete story, i.e. the discussed aspect, the sentiment on it, and the cause of the sentiment. However, no previous ABSA research tried to provide a complete solution in one shot. In this paper, we introduce a new subtask under ABSA, named aspect sentiment triplet extraction (ASTE). Particularly, a solver of this task needs to extract triplets (What, How, Why) from the inputs, which show WHAT the targeted aspects are, HOW their sentiment polarities are and WHY they have such polarities (i.e. opinion reasons). For instance, one triplet from “Waiters are very friendly and the pasta is simply average” could be (‘Waiters’, positive, ‘friendly’). We propose a two-stage framework to address this task. The first stage predicts what, how and why in a unified model, and then the second stage pairs up the predicted what (how) and why from the first stage to output triplets. In the experiments, our framework has set a benchmark performance in this novel triplet extraction task. Meanwhile, it outperforms a few strong baselines adapted from state-of-the-art related methods.
Haiyun Peng, Lu Xu 0007, Lidong Bing, Fei Huang 0002, Wei Lu 0011, Luo Si
AAAI3
2020 GRET: Global Representation Enhanced Transformer
abstract
Transformer, based on the encoder-decoder framework, has achieved state-of-the-art performance on several natural language generation tasks. The encoder maps the words in the input sentence into a sequence of hidden states, which are then fed into the decoder to generate the output sentence. These hidden states usually correspond to the input words and focus on capturing local information. However, the global (sentence level) information is seldom explored, leaving room for the improvement of generation quality. In this paper, we propose a novel global representation enhanced Transformer (GRET) to explicitly model global representation in the Transformer network. Specifically, in the proposed model, an external state is generated for the global representation from the encoder. The global representation is then fused into the decoder during the decoding process to improve generation quality. We conduct experiments in two text generation tasks: machine translation and text summarization. Experimental results on four WMT machine translation tasks and LCSTS text summarization task demonstrate the effectiveness of the proposed approach on natural language generation1.
Rongxiang Weng, Shujian Huang, Heng Yu 0006, Lidong Bing, Weihua Luo, Jiajun Chen 0001
AAAI5
2020 Improving Low-Resource Named Entity Recognition using Joint Sentence and Token Labeling
abstract
Exploiting sentence-level labels, which are easy to obtain, is one of the plausible methods to improve low-resource named entity recognition (NER), where token-level labels are costly to annotate.Current models for jointly learning sentence and token labeling are limited to binary classification.We present a joint model that supports multi-class classification and introduce a simple variant of self-attention that allows the model to learn scaling factors.Our model produces 3.78%, 4.20%, 2.08% improvements in F1 over the BiLSTM-CRF baseline on e-commerce product titles in three different low-resource languages: Vietnamese, Thai, and Indonesian, respectively.
Canasai Kruengkrai, Thien Hai Nguyen, Sharifah Mahani Aljunied, Lidong Bing
ACL4
2020 Review-based Question Generation with Adaptive Instance Transfer and Augmentation
abstract
While online reviews of products and services become an important information source, it remains inefficient for potential consumers to exploit verbose reviews for fulfilling their information need.We propose to explore question generation as a new way of review information exploitation, namely generating questions that can be answered by the corresponding review sentences.One major challenge of this generation task is the lack of training data, i.e. explicit mapping relation between the user-posed questions and review sentences.To obtain proper training instances for the generation model, we propose an iterative learning framework with adaptive instance transfer and augmentation.To generate to the point questions about the major aspects in reviews, related features extracted in an unsupervised manner are incorporated without the burden of aspect annotation.Experiments on data from various categories of a popular E-commerce site demonstrate the effectiveness of the framework, as well as the potentials of the proposed review-based question generation task.
Lidong Bing, Wai Lam, Luo Si
ACL2
2020 Dynamic Topic Tracker for KB-to-Text Generation
abstract
Recently, many KB-to-text generation tasks have been proposed to bridge the gap between knowledge bases and natural language by directly converting a group of knowledge base triples into human-readable sentences.However, most of the existing models suffer from the off-topic problem, namely, the models are prone to generate some unrelated clauses that are somehow involved with certain input terms regardless of the given input data.This problem seriously degrades the quality of the generation results.In this paper, we propose a novel dynamic topic tracker for solving this problem.Different from existing models, our proposed model learns a global hidden representation for topics and recognizes the corresponding topic during each generation step.The recognized topic is used as additional information to guide the generation process and thus alleviates the off-topic problem.The experimental results show that our proposed model can enhance the performance of sentence generation and the off-topic problem is significantly mitigated.
Lidong Bing, Wai Lam, Shoaib Jameel
COLING2
2020 APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task Learning
abstract
Peer review and rebuttal, with rich interactions and argumentative discussions in between, are naturally a good resource to mine arguments.However, few works study both of them simultaneously.In this paper, we introduce a new argument pair extraction (APE) task on peer review and rebuttal in order to study the contents, the structure and the connections between them.We prepare a challenging dataset that contains 4,764 fully annotated review-rebuttal passage pairs from an open review platform to facilitate the study of this task.To automatically detect argumentative propositions and extract argument pairs from this corpus, we cast it as the combination of a sequence labeling task and a text relation classification task.Thus, we propose a multitask learning framework based on hierarchical LSTM networks.Extensive experiments and analysis demonstrate the effectiveness of our multi-task framework, and also show the challenges of the new task as well as motivate future research directions. 1
Liying Cheng, Lidong Bing, Wei Lu 0011, Luo Si
EMNLP (1)2
2020 ENT-DESC: Entity Description Generation by Exploring Knowledge Graph
abstract
Previous works on knowledge-to-text generation take as input a few RDF triples or keyvalue pairs conveying the knowledge of some entities to generate a natural language description.Existing datasets, such as WIKIBIO, WebNLG, and E2E, basically have a good alignment between an input triple/pair set and its output text.However, in practice, the input knowledge could be more than enough, since the output description may only cover the most significant knowledge.In this paper, we introduce a large-scale and challenging dataset to facilitate the study of such a practical scenario in KG-to-text.Our dataset involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter explosion while generating the descriptions.We address these challenges by proposing a multi-graph structure that is able to represent the original graph information more comprehensively.Furthermore, we also incorporate aggregation methods that learn to extract the rich graph information.Extensive experiments demonstrate the effectiveness of our model architecture.1
Liying Cheng, Dekun Wu, Lidong Bing, Yan Zhang 0004, Zhanming Jie, Wei Lu 0011, Luo Si
EMNLP (1)3
2020 DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging Tasks
abstract
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, Chunyan Miao. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Bosheng Ding, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq R. Joty, Luo Si, Chunyan Miao
EMNLP (1)3
2020 Partially-Aligned Data-to-Text Generation with Distant Supervision
abstract
The Data-to-Text task aims to generate humanreadable text for describing some given structured data enabling more interpretability.However, the typical generation task is confined to a few particular domains since it requires wellaligned data which is difficult and expensive to obtain.Using partially-aligned data is an alternative way of solving the dataset scarcity problem.This kind of data is much easier to obtain since it can be produced automatically.However, using this kind of data induces the over-generation problem posing difficulties for existing models, which tends to add unrelated excerpts during the generation procedure.In order to effectively utilize automatically annotated partially-aligned datasets, we extend the traditional generation task to a refined task called Partially-Aligned Data-to-Text Generation (PADTG) which is more practical since it utilizes automatically annotated data for training and thus considerably expands the application domains.To tackle this new task, we propose a novel distant supervision generation framework.It firstly estimates the input data's supportiveness for each target word with an estimator and then applies a supportiveness adaptor and a rebalanced beam search to harness the over-generation problem in the training and generation phases respectively.We also contribute a partially-aligned dataset 1 by sampling sentences from Wikipedia and automatically extracting corresponding KB triples for each sentence from Wikidata.The experimental results show that our framework outperforms all baseline models as well as verify the feasibility of utilizing partially-aligned data. 1 The data and source code of this paper can be obtained from https://github.com/fuzihaofzh/ distant_supervision_nlg Company of Heroes is a strategy video game developed in Canada. Company of Heroes is a strategy video game developed in Canada.Company of Heroes, genre, strategy video game Company of Heroes is a strategy video game developed in Canada.
Bei Shi, Wai Lam, Lidong Bing, Zhiyuan Liu 0001
EMNLP (1)4
2020 Aspect Sentiment Classification with Aspect-Specific Opinion Spans
abstract
Aspect sentiment classification, predicting the sentiment polarity of given aspects, has drawn extensive attention.Previous attention-based models emphasize using aspect semantics to help extract opinion features for classification.However, these works are either not able to capture opinion spans as a whole or capture variable-length opinion spans.In this paper, we present a neat and effective multiple CRFs based structured attention model that is capable of extracting aspect-specific opinion spans.The sentiment polarity of the target is then classified based on the extracted opinion features and contextual information.The experimental results on four datasets demonstrate the effectiveness of the proposed model, and our further analysis shows that our model can capture aspect-specific opinion spans. 1
Lu Xu 0007, Lidong Bing, Wei Lu 0011, Fei Huang 0002
EMNLP (1)2
2020 Position-Aware Tagging for Aspect Sentiment Triplet Extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE)is the task of extracting the triplets of target entities, their associated sentiment, and opinion spans explaining the reason for the sentiment.Existing research efforts mostly solve this problem using pipeline approaches, which break the triplet extraction process into several stages.Our observation is that the three elements within a triplet are highly related to each other, and this motivates us to build a joint model to extract such triplets using a sequence tagging approach.However, how to effectively design a tagging approach to extract the triplets that can capture the rich interactions among the elements is a challenging research question.In this work, we propose the first end-to-end model with a novel positionaware tagging scheme that is capable of jointly extracting the triplets.Our experimental results on several existing datasets show that jointly capturing elements in the triplet using our approach leads to improved performance over the existing approaches.We also conducted extensive experiments to investigate the model effectiveness and robustness 1 .
Lu Xu 0007, Wei Lu 0011, Lidong Bing
EMNLP (1)4
2020 Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-Training
abstract
Adapting pre-trained language models (PrLMs) (e.g., BERT) to new domains has gained much attention recently.Instead of fine-tuning PrLMs as done in most previous work, we investigate how to adapt the features of PrLMs to new domains without fine-tuning.We explore unsupervised domain adaptation (UDA) in this paper.With the features from PrLMs, we adapt the models trained with labeled data from the source domain to the unlabeled target domain.Self-training is widely used for UDA, and it predicts pseudo labels on the target domain data for training.However, the predicted pseudo labels inevitably include noise, which will negatively affect training a robust model.To improve the robustness of self-training, in this paper we present class-aware feature self-distillation (CFd) to learn discriminative features from PrLMs, in which PrLM features are self-distilled into a feature adaptation module and the features from the same class are more tightly clustered.We further extend CFd to a cross-language setting, in which language discrepancy is studied.Experiments on two monolingual and multilingual Amazon review datasets show that CFd can consistently improve the performance of self-training in cross-domain and cross-language settings.
Hai Ye, Ruidan He, Juntao Li 0005, Hwee Tou Ng, Lidong Bing
EMNLP (1)6
2020 Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
abstract
AMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters.
Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing
EMNLP (1)7
2020 An Unsupervised Sentence Embedding Method by Mutual Information Maximization
abstract
BERT is inefficient for sentence-pair tasks such as clustering or semantic search as it needs to evaluate combinatorially many sentence pairs which is very time-consuming.Sentence BERT (SBERT) attempted to solve this challenge by learning semantically meaningful representations of single sentences, such that similarity comparison can be easily accessed.However, SBERT is trained on corpus with high-quality labeled sentence pairs, which limits its application to tasks where labeled data is extremely scarce.In this paper, we propose a lightweight extension on top of BERT and a novel self-supervised learning objective based on mutual information maximization strategies to derive meaningful sentence embeddings in an unsupervised manner.Unlike SBERT, our method is not restricted by the availability of labeled data, such that it can be applied on different domain-specific corpus.Experimental results show that the proposed method significantly outperforms other unsupervised sentence embedding baselines on common semantic textual similarity (STS) tasks and downstream supervised tasks.It also outperforms SBERT in a setting where in-domain labeled data is not available, and achieves performance competitive with supervised methods on various tasks.
Yan Zhang 0004, Ruidan He, Zuozhu Liu, Kwan Hui Lim 0001, Lidong Bing
EMNLP (1)5
2020 Unsupervised Domain Adaptation of a Pretrained Cross-Lingual Language Model
abstract
Recent research indicates that pretraining cross-lingual language models on large-scale unlabeled texts yields significant performance improvements over various cross-lingual and low-resource tasks. Through training on one hundred languages and terabytes of texts, cross-lingual language models have proven to be effective in leveraging high-resource languages to enhance low-resource language processing and outperform monolingual models. In this paper, we further investigate the cross-lingual and cross-domain (CLCD) setting when a pretrained cross-lingual language model needs to adapt to new domains. Specifically, we propose a novel unsupervised feature decomposition method that can automatically extract domain-specific features and domain-invariant features from the entangled pretrained cross-lingual representations, given unlabeled raw texts in the source language. Our proposed model leverages mutual information estimation to decompose the representations computed by a cross-lingual model into domain-invariant and domain-specific parts. Experimental results show that our proposed method achieves significant performance improvements over the state-of-the-art pretrained cross-lingual language model in the CLCD setting.
Juntao Li 0005, Ruidan He, Hai Ye, Hwee Tou Ng, Lidong Bing, Rui Yan 0001
IJCAI5
2019 Generating Distractors for Reading Comprehension Questions from Real Examinations
abstract
We investigate the task of distractor generation for multiple choice reading comprehension questions from examinations. In contrast to all previous works, we do not aim at preparing words or short phrases distractors, instead, we endeavor to generate longer and semantic-rich distractors which are closer to distractors in real reading comprehension from examinations. Taking a reading comprehension article, a pair of question and its correct option as input, our goal is to generate several distractors which are somehow related to the answer, consistent with the semantic context of the question and have some trace in the article. We propose a hierarchical encoderdecoder framework with static and dynamic attention mechanisms to tackle this task. Specifically, the dynamic attention can combine sentence-level and word-level attention varying at each recurrent time step to generate a more readable sequence. The static attention is to modulate the dynamic attention not to focus on question irrelevant sentences or sentences which contribute to the correct option. Our proposed framework outperforms several strong baselines on the first prepared distractor generation dataset of real reading comprehension questions. For human evaluation, compared with those distractors generated by baselines, our generated distractors are more functional to confuse the annotators.
Yifan Gao 0001, Lidong Bing, Piji Li, Irwin King, Michael R. Lyu
AAAI2
2019 Abstractive Text Summarization by Incorporating Reader Comments
Shen Gao, Xiuying Chen, Piji Li, Zhaochun Ren, Lidong Bing, Dongyan Zhao 0001, Rui Yan 0001
AAAI5
2019 A Unified Model for Opinion Target Extraction and Target Sentiment Prediction
abstract
Target-based sentiment analysis involves opinion target extraction and target sentiment classification. However, most of the existing works usually studied one of these two sub-tasks alone, which hinders their practical use. This paper aims to solve the complete task of target-based sentiment analysis in an end-to-end fashion, and presents a novel unified model which applies a unified tagging scheme. Our framework involves two stacked recurrent neural networks: The upper one predicts the unified tags to produce the final output results of the primary target-based sentiment analysis; The lower one performs an auxiliary target boundary prediction aiming at guiding the upper network to improve the performance of the primary task. To explore the inter-task dependency, we propose to explicitly model the constrained transitions from target boundaries to target sentiment polarities. We also propose to maintain the sentiment consistency within an opinion target via a gate mechanism which models the relation between the features for the current word and the previous word. We conduct extensive experiments on three benchmark datasets and our framework achieves consistently superior results.
Xin Li 0056, Lidong Bing, Piji Li, Wai Lam
AAAI2
2019 Learning to Write Stories with Thematic Consistency and Wording Novelty
abstract
Automatic story generation is a challenging task, which involves automatically comprising a sequence of sentences or words with a consistent topic and novel wordings. Although many attention has been paid to this task and prompting progress has been made, there still exists a noticeable gap between generated stories and those created by humans, especially in terms of thematic consistency and wording novelty. To fill this gap, we propose a cache-augmented conditional variational autoencoder for story generation, where the cache module allows to improve thematic consistency while the conditional variational autoencoder part is used for generating stories with less common words by using a continuous latent variable. For combing the cache module and the autoencoder part, we further introduce an effective gate mechanism. Experimental results on ROCStories and WritingPrompts indicate that our proposed model can generate stories with consistency and wording novelty, and outperforms existing models under both automatic metrics and human evaluations.
Juntao Li 0005, Lidong Bing, Lisong Qiu, Dongmin Chen, Dongyan Zhao 0001, Rui Yan 0001
AAAI2
2019 A Knowledge Regularized Hierarchical Approach for Emotion Cause Analysis
abstract
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui, Lidong Bing, Min Yang, Ruifeng Xu, Ruibin Mao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui 0003, Lidong Bing, Min Yang 0007, Ruifeng Xu 0001, Ruibin Mao
EMNLP/IJCNLP (1)5
2019 Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations
abstract
Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao 0001, Rui Yan 0001
EMNLP/IJCNLP (1)5
2019 Improving Question Generation With to the Point Context
abstract
Jingjing Li, Yifan Gao, Lidong Bing, Irwin King, Michael R. Lyu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jingjing Li 0007, Yifan Gao 0001, Lidong Bing, Irwin King, Michael R. Lyu
EMNLP/IJCNLP (1)3
2019 Transferable End-to-End Aspect-based Sentiment Analysis with Selective Adversarial Learning
abstract
Zheng Li, Xin Li, Ying Wei, Lidong Bing, Yu Zhang, Qiang Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zheng Li 0018, Xin Li 0056, Ying Wei 0001, Lidong Bing, Yu Zhang 0006, Qiang Yang 0001
EMNLP/IJCNLP (1)4
2019 Hierarchical Pointer Net Parsing
abstract
Linlin Liu, Xiang Lin, Shafiq Joty, Simeng Han, Lidong Bing. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shafiq R. Joty, Simeng Han, Lidong Bing
EMNLP/IJCNLP (1)5
2019 Semi-supervised Text Style Transfer: Cross Projection in Latent Space
abstract
Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao, Shuming Shi, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao 0001, Shuming Shi 0001, Rui Yan 0001
EMNLP/IJCNLP (1)4
2019 Using Customer Service Dialogues for Satisfaction Analysis with Context-Assisted Multiple Instance Learning
abstract
Kaisong Song, Lidong Bing, Wei Gao, Jun Lin, Lujun Zhao, Jiancheng Wang, Changlong Sun, Xiaozhong Liu, Qiong Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Kaisong Song, Lidong Bing, Wei Gao 0001, Lujun Zhao, Changlong Sun, Xiaozhong Liu 0001, Qi Zhang 0001
EMNLP/IJCNLP (1)2
2019 Tackling Long-Tailed Relations and Uncommon Entities in Knowledge Graph Completion
abstract
Zihao Wang, Kwunping Lai, Piji Li, Lidong Bing, Wai Lam. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Kwun Ping Lai, Piji Li, Lidong Bing, Wai Lam
EMNLP/IJCNLP (1)4
2019 Difficulty Controllable Generation of Reading Comprehension Questions
abstract
We investigate the difficulty levels of questions in reading comprehension datasets such as SQuAD, and propose a new question generation setting, named Difficulty-controllable Question Generation (DQG). Taking as input a sentence in the reading comprehension paragraph and some of its text fragments (i.e., answers) that we want to ask questions about, a DQG method needs to generate questions each of which has a given text fragment as its answer, and meanwhile the generation is under the control of specified difficulty labels---the output questions should satisfy the specified difficulty as much as possible. To solve this task, we propose an end-to-end framework to generate questions of designated difficulty levels by exploring a few important intuitions. For evaluation, we prepared the first dataset of reading comprehension questions with difficulty labels. The results show that the question generated by our framework not only have better quality under the metrics like BLEU, but also comply with the specified difficulty labels.
Yifan Gao 0001, Lidong Bing, Wang Chen 0001, Michael R. Lyu, Irwin King
IJCAI2
2019 Persona-Aware Tips Generation?
abstract
Tips, as a compacted and concise form of reviews, were paid less attention by researchers. In this paper, we investigate the task of tips generation by considering the “persona” information which captures the intrinsic language style of the users or the different characteristics of the product items. In order to exploit the persona information, we propose a framework based on adversarial variational auto-encoders (aVAE) for persona modeling from the historical tips and reviews of users and items. The latent variables from aVAE are regarded as persona embeddings. Besides representing persona using the latent embeddings, we design a persona memory for storing the persona related words for users and items. Pointer Network is used to retrieve persona wordings from the memory when generating tips. Moreover, the persona embeddings are used as latent factors by a rating prediction component to predict the sentiment of a user over an item. Finally, the persona embeddings and the sentiment information are incorporated into a recurrent neural networks based tips generation component. Extensive experimental results are reported and discussed to elaborate the peculiarities of our framework.
Piji Li, Lidong Bing, Wai Lam
WWW3
2018 Transformation Networks for Target-Oriented Sentiment Classification
abstract
Target-oriented sentiment classification aims at classifying sentiment polarities over individual opinion targets in a sentence.RNN with attention seems a good fit for the characteristics of this task, and indeed it achieves the state-of-the-art performance.After re-examining the drawbacks of attention mechanism and the obstacles that block CNN to perform well in this classification task, we propose a new model to overcome these issues.Instead of attention, our model employs a CNN layer to extract salient features from the transformed word representations originated from a bi-directional RNN layer.Between the two layers, we propose a component to generate target-specific representations of words in the sentence, meanwhile incorporate a mechanism for preserving the original contextual information from the RNN layer.Experiments show that our model achieves a new state-of-the-art performance on a few benchmarks. 1
Xin Li 0056, Lidong Bing, Wai Lam, Bei Shi
ACL (1)2
2018 Learning Domain-Sensitive and Sentiment-Aware Word Embeddings
abstract
Word embeddings have been widely used in sentiment classification because of their efficacy for semantic representations of words.Given reviews from different domains, some existing methods for word embeddings exploit sentiment information, but they cannot produce domainsensitive embeddings.On the other hand, some other existing methods can generate domain-sensitive word embeddings, but they cannot distinguish words with similar contexts but opposite sentiment polarity.We propose a new method for learning domain-sensitive and sentimentaware embeddings that simultaneously capture the information of sentiment semantics and domain sensitivity of individual words.Our method can automatically determine and produce domain-common embeddings and domain-specific embeddings.The differentiation of domaincommon and domain-specific words enables the advantage of data augmentation of common semantics from multiple domains and capture the varied semantics of specific words from different domains at the same time.Experimental results show that our model provides an effective way to learn domain-sensitive and sentimentaware word embeddings which benefit sentiment classification at both sentence level and lexicon term level.
Bei Shi, Lidong Bing, Wai Lam
ACL (1)3
2018 Hybrid Neural Attention for Agreement/Disagreement Inference in Online Debates
abstract
Inferring the agreement/disagreement relation in debates, especially in online debates, is one of the fundamental tasks in argumentation mining. The expressions of agreement/disagreement usually rely on argumentative expressions in text as well as interactions between participants in debates. Previous works usually lack the capability of jointly modeling these two factors. To alleviate this problem, this paper proposes a hybrid neural attention model which combines self and cross attention mechanism to locate salient part from textual context and interaction between users. Experimental results on three (dis)agreement inference datasets show that our model outperforms the state-of-the-art models.
Jiachen Du, Lidong Bing, Ruifeng Xu 0001
EMNLP3
2018 Variational Autoregressive Decoder for Neural Response Generation
abstract
Combining the virtues of probability graphic models and neural networks, Conditional Variational Auto-encoder (CVAE) has shown promising performance in many applications such as response generation.However, existing CVAE-based models often generate responses from a single latent variable which may not be sufficient to model high variability in responses.To solve this problem, we propose a novel model that sequentially introduces a series of latent variables to condition the generation of each word in the response sequence.In addition, the approximate posteriors of these latent variables are augmented with a backward Recurrent Neural Network (RNN), which allows the latent variables to capture long-term dependencies of future tokens in generation.To facilitate training, we supplement our model with an auxiliary objective that predicts the subsequent bag of words.Empirical experiments conducted on the OpenSubtitle and Reddit datasets show that the proposed model leads to significant improvements on both relevance and diversity over state-of-the-art baselines.
Jiachen Du, Wenjie Li 0002, Yulan He 0001, Ruifeng Xu 0001, Lidong Bing, Xuan Wang 0002
EMNLP5
2018 QuaSE: Sequence Editing under Quantifiable Guidance
abstract
We propose the task of Quantifiable Sequence Editing (QuaSE): editing an input sequence to generate an output sequence that satisfies a given numerical outcome value measuring a certain property of the sequence, with the requirement of keeping the main content of the input sequence.For example, an input sequence could be a word sequence, such as review sentence and advertisement text.For a review sentence, the outcome could be the review rating; for an advertisement, the outcome could be the click-through rate.The major challenge in performing QuaSE is how to perceive the outcome-related wordings, and only edit them to change the outcome.In this paper, the proposed framework contains two latent factors, namely, outcome factor and content factor, disentangled from the input sentence to allow convenient editing to change the outcome and keep the content.Our framework explores the pseudo-parallel sentences by modeling their content similarity and outcome differences to enable a better disentanglement of the latent factors, which allows generating an output to better satisfy the desired outcome and keep the content.The dual reconstruction structure further enhances the capability of generating expected output by exploiting the couplings of latent factors of pseudo-parallel sentences.For evaluation, we prepared a dataset of Yelp review sentences with the ratings as outcome.Extensive experimental results are reported and discussed to elaborate the peculiarities of our framework.1
Lidong Bing, Piji Li, Shuming Shi 0001, Wai Lam, Tong Zhang 0001
EMNLP2
2018 Estimating Marginal Probabilities of n-grams for Recurrent Neural Language Models
abstract
Recurrent neural network language models (RNNLMs) are the current standard-bearer for statistical language modeling.However, RNNLMs only estimate probabilities for complete sequences of text, whereas some applications require context-independent phrase probabilities instead.In this paper, we study how to compute an RNNLM's marginal probability: the probability that the model assigns to a short sequence of text when the preceding context is not known.We introduce a simple method of altering the RNNLM training to make the model more accurate at marginal estimation.Our experiments demonstrate that the technique is effective compared to baselines including the traditional RNNLM probability and an importance sampling approach.Finally, we show how we can use the marginal estimation to improve an RNNLM by training the marginals to match n-gram probabilities from a larger corpus.
Thanapon Noraset, Doug Downey, Lidong Bing
EMNLP3
2018 Aspect Term Extraction with History Attention and Selective Transformation
abstract
Aspect Term Extraction (ATE), a key sub-task in Aspect-Based Sentiment Analysis, aims to extract explicit aspect expressions from online user reviews. We present a new framework for tackling ATE. It can exploit two useful clues, namely opinion summary and aspect detection history. Opinion summary is distilled from the whole input sentence, conditioned on each current token for aspect prediction, and thus the tailor-made summary can help aspect prediction on this token. On the other hand, the aspect detection history information is distilled from the previous aspect predictions, and it can leverage the coordinate structure and tagging schema constraints to upgrade the aspect prediction. Experimental results over four benchmark datasets clearly demonstrate that our framework can outperform all state-of-the-art methods.
Xin Li 0056, Lidong Bing, Piji Li, Wai Lam, Zhimou Yang
IJCAI2
2018 Semi-Supervised Learning with Declaratively Specified Entropy Constraints
abstract
We propose a technique for declaratively specifying strategies for semi-supervised learning (SSL). SSL methods based on different assumptions perform differently on different tasks, which leads to difficulties applying them in practice. In this paper, we propose to use entropy to unify many types of constraints. Our method can be used to easily specify ensembles of semi-supervised learners, as well as agreement constraints and entropic regularization constraints between these learners, and can be used to model both well-known heuristics such as co-training, and novel domain-specific heuristics. Besides, our model is flexible as to the underlying learning mechanism. Compared to prior frameworks for specifying SSL techniques, our technique achieves consistent improvements on a suite of well-studied SSL benchmarks, and obtains a new state-of-the-art result on a difficult relation extraction task.
Haitian Sun, William W. Cohen, Lidong Bing
NeurIPS3
2018 Learning a unified embedding space of web search from large-scale query log
Lidong Bing, Zhengyu Niu, Piji Li, Wai Lam, Haifeng Wang 0001
Knowl. Based Syst.1
2018 Joint Modeling of Participant Influence and Latent Topics for Recommendation in Event-based Social Networks
abstract
Event-based social networks (EBSNs) are becoming popular in recent years. Users can publish a planned event on an EBSN website, calling for other users to participate in the event. When a user is making a decision on whether to participate in an event in EBSNs, one aspect for consideration is existing participants defined as users who have agreed to join this event. Existing participants of the event may affect the decision of the user, to which we refer as participant influence. However, participant influence is not well studied by previous works. In this article, we propose an event recommendation model that considers participant influence, and exploits the influence of existing participants on the decisions of new participants based on Poisson factorization. The effect of participant influence is associated with the target event, the host group of the event, and the location of the event. Furthermore, our proposed model can extract latent event topics from event text descriptions, and characterize events, groups, and locations by distributions of event topics. Associations between latent event topics and participant influence are exploited for improving event recommendation. Besides making event recommendation, the proposed model is able to reveal the semantic properties of the participant influence between two users semantically. We have conducted extensive experiments on some datasets extracted from a real-world EBSN. Our proposed model achieves superior event recommendation performance over several state-of-the-art models. The results demonstrate that the consideration of participant influence can improve event recommendation.
Wai Lam, Lidong Bing, Xin Shen 0003
ACM Trans. Inf. Syst.3
2017 Bootstrapping Distantly Supervised IE Using Joint Learning and Small Well-Structured Corpora
abstract
We propose a framework to improve the performance of distantly-supervised relation extraction, by jointly learning to solve two related tasks: concept-instance extraction and relation extraction. We further extend this framework to make a novel use of document structure: in some small, well-structured corpora, sections can be identified that correspond to relation arguments, and distantly-labeled examples from such sections tend to have good precision. Using these as seeds we extract additional relation examples by applying label propagation on a graph composed of noisy examples extracted from a large unstructured testing corpus. Combined with the soft constraint that concept examples should have the same type as the second argument of the relation, we get significant improvements over several state-of-the-art approaches to distantly-supervised relation extraction, and reasonable extraction performance even with very small set of distant labels.
Lidong Bing, Bhuwan Dhingra, Kathryn Mazaitis, Jonghyuk Park 0004, William W. Cohen
AAAI1
2017 Salience Estimation via Variational Auto-Encoders for Multi-Document Summarization
abstract
We propose a new unsupervised sentence salience framework for Multi-Document Summarization (MDS), which can be divided into two components: latent semantic modeling and salience estimation. For latent semantic modeling, a neural generative model called Variational Auto-Encoders (VAEs) is employed to describe the observed sentences and the corresponding latent semantic representations. Neural variational inference is used for the posterior inference of the latent variables. For salience estimation, we propose an unsupervised data reconstruction framework, which jointly considers the reconstruction for latent semantic space and observed term vector space. Therefore, we can capture the salience of sentences from these two different and complementary vector spaces. Thereafter, the VAEs-based latent semantic model is integrated into the sentence salience estimation component in a unified fashion, and the whole framework can be trained jointly by back-propagation via multi-task learning. Experimental results on the benchmark datasets DUC and TAC show that our framework achieves better performance than the state-of-the-art models.
Piji Li, Wai Lam, Zhaochun Ren, Lidong Bing
AAAI5
2017 Recurrent Attention Network on Memory for Aspect Sentiment Analysis
abstract
We propose a novel framework based on neural networks to identify the sentiment of opinion targets in a comment/review.Our framework adopts multiple-attention mechanism to capture sentiment features separated by a long distance, so that it is more robust against irrelevant information.The results of multiple attentions are non-linearly combined with a recurrent neural network, which strengthens the expressive power of our model for handling more complications.The weightedmemory mechanism not only helps us avoid the labor-intensive feature engineering work, but also provides a tailor-made memory for different opinion targets of a sentence.We examine the merit of our model on four datasets: two are from Se-mEval2014, i.e. reviews of restaurants and laptops; a twitter dataset, for testing its performance on social media data; and a Chinese news comment dataset, for testing its language sensitivity.The experimental results show that our model consistently outperforms the state-of-the-art methods on different types of data.
Zhongqian Sun, Lidong Bing
EMNLP3
2017 Cascaded Attention based Unsupervised Information Distillation for Compressive Summarization
abstract
When people recall and digest what they have read for writing summaries, the important content is more likely to attract their attention.Inspired by this observation, we propose a cascaded attention based unsupervised model to estimate the salience information from the text for compressive multi-document summarization.The attention weights are learned automatically by an unsupervised data reconstruction framework which can capture the sentence salience.By adding sparsity constraints on the number of output vectors, we can generate condensed information which can be treated as word salience.Fine-grained and coarse-grained sentence compression strategies are incorporated to produce compressive summaries.Experiments on some benchmark data sets show that our framework achieves better results than the state-of-the-art methods.
Piji Li, Wai Lam, Lidong Bing, Weiwei Guo
EMNLP3
2017 Deep Recurrent Generative Decoder for Abstractive Text Summarization
abstract
We propose a new framework for abstractive text summarization based on a sequence-to-sequence oriented encoderdecoder model equipped with a deep recurrent generative decoder (DRGN).Latent structure information implied in the target summaries is learned based on a recurrent latent random model for improving the summarization quality.Neural variational inference is employed to address the intractable posterior inference for the recurrent latent variables.Abstractive summaries are generated based on both the generative latent variables and the discriminative deterministic states.Extensive experiments on some benchmark datasets in different languages show that DRGN achieves improvements over the state-ofthe-art methods.
Piji Li, Wai Lam, Lidong Bing
EMNLP3
2017 Using Graphs of Classifiers to Impose Declarative Constraints on Semi-supervised Learning
abstract
We propose a general approach to modeling semi-supervised learning (SSL) algorithms. Specifically, we present a declarative language for modeling both traditional supervised classification tasks and many SSL heuristics, including both well-known heuristics such as co-training and novel domain-specific heuristics. In addition to representing individual SSL heuristics, we show that multiple heuristics can be automatically combined using Bayesian optimization methods. We experiment with two classes of tasks, link-based text classification and relation extraction. We show modest improvements on well-studied link-based classification benchmarks, and state-of-the-art results on relation-extraction tasks for two realistic domains.
Lidong Bing, William W. Cohen, Bhuwan Dhingra
IJCAI1
2017 Neural Rating Regression with Abstractive Tips Generation for Recommendation
abstract
Recently, some E-commerce sites launch a new interaction box called Tips on their mobile apps. Users can express their experience and feelings or provide suggestions using short texts typically several words or one sentence. In essence, writing some tips and giving a numerical rating are two facets of a user's product assessment action, expressing the user experience and feelings. Jointly modeling these two facets is helpful for designing a better recommendation system. While some existing models integrate text information such as item specifications or user reviews into user and item latent factors for improving the rating prediction, no existing works consider tips for improving recommendation quality. We propose a deep learning based framework named NRT which can simultaneously predict precise ratings and generate abstractive tips with good linguistic quality simulating user experience and feelings. For abstractive tips generation, gated recurrent neural networks are employed to "translate'' user and item latent representations into a concise sentence. Extensive experiments on benchmark datasets from different domains show that NRT achieves significant improvements over the state-of-the-art methods. Moreover, the generated tips can vividly predict the user experience and feelings.
Piji Li, Zhaochun Ren, Lidong Bing, Wai Lam
SIGIR4
2017 Towards a language-independent solution: Knowledge base completion by searching the Web and deriving language pattern
Lidong Bing, Wai Lam, William W. Cohen
Knowl. Based Syst.1
2016 Distant IE by Bootstrapping Using Lists and Document Structure
abstract
Distant labeling for information extraction (IE) suffers from noisy training data. We describe a way of reducing the noise associated with distant IE by identifying coupling constraints between potential instance labels. As one example of coupling,items in a list are likely to have the same label.A second example of coupling comes from analysis of document structure: in some corpora,sections can be identified such that items in the same section are likely to have the same label. Such sections do not exist in all corpora, but we show that augmenting a large corpus with coupling constraints from even a small, well-structured corpus can improve performance substantially, doubling F1 on one task.
Lidong Bing, Mingyang Ling 0001, Richard C. Wang, William W. Cohen
AAAI1
2016 Detecting Common Discussion Topics Across Culture From News Reader Comments
abstract
News reader comments found in many on-line news websites are typically massive in amount.We investigate the task of Cultural-common Topic Detection (CTD), which is aimed at discovering common discussion topics from news reader comments written in different languages.We propose a new probabilistic graphical model called MCTA which can cope with the language gap and capture the common semantics in different languages.We also develop a partially collapsed Gibbs sampler which effectively incorporates the term translation relationship into the detection of cultural-common topics for model parameter learning.Experimental results show improvements over the state-of-the-art model.
Bei Shi, Wai Lam, Lidong Bing, Yinqing Xu
ACL (1)3
2016 Digesting Multilingual Reader Comments via Latent Discussion Topics with Commonality and Specificity
abstract
Many news websites from different regions in the world allow readers to write comments in their own languages about an event. Digesting such enormous amount of comments in different languages is difficult. One elegant way to digest and organize these comments is to detect latent discussion topics with the consideration of language attributes. Some discussion topics are common topics shared between languages whereas some topics are specifically dominated by a particular language. To tackle this task of discovering discussion topics that exhibit commonality or specificity from news reader comments written in different languages, we propose a new model called TDCS based on graphical models, which can cope with the language gap and detect language-common and language-specific latent discussion topics simultaneously. Our TDCS model also exploits comment-oriented clues via a scalable Dirichlet Multinomial Regression method. To learn the model parameters, we develop an inference method which alternates between EM and Gibbs sampling. Experimental results show that our proposed TDCS model can provide an effective way to digest multilingual news reader comments.
Bei Shi, Wai Lam, Lidong Bing, Yinqing Xu
CIKM3
2016 Efficient and Scalable Detection of Overlapping Communities in Big Networks
abstract
Community detection is a hot topic for researchers in the fields including graph theory, social networks and biological networks. Generally speaking, a community refers to a group of densely linked nodes in the network. Nodes usually have more than one community label, indicating their multiple roles or functions in the network. Unfortunately, existing solutions aiming at overlapping-community-detection are not capable of scaling to large-scale networks with millions of nodes and edges. In this paper, we propose a fast overlapping-communitydetection algorithm - FOX. In the experiment on a network with 3.9 millions nodes and 20 millions edges, the detection finishes in 14 minutes and provides the most qualified results. The second fastest algorithm, however, takes ten times longer to run. As for another network with 22 millions nodes and 127 millions edges, our algorithm is the only one that can provide an overlapping community detection result and it only takes 238 minutes. Our algorithm draws lessons from potential games, a concept in game theory. We measure the closeness of a node to a community by counting the number of triangles formed by the node and two other nodes form the community. Potential games ensure that the algorithm can reach convergence. We also extend the exploitation of triangle to open-triangle, which enlarges the scale of the detected communities.
Tianshu Lyu, Lidong Bing, Yan Zhang 0004
ICDM2
2016 Unsupervised Extraction of Popular Product Attributes from E-Commerce Web Sites by Considering Customer Reviews
abstract
We develop an unsupervised learning framework for extracting popular product attributes from product description pages originated from different E-commerce Web sites. Unlike existing information extraction methods that do not consider the popularity of product attributes, our proposed framework is able to not only detect popular product features from a collection of customer reviews but also map these popular features to the related product attributes. One novelty of our framework is that it can bridge the vocabulary gap between the text in product description pages and the text in customer reviews. Technically, we develop a discriminative graphical model based on hidden Conditional Random Fields. As an unsupervised model, our framework can be easily applied to a variety of new domains and Web sites without the need of labeling training samples. Extensive experiments have been conducted to demonstrate the effectiveness and robustness of our framework.
Lidong Bing, Tak-Lam Wong, Wai Lam
ACM Trans. Internet Techn.1
2015 Abstractive Multi-Document Summarization via Phrase Selection and Merging
abstract
Lidong Bing, Piji Li, Yi Liao, Wai Lam, Weiwei Guo, Rebecca Passonneau. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Lidong Bing, Piji Li, Wai Lam, Weiwei Guo, Rebecca J. Passonneau
ACL (1)1
2015 A Unified Posterior Regularized Topic Model with Maximum Margin for Learning-to-Rank
abstract
While most methods for learning-to-rank documents only consider relevance scores as features, better results can often be obtained by taking into account the latent topic structure of the document collection. Existing approaches that consider latent topics follow a two-stage approach, in which topics are discovered in an unsupervised way, as usual, and then used as features for the learning-to-rank task. In contrast, we propose a learning-to-rank framework which integrates the supervised learning of a maximum margin classifier with the discovery of a suitable probabilistic topic model. In this way, the labelled data that is available for the learning-to-rank task can be exploited to identify the most appropriate topics. To this end, we use a unified constrained optimization framework, which can dynamically compute the latent topic similarity score between the query and the document. Our experimental results show a consistent improvement over the state-of-the-art learning-to-rank models.
Shoaib Jameel, Wai Lam, Steven Schockaert, Lidong Bing
CIKM4
2015 Nonparametric Topic Modeling Using Chinese Restaurant Franchise with Buddy Customers
Shoaib Jameel, Wai Lam, Lidong Bing
ECIR3
2015 Improving Distant Supervision for Information Extraction Using Label Propagation Through Lists
abstract
Because of polysemy, distant labeling for information extraction leads to noisy training data.We describe a procedure for reducing this noise by using label propagation on a graph in which the nodes are entity mentions, and mentions are coupled when they occur in coordinate list structures.We show that this labeling approach leads to good performance even when off-the-shelf classifiers are used on the distantly-labeled data.
Lidong Bing, Sneha Chaudhari, Richard C. Wang, William W. Cohen
EMNLP1
2015 Reader-Aware Multi-Document Summarization via Sparse Coding
Piji Li, Lidong Bing, Wai Lam
IJCAI2
2015 Supervised topic models with word order structure for document classification and retrieval learning
Shoaib Jameel, Wai Lam, Lidong Bing
Inf. Retr. J.3
2015 Adaptive Concept Resolution for document representation and its applications in text mining
Lidong Bing, Shan Jiang 0001, Wai Lam, Yan Zhang 0004, Shoaib Jameel
Knowl. Based Syst.1
2015 Web Query Reformulation via Joint Modeling of Latent Topic Dependency and Term Context
abstract
An important way to improve users’ satisfaction in Web search is to assist them by issuing more effective queries. One such approach is query reformulation, which generates new queries according to the current query issued by users. A common procedure for conducting reformulation is to generate some candidate queries first, then a scoring method is employed to assess these candidates. Currently, most of the existing methods are context based. They rely heavily on the context relation of terms in the history queries and cannot detect and maintain the semantic consistency of queries. In this article, we propose a graphical model to score queries. The proposed model exploits a latent topic space, which is automatically derived from the query log, to detect semantic dependency of terms in a query and dependency among topics. Meanwhile, the graphical model also captures the term context in the history query by skip-bigram and n-gram language models. In addition, our model can be easily extended to consider users’ history search interests when we conduct query reformulation for different users. In the task of candidate query generation, we investigate a social tagging data resource—Delicious bookmark—to generate addition and substitution patterns that are employed as supplements to the patterns generated from query log data.
Lidong Bing, Wai Lam, Tak-Lam Wong, Shoaib Jameel
ACM Trans. Inf. Syst.1
2014 Website Community Mining from Query Logs with Two-Phase Clustering
Lidong Bing, Wai Lam, Shoaib Jameel, Chunliang Lu
CICLing (2)1
2014 Web page segmentation with structured prediction and its application in web page classification
abstract
We propose a framework which can perform Web page segmentation with a structured prediction approach. It formulates the segmentation task as a structured labeling problem on a transformed Web page segmentation graph (WPS-graph). WPS-graph models the candidate segmentation boundaries of a page and the dependency relation among the adjacent segmentation boundaries. Each labeling scheme on the WPS-graph corresponds to a possible segmentation of the page. The task of finding the optimal labeling of the WPS-graph is transformed into a binary Integer Linear Programming problem, which considers the entire WPS-graph as a whole to conduct structured prediction. A learning algorithm based on the structured output Support Vector Machine framework is developed to determine the feature weights, which is capable to consider the inter-dependency among candidate segmentation boundaries. Furthermore, we investigate its efficacy in supporting the development of automatic Web page classification.
Lidong Bing, Wai Lam, Zhengyu Niu, Haifeng Wang 0001
SIGIR1
2013 Towards an enhanced and adaptable ontology by distilling and assembling online encyclopedias
abstract
In this paper, we investigate the problem of making better use of semantic knowledge obtained from different encyclopedia sources. We propose a framework to integrate different encyclopedias and reorganize the information. We also utilize Learning to Rank models to distill out more functional knowledge from the encyclopedic information and then align the knowledge with a WordNet-like ontology. Finally as a demonstration, a Chinese semantic knowledge repository named JNet is constructed based on this framework. Experiments show that the proposed methods work well and the three steps reinforce each other towards a more powerful ontology.
Shan Jiang 0001, Lidong Bing, Yan Zhang 0004
CIKM2
2013 Structured positional entity language model for enterprise entity retrieval
abstract
We investigate the problem of general entity retrieval for enterprise websites. Our framework transforms the webpage content into a structured content representation, which captures hierarchical information blocks and semi-structured data records information. To facilitate entity retrieval given a user query, we develop a structured positional entity language model suitable for ranking entities extracted from the webpage content incorporating the structured content representation. Different from existing language models for retrieval, our proposed model considers both the proximity and the structured webpage content in a unified manner. Extensive experiments on the benchmark datasets demonstrate the effectiveness of our proposed framework.
Chunliang Lu, Lidong Bing, Wai Lam
CIKM2
2013 Wikipedia entity expansion and attribute extraction from the web using semi-supervised learning
abstract
We develop a new framework to achieve the goal of Wikipedia entity expansion and attribute extraction from the Web. Our framework takes a few existing entities that are automatically collected from a particular Wikipedia category as seed input and explores their attribute infoboxes to obtain clues for the discovery of more entities for this category and the attribute content of the newly discovered entities. One characteristic of our framework is to conduct discovery and extraction from desirable semi-structured data record sets which are automatically collected from the Web. A semi-supervised learning model with Conditional Random Fields is developed to deal with the issues of extraction learning and limited number of labeled examples derived from the seed entities. We make use of a proximate record graph to guide the semi-supervised learning process. The graph captures alignment similarity among data records. Then the semi-supervised learning process can leverage the unlabeled data in the record set by controlling the label regularization under the guidance of the proximate record graph. Extensive experiments on different domains have been conducted to demonstrate its superiority for discovering new entities and extracting attribute content.
Lidong Bing, Wai Lam, Tak-Lam Wong
WSDM1
2013 Robust detection of semi-structured web records using a DOM structure-knowledge-driven model
abstract
Web data record extraction aims at extracting a set of similar object records from a single webpage. These records have similar attributes or fields and are presented with a regular format in a coherent region of the page. To tackle this problem, most existing works analyze the DOM tree of an input page. One major limitation of these methods is that the lack of a global view in detecting data records from an input page results in a myopic decision. Their brute-force searching manner in detecting various types of records degrades the flexibility and robustness. We propose a Structure-Knowledge-Oriented Global Analysis (Skoga) framework which can perform robust detection of different-kinds of data records and record regions. The major component of the Skoga framework is a DOM structure-knowledge-driven detection model which can conduct a global analysis on the DOM structure to achieve effective detection. The DOM structure knowledge consists of background knowledge as well as statistical knowledge capturing different characteristics of data records and record regions, as exhibited in the DOM structure. The background knowledge encodes the semantics of labels indicating general constituents of data records and regions. The statistical knowledge is represented by some carefully designed features that capture different characteristics of a single node or a node group in the DOM. The feature weights are determined using a development dataset via a parameter estimation algorithm based on a structured output support vector machine. An optimization method based on the divide-and-conquer principle is developed making use of the DOM structure knowledge to quantitatively infer and recognize appropriate records and regions for a page. Extensive experiments have been conducted on four datasets. The experimental results demonstrate that our framework achieves higher accuracy compared with state-of-the-art methods.
Lidong Bing, Wai Lam, Tak-Lam Wong
ACM Trans. Web1
2011 Towards a unified solution: data record region detection and segmentation
abstract
Although the task of data record extraction from Web pages has been studied extensively, yet it fails to handle many pages due to their complexity in format or layout. In this paper, we propose a unified method to tackle this task by addressing several key issues in a uniform manner. A new search structure, named as Record Segmentation Tree (RST), is designed, and several efficient search pruning strategies on the RST structure are proposed to identify the records in a given Web page. Another characteristic of our method which is significantly different from previous works is that it can effectively handle complicated and challenging data record regions. It is achieved by generating subtree groups dynamically from the RST structure during the search process. Furthermore, instead of using string edit distance or tree edit distance, we propose a token-based edit distance which takes each DOM node as a basic unit in the cost calculation. Extensive experiments are conducted on four data sets, including flat, nested, and intertwine records. The experimental results demonstrate that our method achieves higher accuracy compared with three state-of-the-art methods.
Lidong Bing, Wai Lam, Yuan Gu
CIKM1
2011 Using query log and social tagging to refine queries based on latent topics
abstract
An important way to improve users' satisfaction in Web search is to assist them to issue more effective queries. One such approach is query refinement (reformulation), which generates new queries according to the current query issued by users. A common procedure for conducting refinement is to generate some candidate queries first, and then a scoring method is designed to assess the quality of these candidates. Currently, most of the existing methods are context based. They rely heavily on the context relation of terms in the historical queries, and cannot detect and maintain the semantic consistency of queries. In this paper, we propose a graphical model to score queries. The proposed model exploits a latent topic space, which is automatically derived from the query log, to assess the semantic dependency of terms in a query. In the graphical model, both term context dependency and topic context dependency are considered. This also makes it feasible to score some queries which do not have much available historical term context information. We also utilize social tagging data in the candidate query generation process. Based on the observation that different users may tag the same resource with different tags of similar meaning, we propose a method to mine these term pairs for new candidate query construction.
Lidong Bing, Wai Lam, Tak-Lam Wong
CIKM1
2011 Ontology enhancement and concept granularity learning: keeping yourself current and adaptive
abstract
As a well-known semantic repository, WordNet is widely used in many applications. However, due to costly edit and maintenance, WordNet's capability of keeping up with the emergence of new concepts is poor compared with on-line encyclopedias such as Wikipedia. To keep WordNet current with folk wisdom, we propose a method to enhance WordNet automatically by merging Wikipedia entities into WordNet, and construct an enriched ontology, named as WorkiNet. WorkiNet keeps the desirable structure of WordNet. At the same time, it captures abundant information from Wikipedia. We also propose a learning approach which is able to generate a tailor-made semantic concept collection for a given document collection. The learning process takes the characteristics of the given document collection into consideration and the semantic concepts in the tailor-made collection can be used as new features for document representation. The experimental results show that the adaptively generated feature space can outperform a static one significantly in text mining tasks, and WorkiNet dominates WordNet most of the time due to its high coverage.
Shan Jiang 0001, Lidong Bing, Bai Sun, Yan Zhang 0004, Wai Lam
KDD2
2011 Normalizing web product attributes and discovering domain ontology with minimal effort
abstract
We have developed a framework aiming at normalizing product attributes from Web pages collected from different Web sites without the need of labeled training examples. It can deal with pages composed of different layout format and content in an unsupervised manner. As a result, it can handle a variety of different domains with minimal effort. Our model is based on a generative probabilistic graphical model incorporated with Hidden Markov Models (HMM) considering both attribute names and attribute values to extract and normalize text fragments from Web pages in a unified manner. Dirichlet Process is employed to handle the unlimited number of attributes in a domain. An unsupervised inference method is proposed to predict the unobservable variables. We have also developed a method to automatically construct a domain ontology using the normalized product attributes which are the output of the inference on the graphical model. We have conducted extensive experiments and compared with existing works using prouct Web pages collected from real-world Web sites in three different domains to demonstrate the effectiveness of our framework.
Tak-Lam Wong, Lidong Bing, Wai Lam
WSDM2
2010 Learning ontology resolution for document representation and its applications in text mining
abstract
It is well known that synonymous and polysemous terms often bring in some noises when calculating the similarity between documents. Existing ontology-based document representation methods are static, hence, the chosen semantic concept set for representing a document has a fixed resolution and it is not adaptable to the characteristics of a document collection and the text mining problem in hand. We propose an Adaptive Concept Resolution (ACR) model to overcome this issue. ACR can learn a concept border from an ontology taking into consideration of the characteristics of a particular document collection. Then this border can provide a tailor-made semantic concept representation for a document coming from the same domain. Another advantage of ACR is that it is applicable in both classification task where the groups are given in the training document set, and clustering task where no group information is available. Furthermore, the result of this model is not sensitive to the model parameter. The experimental results show that ACR outperforms an existing static method significantly.
Lidong Bing, Bai Sun, Shan Jiang 0001, Yan Zhang 0004, Wai Lam
CIKM1
2008 Weighting Links Using Lexical and Positional Analysis in Web Ranking
abstract
Link analysis has been widely used to evaluate the importance of Web pages. Popular link analysis algorithms are mainly based on the link structure between pages. However, a Web page usually contains various links such as for navigation, decoration or nepotism, which are irrelevant to the topic of the Web page and can not reflect the actual voting relations between pages. In order to improve the performance of Web ranking, we bring out one filtering algorithm to recognize and eliminate these unrelated links using Content Lexical and Positional analysis. Experimental results on different Web domains show that our filtering model can efficiently detect the irrelevant links and effectively help to build a good link graph for the ranking calculation.
Yi Zhang 0012, Yexin Wang, Lidong Bing, Yan Zhang 0004
WAIM3