EDBT 2026 Demo / reviewers in the wild / expert
Xin Li 0056
dblp:09/1365-56
· DBLP profile ↗
32ranked-venue papers
5as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Affordance-Aware Robotic Dexterous Grasping with Human-like PriorsabstractA dexterous hand capable of generalizable grasping objects is fundamental for the development of general-purpose embodied AI. However, previous methods focus narrowly on low-level grasp stability metrics, neglecting affordance-aware positioning and human-like poses which are crucial for downstream manipulation. To address these limitations, we propose AffordDex, a novel framework with two-stage training that learns a universal grasping policy with an inherent understanding of both motion priors and object affordances. In the first stage, a trajectory imitator is pre-trained on a large corpus of human hand motions to instill a strong prior for natural movement. In the second stage, a residual module is trained to adapt these general human-like motions to specific object instances. This refinement is critically guided by two components: our Negative Affordance-aware Segmentation (NAA) module, which identifies functionally inappropriate contact regions, and a privileged teacher-student distillation process that ensures the final vision-based policy is highly successful. Extensive experiments demonstrate that AffordDex not only achieves universal dexterous grasping but also remains remarkably human-like in posture and functionally appropriate in contact location. As a result, AffordDex significantly outperforms state-of-the-art baselines across seen objects, unseen instances, and even entirely novel categories. Linghao Zhuang, Xingyue Zhao, Yuming Jiang 0007, Jun Cen, Kexiang Wang, Jiayan Guo, Siteng Huang, Xin Li 0056, Deli Zhao, Hua Zou 0002 |
AAAI | 11 |
| 2026 | Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive TasksabstractWenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenqi Zhang 0001, Mengna Wang, Gangao Liu, Huixin Xu, Yongliang Shen 0001, Guiyang Hou, Xin Li 0056, Weiming Lu 0001, Peng Li 0031, Yueting Zhuang |
ACL (1) | 10 |
| 2025 | Breaking the Memory Barrier of Contrastive Loss via Tile-Based StrategyabstractContrastive loss is a powerful approach for representation learning, where larger batch sizes enhance performance by providing more negative samples to better distinguish between similar and dissimilar data. However, the full instantiation of the similarity matrix demands substantial GPU memory, making large batch training highly resource-intensive. To address this, we propose a tile-based computation strategy that partitions the contrastive loss calculation into small blocks, avoiding full materialization of the similarity matrix. Additionally, we introduce a multi-level tiling implementation to leverage the hierarchical structure of distributed systems, using ring-based communication at the GPU level to optimize synchronization and fused kernels at the CUDA core level to reduce I/O overhead. Experimental results show that the proposed method significantly reduces GPU memory usage in contrastive loss. For instance, it enables contrastive training of a CLIP-ViT-L/14 model with a batch size of 4M using only 8 A800 80GB GPUs, without sacrificing accuracy. Compared to state-of-the-art memory-efficient solutions, it achieves a two-order-of-magnitude reduction in memory while maintaining comparable speed. The code will be made publicly available.1 Zesen Cheng, Sicong Leng, Deli Zhao, Xin Li 0056, Lidong Bing |
CVPR | 8 |
| 2025 | ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkabstractThe enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code is available at https://github.com/RhDang/ECBench. Ronghao Dang, Yuqian Yuan, Wenqi Zhang 0001, Yifei Xin, Boqiang Zhang, Liuyi Wang, Qinyang Zeng, Xin Li 0056, Lidong Bing |
CVPR | 9 |
| 2025 | VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMabstractVideo Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a largescale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities. Yuqian Yuan, Wentong Li 0001, Zesen Cheng, Boqiang Zhang, Xin Li 0056, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing |
CVPR | 7 |
| 2025 | 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingabstractCompared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook. Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing |
ICCV | 3 |
| 2025 | LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference OptimizationabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities through pretraining and alignment. However, superior short-context LLMs may underperform in long-context scenarios due to insufficient long-context alignment. This alignment process remains challenging due to the impracticality of human annotation for extended contexts and the difficulty in balancing short- and long-context performance. To address these challenges, we introduce LongPO, that enables short-context LLMs to self-evolve to excel on long-context tasks by internally transferring short-context capabilities. LongPO harnesses LLMs to learn from self-generated short-to-long preference data, comprising paired responses generated for identical instructions with long-context inputs and their compressed short-context counterparts, respectively. This preference reveals capabilities and potentials of LLMs cultivated during short-context alignment that may be diminished in under-aligned long-context scenarios. Additionally, LongPO incorporates a short-to-long KL constraint to mitigate short-context performance decline during long-context alignment. When applied to Mistral-7B-Instruct-v0.2 from 128K to 512K context lengths, LongPO fully retains short-context performance and largely outperforms naive SFT and DPO in both long- and short-context tasks. Specifically, LongPO-trained models can achieve results on long-context benchmarks comparable to, or even surpassing, those of superior LLMs (e.g., GPT-4-128K) that involve extensive long-context annotation and larger parameter scales. Our code is available at https://github.com/DAMO-NLP-SG/LongPO. Guanzheng Chen, Xin Li 0056, Michael Shieh, Lidong Bing |
ICLR | 2 |
| 2025 | RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingabstractThe emergence of long-context large language models (LLMs) offers a promising alternative to traditional retrieval-augmented generation (RAG) for processing extensive documents. However, the computational overhead of long-context inference presents significant efficiency challenges. While Speculative Decoding (SD) traditionally accelerates inference using smaller draft models, its effectiveness diminishes substantially in long-context scenarios due to memory-bound KV cache operations. We introduce Retrieval-Augmented Speculative Decoding (RAPID), which leverages RAG for both accelerating and enhancing generation quality in long-context inference. RAPID introduces the RAG drafter—a draft LLM operating on shortened retrieval contexts—to speculate on the generation of long-context target LLMs. Our approach enables a new paradigm where same-scale or even larger LLMs can serve as RAG drafters while maintaining computational efficiency. To fully leverage the potentially superior capabilities from stronger RAG drafters, we develop an inference-time knowledge transfer that enriches the target distribution by RAG. Extensive experiments on the LLaMA-3.1 and Qwen2.5 backbones demonstrate that RAPID effectively integrates the strengths of both RAG and long-context LLMs, achieving significant performance improvements (e.g., from 39.33 to 42.83 on InfiniteBench for LLaMA-3.1-8B) with more than 2$\times$ speedups for long-context inference. Our analyses also reveal the robustness of RAPID across various context lengths and retrieval quality. Guanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li 0056, Michael Shieh |
ICML | 4 |
| 2025 | The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and AudioabstractRecent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most existing LMMs remain vulnerable to hallucinations, the discrepancy between the factual multimodal input and the generated textual output, which has limited their applicability in various real-world scenarios. This paper presents the first systematic investigation of hallucinations in LMMs involving the three most common modalities: language, visual, and audio. Our study reveals two key contributors to hallucinations: overreliance on unimodal priors and spurious inter-modality correlations. To address these challenges, we introduce the benchmark The Curse of Multi-Modalities (CMM), which comprehensively evaluates hallucinations in LMMs, providing a detailed analysis of their underlying issues. Our findings highlight key vulnerabilities, including imbalances in modality integration and biases from training data, underscoring the need for balanced cross-modal learning and enhanced hallucination mitigation strategies. Based on our observations and findings, we suggest potential research directions that could enhance the reliability of LMMs. Sicong Leng, Zesen Cheng, Xin Li 0056, Deli Zhao, Shijian Lu, Chunyan Miao, Lidong Bing |
NeurIPS | 6 |
| 2025 | EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?abstractThe emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions.capabilities in object-level spatiotemporal reasoning required for real-world interactions.To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios.Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types.To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation frameworkBased on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems. Yuqian Yuan, Ronghao Dang, Wentong Li 0001, Xin Li 0056, Deli Zhao, Fan Wang 0019, Wenqiao Zhang, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 6 |
| 2024 | Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingabstractLarge Vision-Language Models (LVLMs) have advanced considerably, intertwining visual recognition and language understanding to generate content that is not only coherent but also contextually attuned. Despite their success, LVLMs still suffer from the issue of object hallucinations, where models generate plausible yet incorrect outputs that include objects that do not exist in the images. To mitigate this issue, we introduce Visual Contrastive Decoding (VCD), a simple and training-free method that contrasts output distributions derived from original and distorted visual inputs. The proposed VCD effectively reduces the over-reliance on statistical bias and unimodal priors, two essential causes of object hallucinations. This adjustment ensures the generated content is closely grounded to visual inputs, resulting in contextually accurate outputs. Our experiments show that VCD, without either additional training or the usage of external tools, significantly mitigates the object hallucination issue across different LVLM families. Beyond mitigating object hallucinations, VCD also excels in general LVLM benchmarks, highlighting its wide-ranging applicability. Sicong Leng, Guanzheng Chen, Xin Li 0056, Shijian Lu, Chunyan Miao, Lidong Bing |
CVPR | 4 |
| 2024 | AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code GenerationabstractThe impressive performance of proprietary LLMs like GPT4 in code generation has led to a trend to replicate these capabilities in open-source models through knowledge distillation (e.g.Code Evol-Instruct).However, these efforts often neglect the crucial aspect of response quality, relying heavily on teacher models for direct response distillation.This paradigm, especially for complex instructions, can degrade the quality of synthesized data, compromising the knowledge distillation process.To this end, our study introduces the Adaptive Modular Response Evolution (AMR-Evol) framework, which employs a two-stage process to refine response distillation.The first stage, modular decomposition, breaks down the direct response into more manageable sub-modules.The second stage, adaptive response evolution, automatically evolves the response with the related function modules.Our experiments with three popular code benchmarks-HumanEval, MBPP, and EvalPlus-attests to the superiority of the AMR-Evol framework over baseline response distillation methods.By comparing with the open-source Code LLMs trained on a similar scale of data, we observed performance enhancements: more than +3.0 points on HumanEval-Plus and +1.0 points on MBPP-Plus, which underscores the effectiveness of our framework.Our codes are available at https://github.com/ChiYeungLaw/ AMR-Evol. Module 3def validate_matrix(matrix: list) -> None:""" Description:Validates if the input matrix is a 3x3 matrix.Parameters:-matrix (list): The input matrix to be validated.Raises:-ValueError: If the matrix is not a 3x3 matrix.""" Coding Tasks Design a function that implements the Laplace expansion theorem to calculate the determinant of a 3x3 matrix, but now with the added condition of utilizing nested for loops and ensuring that the matrix elements are limited to integers only.Teacher Model Response def calculate_determinant(matrix): if len(matrix) != 3 or len(matrix[0]) != 3: raise ValueError('Matrix must be a 3x3 matrix') determinant = 0 for i in range(3): submatrix = [[matrix[j][k] for k in range(3) if k != i] for j in range(1, 3)] sub_determinant = calculate_determinant(submatrix) sign = (-1) ** i determinant += sign * matrix[0][i] * sub_determinant return determinant if len(matrix) == 2 else round(determinant) Direct Distillation Module 2def calculate_minor_matrix(matrix: list, row: int, col: int) -> list:""" Description: Calculates the minor matrix by removing the specified row Parameters:-matrix (list): The input matrix.-row (int): The row index to be removed.-col (int): The column index to be removed. Xin Li 0056, Hongzhan Lin 0001, Jing Ma 0004, Lidong Bing |
EMNLP | 2 |
| 2024 | CLEX: Continuous Length Extrapolation for Large Language ModelsabstractTransformer-based Large Language Models (LLMs) are pioneering advances in many natural language processing tasks, however, their exceptional capabilities are restricted within the preset context window of Transformer. Position Embedding (PE) scaling methods, while effective in extending the context window to a specific length, demonstrate either notable limitations in their extrapolation abilities or sacrificing partial performance within the context window. Length extrapolation methods, although theoretically capable of extending the context window beyond the training sequence length, often underperform in practical long-context applications. To address these challenges, we propose Continuous Length EXtrapolation (CLEX) for LLMs. We generalise the PE scaling approaches to model the continuous dynamics by ordinary differential equations over the length scaling factor, thereby overcoming the constraints of current PE scaling methods designed for specific lengths. Moreover, by extending the dynamics to desired context lengths beyond the training sequence length, CLEX facilitates the length extrapolation with impressive performance in practical tasks. We demonstrate that CLEX can be seamlessly incorporated into LLMs equipped with Rotary Position Embedding, such as LLaMA and GPT-NeoX, with negligible impact on training and inference latency. Experimental results reveal that CLEX can effectively extend the context window to over 4× or almost 8× training length, with no deterioration in performance. Furthermore, when evaluated on the practical LongBench benchmark, our model trained on a 4k length exhibits competitive performance against state-of-the-art open-source models trained on context lengths up to 32k. Our code is available at https://github.com/DAMO-NLP-SG/CLEX. Guanzheng Chen, Xin Li 0056, Zaiqiao Meng, Shangsong Liang, Lidong Bing |
ICLR | 2 |
| 2024 | Stabilize the Latent Space for Image Autoregressive Modeling: A Unified PerspectiveabstractLatent-based image generative models, such as Latent Diffusion Models (LDMs) and Mask Image Models (MIMs), have achieved notable success in image generation tasks. These models typically leverage reconstructive autoencoders like VQGAN or VAE to encode pixels into a more compact latent space and learn the data distribution in the latent space instead of directly from pixels. However, this practice raises a pertinent question: Is it truly the optimal choice? In response, we begin with an intriguing observation: despite sharing the same latent space, autoregressive models significantly lag behind LDMs and MIMs in image generation. This finding contrasts sharply with the field of NLP, where the autoregressive model GPT has established a commanding presence. To address this discrepancy, we introduce a unified perspective on the relationship between latent space and generative models, emphasizing the stability of latent space in image generative modeling. Furthermore, we propose a simple but effective discrete image tokenizer to stabilize the latent space for image generative modeling by applying K-Means on the latent features of self-supervised learning models. Experimental results show that image autoregressive modeling with our tokenizer (DiGIT) benefits both image understanding and image generation with the next token prediction principle, which is inherently straightforward for GPT models but challenging for other generative models. Remarkably, for the first time, a GPT-style autoregressive model for images outperforms LDMs, which also exhibits substantial improvement akin to GPT when scaling up model size. Our findings underscore the potential of an optimized latent space and the integration of discrete tokenization in advancing the capabilities of image generative models. The code is available at \url{https://github.com/DAMO-NLP-SG/DiGIT}. Yongxin Zhu 0003, Bocheng Li, Xin Li 0056, Linli Xu 0002, Lidong Bing |
NeurIPS | 4 |
| 2023 | Improving Self-training for Cross-lingual Named Entity Recognition with Contrastive and Prototype LearningabstractIn cross-lingual named entity recognition (NER), self-training is commonly used to bridge the linguistic gap by training on pseudolabeled target-language data.However, due to sub-optimal performance on target languages, the pseudo labels are often noisy and limit the overall performance.In this work, we aim to improve self-training for cross-lingual NER by combining representation learning and pseudo label refinement in one coherent framework.Our proposed method, namely ContProto mainly comprises two components: (1) contrastive self-training and (2) prototype-based pseudo-labeling.Our contrastive self-training facilitates span classification by separating clusters of different classes, and enhances crosslingual transferability by producing closelyaligned representations between the source and target language.Meanwhile, prototype-based pseudo-labeling effectively improves the accuracy of pseudo labels during training.We evaluate ContProto on multiple transfer pairs, and experimental results show our method brings in substantial improvements over current stateof-the-art methods. 1 Ran Zhou 0004, Xin Li 0056, Lidong Bing, Erik Cambria, Chunyan Miao |
ACL (1) | 2 |
| 2023 | Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsabstractDue to the huge amount of parameters, finetuning of pretrained language models (PLMs) is prone to overfitting in the low resource scenarios.In this work, we present a novel method that operates on the hidden representations of a PLM to reduce overfitting.During fine-tuning, our method inserts random autoencoders between the hidden layers of a PLM, which transform activations from the previous layers into multi-view compressed representations before feeding them into the upper layers.The autoencoders are plugged out after fine-tuning, so our method does not add extra parameters or increase computation cost during inference.Our method demonstrates promising performance improvement across a wide range of sequenceand token-level low-resource NLP tasks.Our code is available at https://github.com/DAMO- NLP-SG/MVCR. Xingxuan Li, Megh Thakkar, Xin Li 0056, Shafiq R. Joty, Luo Si, Lidong Bing |
ACL (1) | 4 |
| 2023 | PeerDA: Data Augmentation via Modeling Peer Relation for Span Identification TasksabstractSpan identification aims at identifying specific text spans from text input and classifying them into pre-defined categories.Different from previous works that merely leverage the Subordinate (SUB) relation (i.e. if a span is an instance of a certain category) to train models, this paper for the first time explores the Peer (PR) relation, which indicates that two spans are instances of the same category and share similar features.Specifically, a novel Peer Data Augmentation (PeerDA) approach is proposed which employs span pairs with the PR relation as the augmentation data for training.PeerDA has two unique advantages: (1) There are a large number of PR span pairs for augmenting the training data.(2) The augmented data can prevent the trained model from over-fitting the superficial span-category mapping by pushing the model to leverage the span semantics.Experimental results on ten datasets over four diverse tasks across seven domains demonstrate the effectiveness of PeerDA.Notably, PeerDA achieves state-of-the-art results on six of them. 1 Weiwen Xu, Xin Li 0056, Yang Deng 0002, Wai Lam, Lidong Bing |
ACL (1) | 2 |
| 2023 | Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal ReasoningabstractOur physical world is constantly evolving over time, rendering challenges for pre-trained language models to understand and reason over the temporal contexts of texts.Existing work focuses on strengthening the direct association between a piece of text and its time-stamp.However, the knowledge-time association is usually insufficient for the downstream tasks that require reasoning over temporal dependencies between knowledge.In this work, we make use of the underlying nature of time, all temporally-scoped sentences are strung together through a one-dimensional time axis, and suggest creating a graph structure based on the relative placements of events along the time axis.Inspired by the graph view, we propose REMEMO (Relative Time Modeling), which explicitly connects all temporally-scoped facts by modeling the time relations between any two sentences.Experimental results show that REMEMO outperforms the baseline T5 on multiple temporal question answering datasets under various settings.Further analysis suggests that REMEMO is especially good at modeling long-range complex temporal dependencies.We release our code and pretrained checkpoints at https://github.com/ DAMO-NLP-SG/RemeMo. Sen Yang 0005, Xin Li 0056, Lidong Bing, Wai Lam |
EMNLP | 2 |
| 2023 | From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine ReaderabstractWe present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data.
PMR can resolve the discrepancy between model pre-training and downstream fine-tuning of existing MLMs.
To build the proposed PMR, we constructed a large volume of general-purpose and high-quality MRC-style training data by using Wikipedia hyperlinks and designed a Wiki Anchor Extraction task to guide the MRC-style pre-training.
Apart from its simplicity, PMR effectively solves extraction tasks, such as Extractive Question Answering and Named Entity Recognition. PMR shows tremendous improvements over existing approaches, especially in low-resource scenarios.
When applied to the sequence classification task in the MRC formulation, PMR enables the extraction of high-quality rationales to explain the classification process, thereby providing greater prediction explainability. PMR also has the potential to serve as a unified model for tackling various extraction and classification tasks in the MRC formulation. Weiwen Xu, Xin Li 0056, Wenxuan Zhang 0001, Wai Lam, Luo Si, Lidong Bing |
NeurIPS | 2 |
| 2023 | A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and ChallengesabstractAs an important fine-grained sentiment analysis problem, aspect-based sentiment analysis (ABSA), aiming to analyze and understand people's opinions at the aspect level, has been attracting considerable interest in the last decade. To handle ABSA in different scenarios, various tasks are introduced for analyzing different sentiment elements and their relations, including the aspect term, aspect category, opinion term, and sentiment polarity. Unlike early ABSA works focusing on a single sentiment element, many compound ABSA tasks involving multiple elements have been studied in recent years for capturing more complete aspect-level sentiment information. However, a systematic review of various ABSA tasks and their corresponding solutions is still lacking, which we aim to fill in this survey. More specifically, we provide a new taxonomy for ABSA which organizes existing studies from the axes of concerned sentiment elements, with an emphasis on recent advances of compound ABSA tasks. From the perspective of solutions, we summarize the utilization of pre-trained language models for ABSA, which improved the performance of ABSA to a new stage. Besides, techniques for building more practical ABSA systems in cross-domain/lingual scenarios are discussed. Finally, we review some emerging topics and discuss some open challenges to outlook potential future directions of ABSA. Wenxuan Zhang 0001, Xin Li 0056, Yang Deng 0002, Lidong Bing, Wai Lam |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERabstractData augmentation is an effective solution to data scarcity in low-resource scenarios.However, when applied to token-level tasks such as NER, data augmentation methods often suffer from token-label misalignment, which leads to unsatsifactory performance.In this work, we propose Masked Entity Language Modeling (MELM) as a novel data augmentation framework for low-resource NER.To alleviate the token-label misalignment issue, we explicitly inject NER labels into sentence context, and thus the fine-tuned MELM is able to predict masked entity tokens by explicitly conditioning on their labels.Thereby, MELM generates high-quality augmented data with novel entities, which provides rich entity regularity knowledge and boosts NER performance.When training data from multiple languages are available, we also integrate MELM with codemixing for further improvement.We demonstrate the effectiveness of MELM on monolingual, cross-lingual and multilingual NER across various low-resource levels.Experimental results show that our MELM presents substantial improvement over the baseline methods. 1 Ran Zhou 0004, Xin Li 0056, Ruidan He, Lidong Bing, Erik Cambria, Luo Si, Chunyan Miao |
ACL (1) | 2 |
| 2022 | Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationabstractWe introduce a new method to improve existing multilingual sentence embeddings with Abstract Meaning Representation (AMR).Compared with the original textual input, AMR is a structured semantic representation that presents the core concepts and relations in a sentence explicitly and unambiguously.It also helps reduce surface variations across different expressions and languages.Unlike most prior work that only evaluates the ability to measure semantic similarity, we present a thorough evaluation of existing multilingual sentence embeddings and our improved versions, which include a collection of five transfer tasks in different downstream applications.Experiment results show that retrofitting multilingual sentence embeddings with AMR leads to better state-of-the-art performance on both semantic textual similarity and transfer tasks.Our codebase and evaluation scripts Deng Cai 0002, Xin Li 0056, Jackie C. S. Ho, Lidong Bing, Wai Lam |
EMNLP | 2 |
| 2022 | Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesabstractKnowledge-enhanced language representation learning has shown promising results across various knowledge-intensive NLP tasks.However, prior methods are limited in efficient utilization of multilingual knowledge graph (KG) data for language model (LM) pretraining.They often train LMs with KGs in indirect ways, relying on extra entity/relation embeddings to facilitate knowledge injection.In this work, we explore methods to make better use of the multilingual annotation and language agnostic property of KG triples, and present novel knowledge based multilingual language models (KMLMs) trained directly on the knowledge triples.We first generate a large amount of multilingual synthetic sentences using the Wikidata KG triples.Then based on the intra-and inter-sentence structures of the generated data, we design pretraining tasks to enable the LMs to not only memorize the factual knowledge but also learn useful logical patterns.Our pretrained KMLMs demonstrate significant performance improvements on a wide range of knowledge-intensive crosslingual tasks, including named entity recognition (NER), factual knowledge retrieval, relation classification, and a newly designed logical reasoning task. 1 Xin Li 0056, Ruidan He, Lidong Bing, Shafiq R. Joty, Luo Si |
EMNLP | 2 |
| 2022 | ConNER: Consistency Training for Cross-lingual Named Entity RecognitionabstractCross-lingual named entity recognition (NER) suffers from data scarcity in the target languages, especially under zero-shot settings.Existing translate-train or knowledge distillation methods attempt to bridge the language gap, but often introduce a high level of noise.To solve this problem, consistency training methods regularize the model to be robust towards perturbations on data or hidden states.However, such methods are likely to violate the consistency hypothesis, or mainly focus on coarse-grain consistency.We propose ConNER as a novel consistency training framework for cross-lingual NER, which comprises of: (1) translation-based consistency training on unlabeled target-language data, and (2) dropoutbased consistency training on labeled sourcelanguage data.ConNER effectively leverages unlabeled target-language data and alleviates overfitting on the source language to enhance the cross-lingual adaptability.Experimental results show our ConNER achieves consistent improvement over various baseline methods. 1 Ran Zhou 0004, Xin Li 0056, Lidong Bing, Erik Cambria, Luo Si, Chunyan Miao |
EMNLP | 2 |
| 2021 | Aspect Sentiment Quad Prediction as Paraphrase GenerationabstractAspect-based sentiment analysis (ABSA) has been extensively studied in recent years, which typically involves four fundamental sentiment elements, including the aspect category, aspect term, opinion term, and sentiment polarity.Existing studies usually consider the detection of partial sentiment elements, instead of predicting the four elements in one shot.In this work, we introduce the Aspect Sentiment Quad Prediction (ASQP) task, aiming to jointly detect all sentiment elements in quads for a given opinionated sentence, which can reveal a more comprehensive and complete aspect-level sentiment structure.We further propose a novel PARAPHRASE modeling paradigm to cast the ASQP task to a paraphrase generation process.On one hand, the generation formulation allows solving ASQP in an end-to-end manner, alleviating the potential error propagation in the pipeline solution.On the other hand, the semantics of the sentiment elements can be fully exploited by learning to generate them in the natural language form.Extensive experiments on benchmark datasets show the superiority of our proposed method and the capacity of crosstask transfer with the proposed unified PARA-PHRASE modeling framework. Wenxuan Zhang 0001, Yang Deng 0002, Xin Li 0056, Yifei Yuan 0002, Lidong Bing, Wai Lam |
EMNLP (1) | 3 |
| 2020 | Relevance-Promoting Language Model for Short-Text ConversationabstractDespite the effectiveness of sequence-to-sequence framework on the task of Short-Text Conversation (STC), the issue of under-exploitation of training data (i.e., the supervision signals from query text is ignored) still remains unresolved. Also, the adopted maximization-based decoding strategies, inclined to generating the generic responses or responses with repetition, are unsuited to the STC task. In this paper, we propose to formulate the STC task as a language modeling problem and tailor-make a training strategy to adapt a language model for response generation. To enhance generation performance, we design a relevance-promoting transformer language model, which performs additional supervised source attention after the self-attention to increase the importance of informative query tokens in calculating the token-level representation. The model further refines the query representation with relevance clues inferred from its multiple references during training. In testing, we adopt a randomization-over-maximization strategy to reduce the generation of generic responses. Experimental results on a large Chinese STC dataset demonstrate the superiority of the proposed model on relevance metrics and diversity metrics.1 Xin Li 0056, Piji Li, Wei Bi, Xiaojiang Liu, Wai Lam |
AAAI | 1 |
| 2019 | A Unified Model for Opinion Target Extraction and Target Sentiment PredictionabstractTarget-based sentiment analysis involves opinion target extraction and target sentiment classification. However, most of the existing works usually studied one of these two sub-tasks alone, which hinders their practical use. This paper aims to solve the complete task of target-based sentiment analysis in an end-to-end fashion, and presents a novel unified model which applies a unified tagging scheme. Our framework involves two stacked recurrent neural networks: The upper one predicts the unified tags to produce the final output results of the primary target-based sentiment analysis; The lower one performs an auxiliary target boundary prediction aiming at guiding the upper network to improve the performance of the primary task. To explore the inter-task dependency, we propose to explicitly model the constrained transitions from target boundaries to target sentiment polarities. We also propose to maintain the sentiment consistency within an opinion target via a gate mechanism which models the relation between the features for the current word and the previous word. We conduct extensive experiments on three benchmark datasets and our framework achieves consistently superior results. Xin Li 0056, Lidong Bing, Piji Li, Wai Lam |
AAAI | 1 |
| 2019 | Exploiting Coarse-to-Fine Task Transfer for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification (ASC) aims at identifying sentiment polarities towards aspects in a sentence, where the aspect can behave as a general Aspect Category (AC) or a specific Aspect Term (AT). However, due to the especially expensive and labor-intensive labeling, existing public corpora in AT-level are all relatively small. Meanwhile, most of the previous methods rely on complicated structures with given scarce data, which largely limits the efficacy of the neural models. In this paper, we exploit a new direction named coarse-to-fine task transfer, which aims to leverage knowledge learned from a rich-resource source domain of the coarse-grained AC task, which is more easily accessible, to improve the learning in a low-resource target domain of the fine-grained AT task. To resolve both the aspect granularity inconsistency and feature mismatch between domains, we propose a Multi-Granularity Alignment Network (MGAN). In MGAN, a novel Coarse2Fine attention guided by an auxiliary task can help the AC task modeling at the same finegrained level with the AT task. To alleviate the feature false alignment, a contrastive feature alignment method is adopted to align aspect-specific feature representations semantically. In addition, a large-scale multi-domain dataset for the AC task is provided. Empirically, extensive experiments demonstrate the effectiveness of the MGAN. Zheng Li 0018, Ying Wei 0001, Yu Zhang 0006, Xiang Zhang 0001, Xin Li 0056 |
AAAI | 5 |
| 2019 | Transferable End-to-End Aspect-based Sentiment Analysis with Selective Adversarial LearningabstractZheng Li, Xin Li, Ying Wei, Lidong Bing, Yu Zhang, Qiang Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zheng Li 0018, Xin Li 0056, Ying Wei 0001, Lidong Bing, Yu Zhang 0006, Qiang Yang 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Transformation Networks for Target-Oriented Sentiment ClassificationabstractTarget-oriented sentiment classification aims at classifying sentiment polarities over individual opinion targets in a sentence.RNN with attention seems a good fit for the characteristics of this task, and indeed it achieves the state-of-the-art performance.After re-examining the drawbacks of attention mechanism and the obstacles that block CNN to perform well in this classification task, we propose a new model to overcome these issues.Instead of attention, our model employs a CNN layer to extract salient features from the transformed word representations originated from a bi-directional RNN layer.Between the two layers, we propose a component to generate target-specific representations of words in the sentence, meanwhile incorporate a mechanism for preserving the original contextual information from the RNN layer.Experiments show that our model achieves a new state-of-the-art performance on a few benchmarks. 1 Xin Li 0056, Lidong Bing, Wai Lam, Bei Shi |
ACL (1) | 1 |
| 2018 | Aspect Term Extraction with History Attention and Selective TransformationabstractAspect Term Extraction (ATE), a key sub-task in Aspect-Based Sentiment Analysis, aims to extract explicit aspect expressions from online user reviews. We present a new framework for tackling ATE. It can exploit two useful clues, namely opinion summary and aspect detection history. Opinion summary is distilled from the whole input sentence, conditioned on each current token for aspect prediction, and thus the tailor-made summary can help aspect prediction on this token. On the other hand, the aspect detection history information is distilled from the previous aspect predictions, and it can leverage the coordinate structure and tagging schema constraints to upgrade the aspect prediction. Experimental results over four benchmark datasets clearly demonstrate that our framework can outperform all state-of-the-art methods. Xin Li 0056, Lidong Bing, Piji Li, Wai Lam, Zhimou Yang |
IJCAI | 1 |
| 2017 | Deep Multi-Task Learning for Aspect Term Extraction with Memory InteractionabstractWe propose a novel LSTM-based deep multi-task learning framework for aspect term extraction from user review sentences.Two LSTMs equipped with extended memories and neural memory operations are designed for jointly handling the extraction tasks of aspects and opinions via memory interactions.Sentimental sentence constraint is also added for more accurate prediction via another LSTM.Experiment results over two benchmark datasets demonstrate the effectiveness of our framework. Xin Li 0056, Wai Lam |
EMNLP | 1 |