VLDB 2026 Research / reviewers in the wild / expert
Bo Xu 0023
dblp:26/1194-23
· DBLP profile ↗
48ranked-venue papers
22as first author
39since 2021 · last 2026
0000-0002-2083-4307ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 15 first-author · 21 since 2021Databases, data management, data science and information retrieval · 15 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaptiveLog: An Adaptive Log Analysis Framework with the Collaboration of Large and Small Language ModelabstractAutomated log analysis is crucial to ensure the high availability and reliability of complex systems. The advent of Large Language Models (LLMs) in Natural Language Processing (NLP) has ushered in a new era of language model-driven automated log analysis, garnering significant interest. Within this field, two primary paradigms based on language models for log analysis have become prominent. Small Language Models (SLMs) (such as BERT) follow the pre-train and fine-tune paradigm, focusing on the specific log analysis task through fine-tuning on supervised datasets. On the other hand, LLMs (such as ChatGPT) following the in-context learning paradigm, analyze logs by providing a few examples in prompt contexts without updating parameters. Despite their respective strengths, both models exhibit inherent limitations. By comparing SLMs and LLMs, we notice that SLMs are more cost-effective but less powerful, whereas LLMs with large parameters are highly powerful but expensive and inefficient. To tradeoff between the performance and inference costs of both models in automated log analysis, this article introduces an adaptive log analysis framework known as AdaptiveLog, which effectively reduces the costs associated with LLM while ensuring superior results. This framework collaborates an LLM and an SLM, strategically allocating the LLM to tackle complex logs while delegating simpler logs to the SLM. Specifically, to efficiently query the LLM, we propose an adaptive selection strategy based on the uncertainty estimation of the SLM, where the LLM is invoked only when the SLM is uncertain. In addition, to enhance the reasoning ability of the LLM in log analysis tasks, we propose a novel prompt strategy by retrieving similar error-prone cases as the reference, enabling the model to leverage past error experiences and learn solutions from these cases. We evaluate AdaptiveLog on different log analysis tasks, Extensive experiments demonstrate that AdaptiveLog achieves state-of-the-art results across different tasks, elevating the overall accuracy of log analysis while maintaining cost efficiency. Our source code and detailed experimental data are available at https://github.com/LeaperOvO/AdaptiveLog-review . Lipeng Ma, Weidong Yang 0001, Ben Fei, Mingjie Zhou, Shuhao Li 0001, Sihang Jiang 0001, Bo Xu 0023, Yanghua Xiao |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2025 | GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionabstractLarge language models (LLMs) demonstrate impressive performance on downstream tasks through in-context learning(ICL). However, there is a significant gap between their performance in Named Entity Recognition (NER) and in fine-tuning methods. We believe this discrepancy is due to inconsistencies in labeling definitions in NER. In addition, recent research indicates that LLMs do not learn the specific input-label mappings from the demonstrations. Therefore, we argue that using examples to implicitly capture the mapping between inputs and labels in in-context learning is not suitable for NER. Instead, it requires explicitly informing the model of the range of entities contained in the labels, such as annotation guidelines. In this paper, we propose GuideNER, which uses LLMs to summarize concise annotation guidelines as contextual information in ICL. We have conducted experiments on widely used NER datasets, and the experimental results indicate that our method can consistently and significantly outperform state-of-the-art methods, while using shorter prompts. Especially on the GENIA dataset, our model outperforms the previous state-of-the-art model by 12.63 F1 scores. Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001 |
AAAI | 2 |
| 2025 | Gen-SQL: Efficient Text-to-SQL By Bridging Natural Language Question And Database Schema With Pseudo-SchemaabstractWith the prevalence of Large Language Models (LLMs), recent studies have shifted paradigms and leveraged LLMs to tackle the challenging task of Text-to-SQL. Because of the complexity of real world databases, previous works adopt the retrieve-then-generate framework to retrieve relevant database schema and then to generate the SQL query. However, efficient embedding-based retriever suffers from lower retrieval accuracy, and more accurate LLM-based retriever is far more expensive to use, which hinders their applicability for broader applications. To overcome this issue, this paper proposes Gen-SQL, a novel generate-ground-regenerate framework, where we exploit prior knowledge from the LLM to enhance embedding-based retriever and reduce cost. Experiments on several datasets are conducted to demonstrate the effectiveness and scalability of our proposed method. We release our code and data at https://github.com/jieshi10/gensql. Jie Shi 0010, Bo Xu 0023, Jiaqing Liang, Yanghua Xiao, Jia Chen 0037, Chenhao Xie 0002, Peng Wang 0027, Wei Wang 0009 |
COLING | 2 |
| 2025 | Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image AugmentationabstractMultimodal named entity recognition (MNER) extends traditional named entity recognition (NER) by integrating visual and textual information. However, current methods still face significant challenges due to the text-image mismatch problem. Recent advancements in text-to-image synthesis provide promising solutions, as synthesized images can introduce additional visual context to enhance MNER model performance. To fully leverage the benefits of both original and synthesized images, we propose an adaptive mixup image augmentation method. This method generates augmented images by determining the mixing ratio based on the matching score between the text and image, utilizing a triplet loss-based Gaussian Mixture Model (TL-GMM). Our approach is highly adaptable and can be seamlessly integrated into existing MNER models. Extensive experiments demonstrate consistent performance improvements, and detailed ablation studies and case studies confirm the effectiveness of our method. Bo Xu 0023, Haiqi Jiang 0004, Hongyu Jing, Ming Du 0002, Hongya Wang, Yanghua Xiao |
COLING | 1 |
| 2025 | Boosting Text-to-SQL through Multi-grained Error IdentificationabstractText-to-SQL is a technology that converts natural language questions into executable SQL queries, allowing users to query and manage relational databases more easily. In recent years, large language models have significantly advanced the development of text-to-SQL. However, existing methods often overlook validation of the generated results during the SQL generation process. Current error identification methods are mainly divided into self-correction approaches based on large models and feedback methods based on SQL execution, both of which have limitations. We categorize SQL errors into three main types: system errors, skeleton errors, and value errors, and propose a multi-grained error identification method. Experimental results demonstrate that this method can be integrated as a plugin into various methods, providing effective error identification and correction capabilities. Bo Xu 0023, Hongyu Jing, Ming Du 0002, Hongya Wang, Yanghua Xiao |
COLING | 1 |
| 2025 | Retrieval-Based Multimodal Data Augmentation for Multimodal Information Extraction in Social Media
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001 |
DASFAA (4) | 2 |
| 2025 | Skeletons Matter: Dynamic Data Augmentation for Text-to-QueryabstractThe task of translating natural language questions into query languages has long been a central focus in semantic parsing.Recent advancements in Large Language Models (LLMs) have significantly accelerated progress in this field.However, existing studies typically focus on a single query language, resulting in methods with limited generalizability across different languages.In this paper, we formally define the Text-to-Query task paradigm, unifying semantic parsing tasks across various query languages.We identify query skeletons as a shared optimization target of Text-to-Query tasks, and propose a general dynamic data augmentation framework that explicitly diagnoses modelspecific weaknesses in handling these skeletons to synthesize targeted training data.Experiments on four Text-to-Query benchmarks demonstrate that our method achieves state-ofthe-art performance using only a small amount of synthesized data, highlighting the efficiency and generality of our approach and laying a solid foundation for unified research on Textto-Query tasks.We release our code Yuchen Ji, Bo Xu 0023, Jie Shi 0010, Jiaqing Liang, Deqing Yang, Hai Chen, Yanghua Xiao |
EMNLP | 2 |
| 2025 | Dialect-SQL: An Adaptive Framework for Bridging the Dialect Gap in Text-to-SQLabstractText-to-SQL is the task of translating natural language questions into SQL queries based on relational databases.Different databases implement their own SQL dialects, leading to variations in syntax.As a result, SQL queries designed for one database may not execute properly in another, creating a dialect gap.Existing Text-to-SQL research primarily focuses on specific database systems, limiting adaptability to different dialects.This paper proposes a novel adaptive framework called Dialect-SQL, which employs Object Relational Mapping (ORM) code as an intermediate language to bridge this gap.Given a question, we guide Large Language Models (LLMs) to first generate ORM code, which is then parsed into SQL queries targeted for specific databases.However, there is a lack of high-quality Textto-Code datasets that enable LLMs to effectively generate ORM code.To address this issue, we propose a bootstrapping approach to synthesize ORM code, where verified ORM code is iteratively integrated into a demonstration pool that serves as in-context examples for ORM code generation.Our experiments demonstrate that Dialect-SQL significantly enhances dialect adaptability, outperforming traditional methods that generate SQL queries directly. Jie Shi 0010, Bo Xu 0023, Jiaqing Liang, Yanghua Xiao, Jia Chen 0037, Peng Wang 0027, Wei Wang 0009 |
EMNLP | 3 |
| 2025 | LogSI: A Benchmark for System-Incremental Log AnalysisabstractAutomated log analysis plays a vital role in software operations, with deep learning methods demonstrating effectiveness for analyzing logs from individual systems. However, existing methods face limitations in efficiency, adaptability, and knowledge preservation in system-incremental log analysis. Continual learning offers a solution by expanding the model’s ability to analyze logs from the increasing number of systems. For evaluating these methods in system-incremental log analysis, we introduce LogSI, a novel benchmark with four essential abilities for system-incremental log analysis. We perform a comprehensive evaluation of various baselines on LogSI, examining their robustness against different system permutations. Additionally, we conduct an in-depth study on the factors that influence their robustness. The datasets and source code of this paper can be found in https://github.com/nonauthor/LogSIbenchmark. Mingjie Zhou, Weidong Yang 0001, Lipeng Ma, Sihang Jiang 0001, Bo Xu 0023, Yanghua Xiao |
ICASSP | 5 |
| 2025 | Hierarchical Prompt Tuning for System-Incremental Log AnalysisabstractSystem-incremental log analysis, involves the ongoing training of a model using logs from diverse systems to enable effective resolution of log analysis tasks across an expanding array of systems. Existing continual learning methods, which are based on prompt tuning, have shown challenges in insufficient knowledge transfer and increasing catastrophic forgetting. To tackle these challenges, we present LogHPT, a novel continual learning method based on a hierarchical prompt tuning frame-work specifically tailored for system-incremental log analysis. LogHPT incorporates four types of prompt meticulously crafted to capture log knowledge across various granularities, thereby enhancing knowledge transfer. Subsequently, we employ a key-value mechanism to discern the most suitable prompts for the input logs. Additionally, we use general prompt learning based on knowledge distillation to mitigate catastrophic forgetting. To evaluate the performance of LogHPT, we conduct comprehensive experiments focusing on two fundamental subtasks: log parsing and log anomaly detection. The results show that LogHPT achieves state-of-the-art (SOTA) performance. The source code and datasets for this paper are accessible at the following link: https://github.com/nonauthor/LogHPT. Mingjie Zhou, Weidong Yang 0001, Lipeng Ma, Sihang Jiang 0001, Bo Xu 0023, Yanghua Xiao |
ICASSP | 5 |
| 2025 | Fast Locality Sensitive Hashing with Theoretical Guarantee
Zongyuan Tan, Hongya Wang, Bo Xu 0023, Minjie Luo, Ming Du 0002 |
ICCBR | 3 |
| 2025 | Low-Redundancy Knowledge Generation and Modality-Aware Interaction for Multimodal Information Extraction in Social MediaabstractMultimodal information extraction (MIE) has gained increasing attention, as it helps to accomplish information extraction by adding images as auxiliary information. By acquiring entity-related knowledge, knowledge generation methods can effectively enhance the performance of information extraction models. However, current knowledge generation methods have two weaknesses: (1) they often generate knowledge that includes task-irrelevant information causing redundancy and negatively impacting model performance; (2) they typically concatenate knowledge and text input directly together, ignoring the stylistic and contextual differences arising from their different sources. To address these issues, we propose Low-Redundancy Knowledge Generation and Modality-Aware Interaction (LRKG-MAI). Our approach leverages a large language model to generate task-relevant knowledge with minimal redundancy, while treating knowledge as a distinct modality that interacts with text within its own representation space. Extensive experiments demonstrate the effectiveness of our approach. The source code can be found at https://github.com/JinFish/LRKG-MAI. Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001 |
ICME | 2 |
| 2025 | Bridging the Unseen Gap: Label-Enhanced Information Bottleneck Distillation for Multimodal Named Entity RecognitionabstractMultimodal Named Entity Recognition (MNER) integrates visual information to resolve textual ambiguities but struggles with generalizing to unseen entities (out-of-vocabulary, OOV), particularly in social media. To bridge this gap, we leveraging internal label knowledge and visual information and propose a Label-Enhanced Information Bottleneck Distillation (LIBD) framework, which transfers label-aware generalization capabilities via a teacher-student architecture. Our method introduces Dual-level Label Augmentation (DLA), enhancing the teacher model by integrating word-level entity replacement with labels and embedding-level learnable label vectors. This is paired with Information Bottleneck Distillation (IBD), selectively distilling critical knowledge from the teacher while suppressing irrelevant noise. Experiments on benchmark datasets demonstrate that LIBD outperforms state-of-the-art methods, especially in identifying OOV entities. Bo Xu 0023, Hongya Wang, Ming Du 0002, Yanghua Xiao |
ACM Multimedia | 1 |
| 2025 | A Multi-expert Collaborative Framework for Multimodal Named Entity Recognition
Bo Xu 0023, Haiqi Jiang 0004, Shouang Wei, Ming Du 0002, Hongya Wang |
MMM (1) | 1 |
| 2025 | LUK: Empowering Log Understanding With Expert Knowledge From Large Language ModelsabstractLogs play a critical role in providing essential information for system monitoring and troubleshooting. Recently, with the success of pre-trained language models (PLMs) and large language models (LLMs) in natural language processing (NLP), smaller PLMs (such as BERT) and LLMs (like GPT-4) have become the current mainstream approaches for log analysis. Despite the remarkable capabilities of LLMs, their higher cost and inefficient inference present significant challenges in leveraging the full potential of LLMs to analyze logs. In contrast, smaller PLMs can be fine-tuned for specific tasks even with limited computational resources, making them more practical. However, these smaller PLMs face challenges in understanding logs comprehensively due to their limited expert knowledge. To address the lack of expert knowledge and enhance log understanding for smaller PLMs, this paper introduces a novel and practical knowledge enhancement framework, called LUK, which acquires expert knowledge from LLMs automatically and then enhances the smaller PLM for log analysis with the expert knowledge. LUK can take full advantage of both types of models. Specifically, we design a multi-expert collaboration framework based on LLMs with different roles to acquire expert knowledge. In addition, we propose two novel pre-training tasks to enhance the log pre-training with expert knowledge. LUK achieves state-of-the-art results on different log analysis tasks, and extensive experiments demonstrate that expert knowledge from LLMs can be utilized more effectively to understand logs. Our source code and detailed experimental data are available athttps://github.com/LeaperOvO/LUK. Lipeng Ma, Weidong Yang 0001, Sihang Jiang 0001, Ben Fei, Mingjie Zhou, Shuhao Li 0001, Bo Xu 0023, Yanghua Xiao |
IEEE Trans. Software Eng. | 8 |
| 2024 | Adaptive Reinforcement Tuning Language Models as Hard Data Generators for Sentence RepresentationabstractSentence representation learning is a fundamental task in NLP. Existing methods use contrastive learning (CL) to learn effective sentence representations, which benefit from high-quality contrastive data but require extensive human annotation. Large language models (LLMs) like ChatGPT and GPT4 can automatically generate such data. However, this alternative strategy also encounters challenges: 1) obtaining high-quality generated data from small-parameter LLMs is difficult, and 2) inefficient utilization of the generated data. To address these challenges, we propose a novel adaptive reinforcement tuning (ART) framework. Specifically, to address the first challenge, we introduce a reinforcement learning approach for fine-tuning small-parameter LLMs, enabling the generation of high-quality hard contrastive data without human feedback. To address the second challenge, we propose an adaptive iterative framework to guide the small-parameter LLMs to generate progressively harder samples through multiple iterations, thereby maximizing the utility of generated data. Experiments conducted on seven semantic text similarity tasks demonstrate that the sentence representation models trained using the synthetic data generated by our proposed method achieve state-of-the-art performance. Our code is available at https://github.com/WuNein/AdaptCL. Bo Xu 0023, Shouang Wei, Ming Du 0002, Hongya Wang |
LREC/COLING | 1 |
| 2024 | MNER-MI: A Multi-image Dataset for Multimodal Named Entity Recognition in Social MediaabstractRecently, multimodal named entity recognition (MNER) has emerged as a vital research area within named entity recognition. However, current MNER datasets and methods are predominantly based on text and a single accompanying image, leaving a significant research gap in MNER scenarios involving multiple images. To address the critical research gap and enhance the scope of MNER for real-world applications, we propose a novel human-annotated MNER dataset with multiple images called MNER-MI. Additionally, we construct a dataset named MNER-MI-Plus, derived from MNER-MI, to ensure its generality and applicability. Based on these datasets, we establish a comprehensive set of strong and representative baselines and we further propose a simple temporal prompt model with multiple images to address the new challenges in multi-image scenarios. We have conducted extensive experiments to demonstrate that considering multiple images provides a significant improvement over a single image and can offer substantial benefits for MNER. Furthermore, our proposed method achieves state-of-the-art results on both MNER-MI and MNER-MI-Plus, demonstrating its effectiveness. The datasets and source code can be found at https://github.com/JinFish/MNER-MI. Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001 |
LREC/COLING | 2 |
| 2024 | Few-Shot Log Analysis with Prompt-Based Multi-task Transfer Learning
Mingjie Zhou, Weidong Yang 0001, Lipeng Ma, Sihang Jiang 0001, Bo Xu 0023, Yanghua Xiao |
DASFAA (2) | 5 |
| 2024 | KnowLog: Knowledge Enhanced Pre-trained Language Model for Log UnderstandingabstractLogs as semi-structured text are rich in semantic information, making their comprehensive understanding crucial for automated log analysis. With the recent success of pre-trained language models in natural language processing, many studies have leveraged these models to understand logs. Despite their successes, existing pre-trained language models still suffer from three weaknesses. Firstly, these models fail to understand domain-specific terminology, especially abbreviations. Secondly, these models struggle to adequately capture the complete log context information. Thirdly, these models have difficulty in obtaining universal representations of different styles of the same logs. To address these challenges, we introduce KnowLog, a knowledge-enhanced pre-trained language model for log understanding. Specifically, to solve the previous two challenges, we exploit abbreviations and natural language descriptions of logs from public documentation as local and global knowledge, respectively, and leverage this knowledge by designing novel pre-training tasks for enhancing the model. To solve the last challenge, we design a contrastive learning-based pre-training task to obtain universal representations. We evaluate KnowLog by fine-tuning it on six different log understanding tasks. Extensive experiments demonstrate that KnowLog significantly enhances log understanding and achieves state-of-the-art results compared to existing pre-trained language models without knowledge enhancement. Moreover, we conduct additional experiments in transfer learning and low-resource scenarios, showcasing the substantial advantages of KnowLog. Our source code and detailed experimental data are available at https://github.com/LeaperOvO/KnowLog. Lipeng Ma, Weidong Yang 0001, Bo Xu 0023, Sihang Jiang 0001, Ben Fei, Jiaqing Liang, Mingjie Zhou, Yanghua Xiao |
ICSE | 3 |
| 2024 | Chain-of-Program Prompting with Open-Source Large Language Models for Text-to-SQLabstractText-to-SQL is a fundamental natural language processing (NLP) task that involves translating natural language queries related to a specified relational database into SQL queries. Recently, large language models (LLMs) have emerged as a crucial paradigm in the Text-to-SQL task. Despite their success, current methods heavily depend on closed-source LLMs with a large number of parameters, such as ChatGPT and GPT4, resulting in significant API costs and privacy concerns. Therefore, a more cost-effective strategy is to fine-tune open-source LLMs with smaller parameters for SQL generation. However, this alternative strategy faces challenges due to the weaker reasoning capabilities of open-source LLMs, particularly in generating complex SQL queries. To address this issue, we propose a Chain-of-Programs (COP) prompting framework for Text-to-SQL. Different from the conventional Chain-of-Thoughts (COT), we utilize Pandas code as an intermediate representation aligned with the step-wise nature of human thinking. This decomposition transforms complex SQL queries into a series of simple Pandas queries. Each step in the COP can be validated using a Python interpreter. Finally, we use the COP prompting to generate SQL queries. Experiments conducted on the Spider dataset using two open-source large language models have demonstrated that our performances are comparable to GPT4 in zero-shot scenario. Bo Xu 0023, Shouang Wei, Ming Du 0002, Hongya Wang |
IJCNN | 1 |
| 2024 | A Decomposition Framework for Class Incremental Text Classification with Dual Prompt TuningabstractText classification models need to be capable of learning new categories continually as new topics emerge over time. Class incremental text classification provides a solution that enables models to sequentially learn new classes while retaining previously acquired knowledge. However, existing methods have limitations such as high computational overhead, substantial storage requirements, and privacy concerns. To address these issues and leverage the powerful capabilities of prompt-based continual learning methods, we propose a decomposition framework for class incremental text classification with two key subtasks: task ID identification and continual text classification. For task identification, we generate synthetic samples using large language models, then train a sentence encoder with supervised contrastive learning. This allows task retrieval with minimal replay data and no privacy concerns. For classification, we introduce a novel dual prompt tuning approach. It employs a unified prompt decoupling strategy to capture both task-general and task-specific knowledge. We also propose a task-aware prompt initialization method utilizing relationships between tasks. Experiments on benchmark datasets demonstrate state-of-the-art performance. Our method proposed reduces reliance on replayed data and optimally leverages knowledge transfer. Bo Xu 0023, Ming Du 0002, Hongya Wang |
IJCNN | 1 |
| 2024 | A Sentimental Prompt Framework with Visual Text Encoder for Multimodal Sentiment AnalysisabstractRecently, multimodal sentiment analysis from social media posts has received increasing attention, as it can effectively improve single-modality-based sentiment analysis by leveraging the complementary information between text and images. Despite their success, current methods still suffer from two weaknesses: (1) the current methods for obtaining image representations do not obtain sentiment information, which leads to a significant gap between image representations and results; (2) the current methods ignore the sentiments expressed by the symbols (emoticons, emojis) in the text, but these symbols can effectively reflect the user's sentiments. To address these issues, we propose a sentimental prompt framework with visual text encoder (SPFVTE). Specifically, for the first problem, instead of using the image representation directly, we project the image representation as a prompt and utilize the prompt learning to capture sentimental information in images by learning a sentiment-specific prompt. For the second problem, considering that people get the meanings of emojis and emoticons from their graphics, we propose to render the text as an image and use a visual text encoder to capture the sentiments contained in emojis and emoticons. We have conducted experiments on three public multimodal sentiment datasets, and the experimental results show that our method can significantly and consistently outperform the state-of-the-art methods. The datasets and source code can be found at https://github.com/JinFish/SPFVTE. Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001 |
ICMR | 2 |
| 2023 | A Unified Visual Prompt Tuning Framework with Mixture-of-Experts for Multimodal Information Extraction
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Yanghua Xiao, Xin Lin 0001 |
DASFAA (3) | 1 |
| 2023 | Semi-supervised Learning for Fine-Grained Entity Typing with Mixed Label Smoothing and Pseudo Labeling
Bo Xu 0023, Zhengqi Zhang, Ming Du 0002, Hongya Wang, Yanghua Xiao |
DASFAA (3) | 1 |
| 2023 | Knowledge Graph Enhanced Sentential Relation Extraction via Dual Heterogeneous Graph Context SelectionabstractSentential relation extraction is a type of relation extraction task whose goal is to extract semantic relations between entities from a single sentence. Compared with other variants of relation extraction, it often suffers from limitations of semantic contextual information. Due to the presence of knowledge graphs, many approaches propose to augment the semantics of sentences with the knowledge of entities, thus improving the performance of relation extraction. Despite their success, existing methods still suffer from two weaknesses: (1) existing approaches aggregate sentences, entities and their attribute values into a heterogeneous information graph, but do not consider the types of edges; (2) existing methods dynamically select knowledge based only on the structural features of the graph, without considering the features of the nodes themselves. To address these two problems, we propose a dual heterogeneous graph context selection method for knowledge graph enhanced sentential relation extraction. Specifically, to solve the first problem, we employ an edge-aware graph convolutional network to learn the representations of the heterogeneous graph with considering the types of edges. To solve the second problem, we propose dual graph context selection to select the useful context by considering the graph structure and node feature representation together. Experiments conducted on the Wikidata-RE dataset demonstrate the effectiveness of the method. Bo Xu 0023, Luyi Cheng, Shizhou Huang, Shouang Wei, Ming Du 0002, Hongya Wang |
IJCNN | 1 |
| 2023 | HSimCSE: Improving Contrastive Learning of Unsupervised Sentence Representation with Adversarial Hard Positives and Dual Hard NegativesabstractRecently, contrastive learning (CL) has emerged as the fundamental framework for learning better sentence representations. In the unsupervised sentence representation task, due to the lack of labeled data, current CL-based approaches generally use various methods to generate or select positive and negative samples for the given sentence. Despite their success, existing CL-based unsupervised sentence representation methods underestimate hard positive samples and hard negative samples, which do not fully exploit the power of contrastive learning. In this paper, we argue that we need to focus more on hard positive and hard negative samples. To this end, we propose a novel contrastive learning model, HSimCSE, that extends SimCSE by considering both the hard positive and hard negative samples. Specifically, we first propose a novel adversarial positive sample generation module to generate an adversarial hard positive sample, then we propose a dual negative sample selection module to select hard negative samples from the in-batch samples and the entire training corpus. Finally, we propose a quadruplet loss to minimize the distance between the anchor sample and the adversarial hard positive sample and maximize the distance between the anchor sample and the two hard negative samples. Experiments conducted on seven semantic text similarity tasks demonstrate the effectiveness of our method. The source code can be found at https://github.com/xubodhu/HSimCSE. Bo Xu 0023, Shouang Wei, Luyi Cheng, Shizhou Huang, Ming Du 0002, Hongya Wang |
IJCNN | 1 |
| 2023 | ShellGPT: Generative Pre-trained Transformer Model for Shell Language UnderstandingabstractThis paper presents ShellGPT, a pre-trained language model specifically designed to enhance the understanding of shell language which plays a crucial role in IT operations. Based on the GPT series of models, ShellGPT is trained on a corpus that aligns shell language with natural language, aiming to inject domain-specific knowledge into the model. The technique of pre-tokenization is employed to maximize the reuse of a general-purpose vocabulary, facilitating effective model transfer from general domain. Furthermore, a new pre-training objective, named equivalent command learning, is proposed to refine the command representations through modeling function equivalence of commands. To evaluate the performance of ShellGPT, we conduct fine-tuning on various downstream tasks related to shell language understanding. These tasks include command recommendation, command correction, and translation from natural language to shell command. Our experimental results demonstrate that ShellGPT outperforms other baseline models in terms of performance across almost all evaluated tasks. The findings from our experiments highlight the effectiveness of Shell-GPT in enhancing shell language understanding and demonstrate its potential for practical applications in IT operations. Jie Shi 0010, Sihang Jiang 0001, Bo Xu 0023, Jiaqing Liang, Yanghua Xiao, Wei Wang 0009 |
ISSRE | 3 |
| 2023 | ServerRCA: Root Cause Analysis for Server Failure using Operating System LogsabstractThe development of the information technology industry has made servers an essential infrastructure for enterprises. Server failure may result in significant economic losses. Therefore, it is essential to conduct root cause analysis (RCA) on server failure to improve server reliability. However, existing RCA approaches suffer from limitations in analysis granularity, adaptation difficulties, and data acquisition constraints. To overcome the limitations, we propose ServerRCA, an automated solution that utilizes operating system (OS) logs for accurate and efficient root cause analysis of server failures. OS logs provide detailed information and are easily accessible. Firstly, ServerRCA employs log parsing to transform raw logs into log templates. Next, we propose a hierarchical matching approach that leverages the hierarchical structure of fault logs to accurately identify fault events. Furthermore, we also introduce a human-in-the-loop feedback mechanism to enhance the ability of ServerRCA. Finally, ServerRCA constructs the fault propagation chain using the fault events identified earlier. Extensive experiments on real server failures demonstrate the effectiveness of ServerRCA, achieving significant improvements in F1-score, HR@1, and HR@3 over comparative methods. Our work contributes to the automated RCA of server failures using OS logs and provides a novel framework for accurate fault event identification in server failure analysis. Sihang Jiang 0001, Bo Xu 0023, Yanghua Xiao |
ISSRE | 3 |
| 2023 | Dialogue State Tracking with a Dialogue-Aware Slot-Level Schema Graph Approach
Bo Xu 0023 |
KSEM (3) | 3 |
| 2022 | Different Data, Different Modalities! Reinforced Data Splitting for Effective Multimodal Information Extraction from Social Media PostsabstractRecently, multimodal information extraction from social media posts has gained increasing attention in the natural language processing community. Despite their success, current approaches overestimate the significance of images. In this paper, we argue that different social media posts should consider different modalities for multimodal information extraction. Multimodal models cannot always outperform unimodal models. Some posts are more suitable for the multimodal model, while others are more suitable for the unimodal model. Therefore, we propose a general data splitting strategy to divide the social media posts into two sets so that these two sets can achieve better performance under the information extraction models of the corresponding modalities. Specifically, for an information extraction task, we first propose a data discriminator that divides social media posts into a multimodal and a unimodal set. Then we feed these sets into the corresponding models. Finally, we combine the results of these two models to obtain the final extraction results. Due to the lack of explicit knowledge, we use reinforcement learning to train the data discriminator. Experiments on two different multimodal information extraction tasks demonstrate the effectiveness of our method. The source code of this paper can be found in https://github.com/xubodhu/RDS. Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Chaofeng Sha, Yanghua Xiao |
COLING | 1 |
| 2022 | A Three-Stage Curriculum Learning Framework with Hierarchical Label Smoothing for Fine-Grained Entity Typing
Bo Xu 0023, Zhengqi Zhang, Chaofeng Sha, Ming Du 0002, Hongya Wang |
DASFAA (3) | 1 |
| 2022 | TS-DST: A Two-Stage Framework for Schema-Guided Dialogue State Tracking with Selected Dialogue HistoryabstractThe task-oriented dialogue systems aim to assist the users in completing specific tasks through natural language dialogue. Recently, word-level dialogue state tracking (DST) has become a core component of task-oriented dialogue systems. In this paper, we study the word-level DST task at the 8th dialogue system technology challenge (DSTC8), namely schema-guided dialogue state tracking, which focuses on cross-domain dialogue state tracking and zero-shot generalization to new services. Many approaches have been proposed to exploit the schema description for dialogue modeling, especially on unseen services. Despite their success, existing methods still suffer from two weaknesses: (1) the current methods do not fully exploit the dialogue history, which makes it difficult to solve the slot carryover problem from the multi-domain dialogues; (2) the current method treats the task as four independent sub tasks without considering the relevance of the subtasks. To address these issues, we propose a novel two-stage framework for schema-guided dialogue state tracking with selected dialogue history (TS-DST). Specifically, to solve the first issue, we propose a novel utterance selection module to select the most related previous utterances from the dialogue history by considering the specific schema element. To solve the second issue, we propose a two-stage framework to solve the four subtasks. Experiments conducted on the SGD dataset show that our method achieves new state-of-the-art performance. We also conduct ablation studies to demonstrate the effectiveness of the utterance selection module and the two-stage strategy. Ming Du 0002, Luyi Cheng, Bo Xu 0023, Sufen Wang, Junyi Yuan, Changqing Pan |
IJCNN | 3 |
| 2022 | Knowledge Base Entity Typing From Text via Entity-Aware Heterogeneous Graph Attention NetworkabstractKnowledge base entity typing from the description text (KBET-X) has become an important research direction, which takes semantically richer descriptive text as input to obtain better typing results. However, existing approaches either consider all sentences in the text but ignore the entities themselves or consider only the sentences in which the entities are mentioned without considering the other sentences in the text. To address these issues, we propose a novel framework for KBET-X based on an entity-aware heterogeneous graph attention network that makes full use of all sentences in the description text and considers the entities themselves. Specifically, we construct a heterogeneous graph for the description text with the node being a word or sentence. Node embeddings are initialized with an entity-aware encoder. Then we use a context encoder to obtain a contextual node representation of each word and sentence, consisting of a heterogeneous graph attention network and a gated recurrent unit (GRU) network. Finally, we use a type decoder based on a multilayer perceptron (MLP) network to obtain the types of each entity. Experiments conducted on the DBpedia dataset show that our method achieves the new state-of-the-art performance. We also conduct an ablation study to demonstrate that each component plays an essential role in our framework. Bo Xu 0023, Zhong Sun, Ming Du 0002, Hongya Wang |
IJCNN | 1 |
| 2022 | Revisiting Performance Measures for Cross-Modal HashingabstractRecently, cross-modal hashing has attracted much attention due to its low storage cost and fast query speed. Mean Average Precision (MAP) is the most widely used performance measure for cross-modal hashing. However, we found that the MAP scores do not fully reflect the quality of the top-K results for cross-modal retrieval because it neglects multi-label information and overlooks the label semantic hierarchy. In view of this, we propose a new performance measure named Normalized Weighted Discounted Cumulative Gains (NWDCG) by extending Normalized Discounted Cumulative Gains (NDCG) using co-occurrence probability matrix. To verify the effectiveness of NWDCG, we conduct extensive experiments using three popular cross-modal hashing schemes over two publically available datasets. Hongya Wang, Shunxin Dai, Ming Du 0002, Bo Xu 0023, Mingyong Li |
ICMR | 4 |
| 2022 | MAF: A General Matching and Alignment Framework for Multimodal Named Entity RecognitionabstractIn this paper, we study multimodal named entity recognition in social media posts. Existing works mainly focus on using a cross-modal attention mechanism to combine text representation with image representation. However, they still suffer from two weaknesses: (1) the current methods are based on a strong assumption that each text and its accompanying image are matched, and the image can be used to help identify named entities in the text. However, this assumption is not always true in real scenarios, and the strong assumption may reduce the recognition effect of theMNER model; (2) the current methods fail to construct a consistent representation to bridge the semantic gap between two modalities, which prevents the model from establishing a good connection between the text and image. To address these issues, we propose a general matching and alignment framework (MAF) for multimodal named entity recognition in social media posts. Specifically, to solve the first issue, we propose a novel cross-modal matching (CM) module to calculate the similarity score between text and image, and use the score to determine the proportion of visual information that should be retained. To solve the second issue, we propose a novel cross-modal alignment (CA) module to make the representations of the two modalities more consistent. We conduct extensive experiments, ablation studies, and case studies to demonstrate the effectiveness and efficiency of our method.The source code of this paper can be found in https://github.com/xubodhu/MAF. Bo Xu 0023, Shizhou Huang, Chaofeng Sha, Hongya Wang |
WSDM | 1 |
| 2022 | Micro-behaviour with Reinforcement Knowledge-aware Reasoning for Explainable Recommendation
Shaohua Tao, Runhe Qiu, Bo Xu 0023, Yuan Ping 0003 |
Knowl. Based Syst. | 3 |
| 2021 | Joint Entity and Relation Extraction for Long Text
Xianglong He, Bo Xu 0023 |
KSEM | 4 |
| 2021 | Reinforced Natural Language Inference for Distantly Supervised Relation Classification
Bo Xu 0023, Xiangsan Zhao, Chaofeng Sha, Minjun Zhang |
PAKDD (3) | 1 |
| 2021 | Improving Sentence-Level Relation Classification via Machine Reading Comprehension and Reinforcement Learning
Bo Xu 0023, Zhengqi Zhang, Xiangsan Zhao, Ming Du 0002 |
PRICAI (2) | 1 |
| 2020 | Mining Verb-Oriented Commonsense KnowledgeabstractCommonsense knowledge acquisition is one of the fundamental issues in the implementation of human-level AI. However, commonsense is difficult to obtain, because it is a human consensus and rarely explicitly appears in texts or other data. In this paper, we focus on the automatic acquisition of a typical kind of implicit verb-oriented commonsense knowledge (e.g., "person eats food"), which is the concept level knowledge of verb phrases. For this purpose, we propose a knowledge-driven approach to mine verb-oriented commonsense knowledge from verb phrases with the help of taxonomy. First, we design an entropy-based filter to cope with noisy input verb phrases. Then, we propose a joint model based on minimum description length and a neural language model to generate verb-oriented common-sense knowledge. We conduct extensive experiments to show that our solution is more effective to mine verb-oriented commonsense knowledge than competitors, and finally, we harvest 18K verb-oriented commonsense knowledge. Yuanfu Zhou, Chao Wang 0095, Haiyun Jiang, Sheng Zhang 0027, Bo Xu 0023, Yanghua Xiao |
ICDE | 7 |
| 2020 | Using Active Learning to Improve Distantly Supervised Entity Typing in Multi-source Knowledge Bases
Bo Xu 0023, Xiangsan Zhao, Qingxuan Kong |
NLPCC (1) | 1 |
| 2019 | A Community-Based Collaborative Filtering Method for Social Recommender SystemsabstractRecommender systems have become indispensable for recommending items of interest to users and have been successfully deployed in a wide range of real-world applications. In this paper, we exploit the community influence of users in social networks to improve the recommendation accuracy. Depending on the community in which the user is located, the user's preferences are defined more accurately. Specifically, we propose a community-based collaborative filtering method for social recommender systems, which makes full use of the rich link/community structure within a social network. We first group users in a social network into overlapping communities. Then we explicitly incorporate the community preference into the latent factor model. Compared with seven state-of-the-art methods on four real-world datasets, our method achieves the best performance. Bo Xu 0023, Deqing Yang, Yanghua Xiao, Wei Wang 0009 |
ICWS | 2 |
| 2019 | Feature-Level Attention Based Sentence Encoding for Neural Relation Extraction
Longqi Dai, Bo Xu 0023 |
NLPCC (1) | 2 |
| 2018 | METIC: Multi-Instance Entity Typing from CorpusabstractThis paper addresses the problem ofmulti-instance entity typing from corpus. Current approaches mainly rely on the structured features (\textitattributes, attribute-value pairs andtags ) of the entities. However, their effectiveness is largely dependent on the completeness of structured features, which unfortunately is not guaranteed in KBs. In this paper, we therefore propose to use the text corpus of an entity to infer its types, and propose a multi-instance method to tackle this problem. We take each mention of an entity in KBs as an instance of the entity, and learn the types of these entities from multiple instances. Specifically, we first use an end-to-end neural network model to type each instance of an entity, and then use an integer linear programming (ILP) method to aggregate the predicted type results from multiple instances. Experimental results show the effectiveness of our method. Bo Xu 0023, Luyang Huang, Yanghua Xiao, Deqing Yang, Wei Wang 0009 |
CIKM | 1 |
| 2017 | CN-DBpedia: A Never-Ending Chinese Knowledge Extraction System
Bo Xu 0023, Jiaqing Liang, Chenhao Xie 0002, Wanyun Cui, Yanghua Xiao |
IEA/AIE (2) | 1 |
| 2016 | Cross-Lingual Type Inference
Bo Xu 0023, Jiaqing Liang, Yanghua Xiao, Seung-won Hwang, Wei Wang 0009 |
DASFAA (1) | 1 |
| 2016 | Learning Defining Features for Categories
Bo Xu 0023, Chenhao Xie 0002, Yanghua Xiao, Haixun Wang, Wei Wang 0009 |
IJCAI | 1 |
| 2012 | Which Topic Will You Follow?
Deqing Yang, Yanghua Xiao, Bo Xu 0023, Hanghang Tong, Wei Wang 0009 |
ECML/PKDD (2) | 3 |