Ming Du 0002

dblp:78/105-2 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
22since 2021 · last 2025
0000-0002-8617-8939ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation
abstract
Multimodal named entity recognition (MNER) extends traditional named entity recognition (NER) by integrating visual and textual information. However, current methods still face significant challenges due to the text-image mismatch problem. Recent advancements in text-to-image synthesis provide promising solutions, as synthesized images can introduce additional visual context to enhance MNER model performance. To fully leverage the benefits of both original and synthesized images, we propose an adaptive mixup image augmentation method. This method generates augmented images by determining the mixing ratio based on the matching score between the text and image, utilizing a triplet loss-based Gaussian Mixture Model (TL-GMM). Our approach is highly adaptable and can be seamlessly integrated into existing MNER models. Extensive experiments demonstrate consistent performance improvements, and detailed ablation studies and case studies confirm the effectiveness of our method.
Bo Xu 0023, Haiqi Jiang 0004, Hongyu Jing, Ming Du 0002, Hongya Wang, Yanghua Xiao
COLING5
2025 Boosting Text-to-SQL through Multi-grained Error Identification
abstract
Text-to-SQL is a technology that converts natural language questions into executable SQL queries, allowing users to query and manage relational databases more easily. In recent years, large language models have significantly advanced the development of text-to-SQL. However, existing methods often overlook validation of the generated results during the SQL generation process. Current error identification methods are mainly divided into self-correction approaches based on large models and feedback methods based on SQL execution, both of which have limitations. We categorize SQL errors into three main types: system errors, skeleton errors, and value errors, and propose a multi-grained error identification method. Experimental results demonstrate that this method can be integrated as a plugin into various methods, providing effective error identification and correction capabilities.
Bo Xu 0023, Hongyu Jing, Ming Du 0002, Hongya Wang, Yanghua Xiao
COLING4
2025 Fast Locality Sensitive Hashing with Theoretical Guarantee
Zongyuan Tan, Hongya Wang, Bo Xu 0023, Minjie Luo, Ming Du 0002
ICCBR5
2025 Bridging the Unseen Gap: Label-Enhanced Information Bottleneck Distillation for Multimodal Named Entity Recognition
abstract
Multimodal Named Entity Recognition (MNER) integrates visual information to resolve textual ambiguities but struggles with generalizing to unseen entities (out-of-vocabulary, OOV), particularly in social media. To bridge this gap, we leveraging internal label knowledge and visual information and propose a Label-Enhanced Information Bottleneck Distillation (LIBD) framework, which transfers label-aware generalization capabilities via a teacher-student architecture. Our method introduces Dual-level Label Augmentation (DLA), enhancing the teacher model by integrating word-level entity replacement with labels and embedding-level learnable label vectors. This is paired with Information Bottleneck Distillation (IBD), selectively distilling critical knowledge from the teacher while suppressing irrelevant noise. Experiments on benchmark datasets demonstrate that LIBD outperforms state-of-the-art methods, especially in identifying OOV entities.
Bo Xu 0023, Hongya Wang, Ming Du 0002, Yanghua Xiao
ACM Multimedia4
2025 A Multi-expert Collaborative Framework for Multimodal Named Entity Recognition
Bo Xu 0023, Haiqi Jiang 0004, Shouang Wei, Ming Du 0002, Hongya Wang
MMM (1)4
2024 Adaptive Reinforcement Tuning Language Models as Hard Data Generators for Sentence Representation
abstract
Sentence representation learning is a fundamental task in NLP. Existing methods use contrastive learning (CL) to learn effective sentence representations, which benefit from high-quality contrastive data but require extensive human annotation. Large language models (LLMs) like ChatGPT and GPT4 can automatically generate such data. However, this alternative strategy also encounters challenges: 1) obtaining high-quality generated data from small-parameter LLMs is difficult, and 2) inefficient utilization of the generated data. To address these challenges, we propose a novel adaptive reinforcement tuning (ART) framework. Specifically, to address the first challenge, we introduce a reinforcement learning approach for fine-tuning small-parameter LLMs, enabling the generation of high-quality hard contrastive data without human feedback. To address the second challenge, we propose an adaptive iterative framework to guide the small-parameter LLMs to generate progressively harder samples through multiple iterations, thereby maximizing the utility of generated data. Experiments conducted on seven semantic text similarity tasks demonstrate that the sentence representation models trained using the synthetic data generated by our proposed method achieve state-of-the-art performance. Our code is available at https://github.com/WuNein/AdaptCL.
Bo Xu 0023, Shouang Wei, Ming Du 0002, Hongya Wang
LREC/COLING4
2024 Chain-of-Program Prompting with Open-Source Large Language Models for Text-to-SQL
abstract
Text-to-SQL is a fundamental natural language processing (NLP) task that involves translating natural language queries related to a specified relational database into SQL queries. Recently, large language models (LLMs) have emerged as a crucial paradigm in the Text-to-SQL task. Despite their success, current methods heavily depend on closed-source LLMs with a large number of parameters, such as ChatGPT and GPT4, resulting in significant API costs and privacy concerns. Therefore, a more cost-effective strategy is to fine-tune open-source LLMs with smaller parameters for SQL generation. However, this alternative strategy faces challenges due to the weaker reasoning capabilities of open-source LLMs, particularly in generating complex SQL queries. To address this issue, we propose a Chain-of-Programs (COP) prompting framework for Text-to-SQL. Different from the conventional Chain-of-Thoughts (COT), we utilize Pandas code as an intermediate representation aligned with the step-wise nature of human thinking. This decomposition transforms complex SQL queries into a series of simple Pandas queries. Each step in the COP can be validated using a Python interpreter. Finally, we use the COP prompting to generate SQL queries. Experiments conducted on the Spider dataset using two open-source large language models have demonstrated that our performances are comparable to GPT4 in zero-shot scenario.
Bo Xu 0023, Shouang Wei, Ming Du 0002, Hongya Wang
IJCNN5
2024 A Decomposition Framework for Class Incremental Text Classification with Dual Prompt Tuning
abstract
Text classification models need to be capable of learning new categories continually as new topics emerge over time. Class incremental text classification provides a solution that enables models to sequentially learn new classes while retaining previously acquired knowledge. However, existing methods have limitations such as high computational overhead, substantial storage requirements, and privacy concerns. To address these issues and leverage the powerful capabilities of prompt-based continual learning methods, we propose a decomposition framework for class incremental text classification with two key subtasks: task ID identification and continual text classification. For task identification, we generate synthetic samples using large language models, then train a sentence encoder with supervised contrastive learning. This allows task retrieval with minimal replay data and no privacy concerns. For classification, we introduce a novel dual prompt tuning approach. It employs a unified prompt decoupling strategy to capture both task-general and task-specific knowledge. We also propose a task-aware prompt initialization method utilizing relationships between tasks. Experiments on benchmark datasets demonstrate state-of-the-art performance. Our method proposed reduces reliance on replayed data and optimally leverages knowledge transfer.
Bo Xu 0023, Ming Du 0002, Hongya Wang
IJCNN5
2023 A Unified Visual Prompt Tuning Framework with Mixture-of-Experts for Multimodal Information Extraction
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Yanghua Xiao, Xin Lin 0001
DASFAA (3)3
2023 Semi-supervised Learning for Fine-Grained Entity Typing with Mixed Label Smoothing and Pseudo Labeling
Bo Xu 0023, Zhengqi Zhang, Ming Du 0002, Hongya Wang, Yanghua Xiao
DASFAA (3)3
2023 Knowledge Graph Enhanced Sentential Relation Extraction via Dual Heterogeneous Graph Context Selection
abstract
Sentential relation extraction is a type of relation extraction task whose goal is to extract semantic relations between entities from a single sentence. Compared with other variants of relation extraction, it often suffers from limitations of semantic contextual information. Due to the presence of knowledge graphs, many approaches propose to augment the semantics of sentences with the knowledge of entities, thus improving the performance of relation extraction. Despite their success, existing methods still suffer from two weaknesses: (1) existing approaches aggregate sentences, entities and their attribute values into a heterogeneous information graph, but do not consider the types of edges; (2) existing methods dynamically select knowledge based only on the structural features of the graph, without considering the features of the nodes themselves. To address these two problems, we propose a dual heterogeneous graph context selection method for knowledge graph enhanced sentential relation extraction. Specifically, to solve the first problem, we employ an edge-aware graph convolutional network to learn the representations of the heterogeneous graph with considering the types of edges. To solve the second problem, we propose dual graph context selection to select the useful context by considering the graph structure and node feature representation together. Experiments conducted on the Wikidata-RE dataset demonstrate the effectiveness of the method.
Bo Xu 0023, Luyi Cheng, Shizhou Huang, Shouang Wei, Ming Du 0002, Hongya Wang
IJCNN6
2023 HSimCSE: Improving Contrastive Learning of Unsupervised Sentence Representation with Adversarial Hard Positives and Dual Hard Negatives
abstract
Recently, contrastive learning (CL) has emerged as the fundamental framework for learning better sentence representations. In the unsupervised sentence representation task, due to the lack of labeled data, current CL-based approaches generally use various methods to generate or select positive and negative samples for the given sentence. Despite their success, existing CL-based unsupervised sentence representation methods underestimate hard positive samples and hard negative samples, which do not fully exploit the power of contrastive learning. In this paper, we argue that we need to focus more on hard positive and hard negative samples. To this end, we propose a novel contrastive learning model, HSimCSE, that extends SimCSE by considering both the hard positive and hard negative samples. Specifically, we first propose a novel adversarial positive sample generation module to generate an adversarial hard positive sample, then we propose a dual negative sample selection module to select hard negative samples from the in-batch samples and the entire training corpus. Finally, we propose a quadruplet loss to minimize the distance between the anchor sample and the adversarial hard positive sample and maximize the distance between the anchor sample and the two hard negative samples. Experiments conducted on seven semantic text similarity tasks demonstrate the effectiveness of our method. The source code can be found at https://github.com/xubodhu/HSimCSE.
Bo Xu 0023, Shouang Wei, Luyi Cheng, Shizhou Huang, Ming Du 0002, Hongya Wang
IJCNN6
2023 An efficient indexing technique for billion-scale nearest neighbor search
Hongya Wang, Ming Du 0002, Zhizheng Wang, Zongyuan Tan, Jie Zhang 0077, Yingyuan Xiao
Multim. Tools Appl.3
2023 Fast Reachability Queries Answering Based on $\mathsf{RCN}$RCN Reduction
abstract
Answering reachability queries is a fundamental graph operation. Considering that the size of the input graph has a great impact on query performance, there are studies focusing on reducing the graph size, such that queries can be answered over a smaller graph. Although the input graph can be compressed significantly by existing approaches, a good compression ratio does not always mean a positive effect on query performance. In this paper, we study graph reduction to accelerate reachability queries answering. We propose a novel graph reduction approach, namely RCN reduction, to compress the input graph into a smaller one. Let a be the compression ratio of the number of nodes in the reduced graph over that of the input graph, we show that based on our approach, the lower bound probability that a query q can be answered in constant time is 1-a^2. We show the difficulties of RCN reduction and propose efficient algorithms to improve the compression ratio. Based on the result of RCN reduction, we further propose a novel labeling scheme to accelerate queries answering. We confirm the efficiency of our approach by extensive experimental results for graph reduction and reachability queries processing using 20 real datasets.
Junfeng Zhou, Jeffrey Xu Yu, Yaxian Qiu, Xian Tang, Ming Du 0002
IEEE Trans. Knowl. Data Eng.6
2022 Different Data, Different Modalities! Reinforced Data Splitting for Effective Multimodal Information Extraction from Social Media Posts
abstract
Recently, multimodal information extraction from social media posts has gained increasing attention in the natural language processing community. Despite their success, current approaches overestimate the significance of images. In this paper, we argue that different social media posts should consider different modalities for multimodal information extraction. Multimodal models cannot always outperform unimodal models. Some posts are more suitable for the multimodal model, while others are more suitable for the unimodal model. Therefore, we propose a general data splitting strategy to divide the social media posts into two sets so that these two sets can achieve better performance under the information extraction models of the corresponding modalities. Specifically, for an information extraction task, we first propose a data discriminator that divides social media posts into a multimodal and a unimodal set. Then we feed these sets into the corresponding models. Finally, we combine the results of these two models to obtain the final extraction results. Due to the lack of explicit knowledge, we use reinforcement learning to train the data discriminator. Experiments on two different multimodal information extraction tasks demonstrate the effectiveness of our method. The source code of this paper can be found in https://github.com/xubodhu/RDS.
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Chaofeng Sha, Yanghua Xiao
COLING3
2022 A Three-Stage Curriculum Learning Framework with Hierarchical Label Smoothing for Fine-Grained Entity Typing
Bo Xu 0023, Zhengqi Zhang, Chaofeng Sha, Ming Du 0002, Hongya Wang
DASFAA (3)4
2022 Fast Reachability Queries Answering based on RCN Reduction (Extended abstract)
abstract
We study graph reduction to accelerate reachability queries answering. We propose a novel graph reduction approach, namely RCN reduction, to reduce the input graph$G$of$\vert V\vert$nodes into a smaller one with$\vert V^{r}\vert$nodes. Assume that the probability of a node of$G$to be a query node is$1/\vert V\vert$, we show that based on our approach, the lower bound probability that a query$q$can be answered in constant time is$1-(\frac{\vert V^{r}\vert}{\vert V\vert})^{2}$, denoting that the smaller the reduced graph, the larger the probability that$q$can be answered in constant time. We show the difficulties of RCN reduction and propose efficient algorithms to improve the reduction ratio. We confirm the benefits of our approach by rich experimental results using real datasets.
Junfeng Zhou, Jeffrey Xu Yu, Yaxian Qiu, Xian Tang, Ming Du 0002
ICDE6
2022 TS-DST: A Two-Stage Framework for Schema-Guided Dialogue State Tracking with Selected Dialogue History
abstract
The task-oriented dialogue systems aim to assist the users in completing specific tasks through natural language dialogue. Recently, word-level dialogue state tracking (DST) has become a core component of task-oriented dialogue systems. In this paper, we study the word-level DST task at the 8th dialogue system technology challenge (DSTC8), namely schema-guided dialogue state tracking, which focuses on cross-domain dialogue state tracking and zero-shot generalization to new services. Many approaches have been proposed to exploit the schema description for dialogue modeling, especially on unseen services. Despite their success, existing methods still suffer from two weaknesses: (1) the current methods do not fully exploit the dialogue history, which makes it difficult to solve the slot carryover problem from the multi-domain dialogues; (2) the current method treats the task as four independent sub tasks without considering the relevance of the subtasks. To address these issues, we propose a novel two-stage framework for schema-guided dialogue state tracking with selected dialogue history (TS-DST). Specifically, to solve the first issue, we propose a novel utterance selection module to select the most related previous utterances from the dialogue history by considering the specific schema element. To solve the second issue, we propose a two-stage framework to solve the four subtasks. Experiments conducted on the SGD dataset show that our method achieves new state-of-the-art performance. We also conduct ablation studies to demonstrate the effectiveness of the utterance selection module and the two-stage strategy.
Ming Du 0002, Luyi Cheng, Bo Xu 0023, Sufen Wang, Junyi Yuan, Changqing Pan
IJCNN1
2022 Knowledge Base Entity Typing From Text via Entity-Aware Heterogeneous Graph Attention Network
abstract
Knowledge base entity typing from the description text (KBET-X) has become an important research direction, which takes semantically richer descriptive text as input to obtain better typing results. However, existing approaches either consider all sentences in the text but ignore the entities themselves or consider only the sentences in which the entities are mentioned without considering the other sentences in the text. To address these issues, we propose a novel framework for KBET-X based on an entity-aware heterogeneous graph attention network that makes full use of all sentences in the description text and considers the entities themselves. Specifically, we construct a heterogeneous graph for the description text with the node being a word or sentence. Node embeddings are initialized with an entity-aware encoder. Then we use a context encoder to obtain a contextual node representation of each word and sentence, consisting of a heterogeneous graph attention network and a gated recurrent unit (GRU) network. Finally, we use a type decoder based on a multilayer perceptron (MLP) network to obtain the types of each entity. Experiments conducted on the DBpedia dataset show that our method achieves the new state-of-the-art performance. We also conduct an ablation study to demonstrate that each component plays an essential role in our framework.
Bo Xu 0023, Zhong Sun, Ming Du 0002, Hongya Wang
IJCNN3
2022 Revisiting Performance Measures for Cross-Modal Hashing
abstract
Recently, cross-modal hashing has attracted much attention due to its low storage cost and fast query speed. Mean Average Precision (MAP) is the most widely used performance measure for cross-modal hashing. However, we found that the MAP scores do not fully reflect the quality of the top-K results for cross-modal retrieval because it neglects multi-label information and overlooks the label semantic hierarchy. In view of this, we propose a new performance measure named Normalized Weighted Discounted Cumulative Gains (NWDCG) by extending Normalized Discounted Cumulative Gains (NDCG) using co-occurrence probability matrix. To verify the effectiveness of NWDCG, we conduct extensive experiments using three popular cross-modal hashing schemes over two publically available datasets.
Hongya Wang, Shunxin Dai, Ming Du 0002, Bo Xu 0023, Mingyong Li
ICMR3
2022 Index-based top kα-maximal-clique enumeration over uncertain graphs
Jing Bai 0011, Junfeng Zhou, Ming Du 0002
J. Supercomput.3
2021 Improving Sentence-Level Relation Classification via Machine Reading Comprehension and Reinforcement Learning
Bo Xu 0023, Zhengqi Zhang, Xiangsan Zhao, Ming Du 0002
PRICAI (2)5