VLDB 2026 Research / reviewers in the wild / expert
Ming Dong 0004
dblp:22/2379-4
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-3700-0154ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A lightweight framework for self-reflective knowledge switching in misinformation detection
Peilin Lv, Ming Dong 0004, Po Hu 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Noise-robust item modeling and dynamic multiview contrastive learning for multimodal recommendation
Zihao Gong, Jiawei Wang 0027, Po Hu 0001, Ming Dong 0004, Zhifei Li 0009, Yan Zhang 0077, Miao Zhang 0036 |
Inf. Process. Manag. | 4 |
| 2026 | P-CLIP: Progressive Discrepancy Learning for One-Shot Text-to-Image Person Re-IdentificationabstractOne-shot Text-to-Image Person Re-Identification (One-shot TIReID) aims to construct a TIReID model using only a single labeled image-text pair per identity, along with a large pool of unlabeled person images. While supervised learning in text-to-image person re-identification has demonstrated high effectiveness, the requirement for extensive annotated data, both in terms of identities and corresponding textual descriptions, makes it impractical for large-scale camera networks. One-shot TIReID presents a promising approach to reduce the annotation burden. The primary challenge in one-shot TIReID lies in establishing consistent visual-textual correspondences across diverse viewing conditions, particularly in the absence of cross-view paired data. To address this challenge, we propose a novel progressive discrepancy learning framework, termed P-CLIP, which aims to establish a shared embedding space that is robust to view-specific biases. To achieve this goal, we dynamically construct multi-view image-text pairs based on a single labeled pair and simultaneously project the multi-view data into a unified embedding space. Specifically, we propose a Progressive Multi-View Generation method (MVG) to generate multiple noisy views from a single labeled instance for training. To mitigate cross-view ambiguities, we introduce a Cross-View Discrepancy Learning module (CDL) that leverages the discrepancies among different views to guide the learning of cross-view visual-textual correspondences. This approach effectively integrates multimodal error correction into the person re-identification domain. Furthermore, to enhance the effectiveness of visual-textual correspondence learning, we propose a Compact Cross-Modal Matching Loss (CCM), which suppresses unmatched pairs while emphasizing matched ones. Extensive experiments were conducted on three benchmark datasets, and the experimental results demonstrate the effectiveness of our proposed method. The data and codes are available at https://github.com/Itachjw/P-CLIP/tree/main. Chengji Wang, Ming Dong 0004, Mang Ye, Hao Sun 0014, Xingpeng Jiang |
IEEE Trans. Image Process. | 2 |
| 2025 | SMILE: Semantic Multi-Scale Integration and LLM-Enhanced Influenza-Like Illness ForecastingabstractInfluenza-like illness (ILI) forecasting is crucial for effective public health intervention, but existing models often fail to capture the complex temporal and semantic patterns inherent in epidemic data. Traditional statistical techniques and even advanced deep learning methods predominantly leverage numerical time series data, thereby overlooking contextual medical and epidemiological insights that could enhance the performance. Recent progress in large language models (LLMs) has illustrated their exceptional effectiveness in integrating semantic understanding into natural language processing tasks within the medical context. Motivated by these developments, we propose SMILE (Semantic Multi-scale Integration and LLM-Enhanced network), a novel multi-modal forecasting framework designed to integrate LLM-derived semantic features with multi-scale temporal analysis. Built upon TimeMixer architecture, SMILE introduces an automatic semantic feature extraction system using LLMs, adaptive fusion mechanisms for integrating textual and temporal data, and demonstrates robust performance improvements. Extensive experiments on ILI and benchmark datasets confirm that SMILE significantly outperforms state-of-the-art forecasting methods, highlighting the value of incorporating semantic context into time series disease prediction. Ming Dong 0004, Qianxiao Fang, Hao Sun 0014, Weizhong Zhao, Tingting He 0003 |
BIBM | 1 |
| 2025 | Retrieval-Augmented Generation for Large Language Model based Few-shot Chinese Spell CheckingabstractLarge language models (LLMs) are naturally suitable for Chinese spelling check (CSC) task in few-shot scenarios due to their powerful semantic understanding and few-shot learning capabilities. Recent CSC research has begun to use LLMs as foundational models. However, most current datasets are primarily focused on errors generated during the text generation process, with little attention given to errors occurring in the modal conversion process. Furthermore, existing LLM-based CSC methods often rely on fixed prompt samples, which limits the performance of LLMs. Therefore, we propose a framework named RagID (Retrieval-Augment Generation and Iterative Discriminator Strategy). By utilizing semantic-based similarity search and an iterative discriminator mechanism, RagID can provide well-chosen prompt samples and reduce over-correction issues in LLM-based CSC. RagID demonstrates excellent effectiveness in few-shot scenarios. We conducted comprehensive experiments, and the results show that RagID achieves the best performance on dataset that include data from multiple domains and dataset containing modal conversion spelling errors. The dataset and method are available online. Ming Dong 0004, Changyin Luo, Tingting He 0003 |
COLING | 1 |
| 2025 | DSCD: Large Language Model Detoxification with Self-Constrained DecodingabstractDetoxification in large language models (LLMs) remains a significant research challenge.Existing decoding detoxification methods are all based on external constraints, which require additional resource overhead and lose generation fluency.This work innovatively proposes Detoxification with Self-Constrained Decoding (DSCD), a novel method for LLMs detoxification without parameter fine-tuning.DSCD strengthens the inner next-token distribution of the safety layer while weakening that of hallucination and toxic layer during output generation.This effectively diminishes toxicity and enhances output safety.DSCD offers lightweight, high compatibility, and plug-andplay capabilities, readily integrating with existing detoxification methods for further performance improvement.Extensive experiments on representative open-source LLMs and public datasets validate DSCD's effectiveness, demonstrating state-of-the-art (SOTA) performance in both detoxification and generation fluency, with superior efficiency compared to existing methods.These results highlight DSCD's potential as a practical and scalable solution for safer LLM deployments.For more details, please refer to the project repository: https://github.com/ZHANGJINKUI/DSCD. Ming Dong 0004, Jinkui Zhang, Bolong Zheng, Xinhui Tu, Po Hu 0001, Tingting He 0003 |
EMNLP | 1 |
| 2025 | PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkabstractWith the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https://github.com/mingyao1120/PC-Net. Mingyao Zhou, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Chengji Wang, Mang Ye |
NeurIPS | 4 |
| 2025 | Multi-faceted data augmentation for aspect-based sentiment analysis via large language models
Rui Fan 0005, Tingting He 0003, Ming Dong 0004 |
Knowl. Based Syst. | 3 |
| 2025 | MRBalance: A framework for enhancing event causality identification in multi-agent debates via role assignment
Xuanhong Li, Po Hu 0001, Ming Dong 0004 |
Knowl. Based Syst. | 4 |
| 2025 | Dual Causes Generation Assisted Model for Multimodal Aspect-Based Sentiment ClassificationabstractMultimodal aspect-based sentiment classification (MABSC) aims to identify the sentiment polarity toward specific aspects in multimodal data. It has gained significant attention with the increasing use of social media platforms. Existing approaches primarily focus on analyzing the content of posts to predict sentiment. However, they often struggle with limited contextual information inherent in social media posts, hindering accurate sentiment detection. To overcome this issue, we propose a novel multimodal dual cause analysis (MDCA) method to track the underlying causes behind expressed sentiments. MDCA can provide additional reasoning cause (RC) and direct cause (DC) to explain why users express certain emotions, thus helping improve the accuracy of sentiment prediction. To develop a model with MDCA, we construct MABSC datasets with RC and DC by utilizing large language models (LLMs) and visual-language models. Subsequently, we devise a multitask learning framework that leverages the datasets with cause data to train a small generative model, which can generate RC and DC, and predict the sentiment assisted by these causes. Experimental results on MABSC benchmark datasets demonstrate that our MDCA model achieves the state-of-the-art performance, and the small fine-tuned model exhibits superior adaptability to MABSC compared to large models like ChatGPT and BLIP-2. Rui Fan 0005, Tingting He 0003, Xinhui Tu, Ming Dong 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Spatial Dual Context Learning for Weakly-supervised Group Activity Recognition in Still-imagesabstractThis paper investigates a new task, Weakly- supervised Group Activity Recognition in Still-images (WGARS), which aims to extend the applicability of Group Activity Recognition (GAR) to broader scenarios, such as low-latency domains. To tackle this challenge, we propose a Spatial Dual Context Transformer (SDCT), comprising a Dual Context Encoder (DCE) and a Dual Context Decoder (DCD). The DCE module individually encodes holistic context with integral relations of overall actors, and encodes partial context with individual features in still images. Subsequently, the DCD module explores the complementarity between holistic and partial contexts, and alternatively updates these encoded contexts to enhance the interaction of actors. Additionally, auxiliary supervised contrastive learning is incorporated to mitigate activity confusion. The proposed SDCT attains state-of-the-art performance on Volleyball and NBA datasets in WGARS. Notably, SDCT even outperforms recent methods when extended to the weakly-supervised GAR in videos task on Volleyball dataset. Dunbo Ning, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004 |
ICME | 6 |
| 2024 | Meta-path reasoning of knowledge graph for commonsense question answering
Miao Zhang 0036, Tingting He 0003, Ming Dong 0004 |
Frontiers Comput. Sci. | 3 |
| 2024 | Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu |
Knowl. Based Syst. | 5 |
| 2023 | A Hybrid Corpus based Fine-grained Semantic Alignment Method for Pre-trained Language Model of Ancient Chinese PoetryabstractAncient Chinese poetry (ACP) is a vital component of Chinese traditional culture. Enhancing the performance of related downstream tasks demands the development of high-quality pre-trained language models (PLMs) dedicated to ACP. Notably, the semantics of ACP significantly differ from modern Chinese. Existing PLMs have limited knowledge of ACP and are inadequately aligned with the semantic space of modern Chinese, which constrains the utility for tasks related to ACP. In this paper, we propose a fine-tuning strategy to establish a precise alignment between ACP and modern Chinese semantics on sentence level. This strategy involves the inclusion of corresponding modern Chinese translations alongside original ancient poems, creating a hybrid corpus. This corpus facilitates a more effective transfer of knowledge from existing PLMs to the domain of ACP. Furthermore, we employ a training strategy based on a glyph-based foundational PLM, enabling meticulous fine-tuning. Consequently, we develop a specialized PLM named CP-ChineseBERT. To evaluate the effectiveness of our proposed strategies, we conducted experiments on two real-world datasets, focusing on tasks related to ACP sentiment classification and ACP title prediction. The experimental results demonstrate the significant improvements in performance achieved through our innovative approaches. Tingting He 0003, Ming Dong 0004, Zheming Zhang, Xinhui Tu |
IEEE Big Data | 4 |
| 2023 | Incorporating BERT With Probability-Aware Gate for Spoken Language UnderstandingabstractSpoken language understanding (SLU) is an essential part of a task-oriented dialogue system, which mainly includes intent detection and slot filling. Some existing approaches obtain enhanced semantic representation by establishing the correlation between two tasks. However, those methods show little improvement when applied to BERT, since BERT has learned rich semantic features. In this paper, we propose a BERT-based model with the probability-aware gate mechanism, called PAGM (ProbabilityAwareGatedModel). PAGM aims to learn the correlation between intent and slot from the perspective of probability distribution, which explicitly utilizes intent information to guide slot filling. Besides, in order to efficiently incorporate BERT with the probability-aware gate, we design the stacked fine-tuning strategy. This approach introduces a mid-stage before target model training, which enables BERT to get better initialization for final training. Experiments show that PAGM achieves significant improvement on two benchmark datasets, and outperforms the previous state-of-the-art results. Xinhui Tu, Ming Dong 0004, Tingting He 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | DRLK: Dynamic Hierarchical Reasoning with Language Model and Knowledge Graph for Question AnsweringabstractIn recent years, Graph Neural Network (GNN) approaches with enhanced knowledge graphs (KG) perform well in question answering (QA) tasks.One critical challenge is how to effectively utilize interactions between the QA context and KG.However, existing work only adopts the identical QA context representation to interact with multiple layers of KG, which results in a restricted interaction.In this paper, we propose DRLK (Dynamic Hierarchical Reasoning with Language Model and Knowledge Graphs), a novel model that utilizes dynamic hierarchical interactions between the QA context and KG for reasoning.DRLK extracts dynamic hierarchical features in the QA context, and performs inter-layer and intra-layer interactions on each iteration, allowing the KG representation to be grounded with the hierarchical features of the QA context.We conduct extensive experiments on four benchmark datasets in medical QA and commonsense reasoning.The experimental results demonstrate that DRLK achieves state-of-theart performances on two benchmark datasets and performs competitively on the others 1 . Miao Zhang 0036, Rufeng Dai, Ming Dong 0004, Tingting He 0003 |
EMNLP | 3 |
| 2022 | Deep reinforcement learning based ensemble model for rumor tracking
Guohui Li 0001, Ming Dong 0004, Lingfeng Ming, Changyin Luo, Xiaofei Hu, Bolong Zheng |
Inf. Syst. | 2 |
| 2020 | An Effective Fleet Management Strategy for Collaborative Spatio-Temporal Searching: GIS CupabstractThe ACM SIGSPATIAL GIS Cup 2020 focuses on the Collaborative Spatio-Temporal Searching (CSTS) problem, in which a fleet of mobile agents search for stationary resources on a road network. While each resource can be obtained by exactly one agent, agents can collaborate to obtain resources as quickly as possible. The key of solving CSTS is to guide agents to "hotspot" areas and to avoid the competition by considering agent collaboration. We propose a fleet management method by formulating CSTS as a minimum cost flow problem, called MCF-FM. In addition, we develop a continuous order dispatch strategy. Our submission is the top performer in the agent utilization scenario and runner-up in the customer experience scenario. Our source code is available at: https://github.com/Chriszblong/MCF-FM. Lingfeng Ming, Ming Dong 0004, Bolong Zheng |
SIGSPATIAL/GIS | 3 |
| 2020 | Misinformation-oriented expert finding in social networks
Guohui Li 0001, Ming Dong 0004, Fuming Yang, Jiansen Yuan, Congyuan Jin, Nguyen Quoc Viet Hung, Phan Thanh Cong, Bolong Zheng |
World Wide Web | 2 |
| 2019 | Multiple Rumor Source Detection with Graph Convolutional NetworksabstractDetecting rumor source in social networks is one of the key issues for defeating rumors automatically. Although many efforts have been devoted to defeating online rumors, most of them are proposed based an assumption that the underlying propagation model is known in advance. However, this assumption may lead to impracticability on real data, since it is usually difficult to acquire the actual underlying propagation model. Some attempts are developed by using label propagation to avoid the limitation caused by lack of prior knowledge on the underlying propagation model. Nonetheless, they still suffer from the shortcoming that the node label is simply an integer which may restrict the prediction precision. In this paper, we propose a deep learning based model, namely GCNSI (Graph Convolutional Networks based Source Identification), to locate multiple rumor sources without prior knowledge of underlying propagation model. By adopting spectral domain convolution, we build node representation by utilizing its multi-order neighbors information such that the prediction precision on the sources is improved. We conduct experiments on several real datasets and the results demonstrate that our model outperforms state-of-the-art model. Ming Dong 0004, Bolong Zheng, Nguyen Quoc Viet Hung, Han Su 0001, Guohui Li 0001 |
CIKM | 1 |