EDBT 2026 Demo / reviewers in the wild / expert
Mingying Xu
dblp:289/1484
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal RetrievalabstractCross-modal hashing has emerged as a pivotal solution for efficient retrieval across diverse modalities, such as images and texts, by mapping them into compact binary hash spaces. However, in real-world scenarios, the modalities data is often missing or misaligned. Existing methods are most rely on fully paired training data and ignore missing or misaligned modalities data, resulting in the semantic inconsistencies. To address these challenges, we propose an Adaptive Graph Attention-Based Discrete Hashing (AGADH) method, which consists of three parts. First, to solve the problem of missing modalities, AGADH employs a masked completion strategy to reconstruct missing modalities. Second, to mitigate semantic misalignment, AGADH leverages a Graph Attention Network (GAT) encoder-decoder architecture with alignment module to construct features from different modalities. Additionally, to enhance the fusion performance, an adaptive fusion module dynamically adjusting the contributions of image and text modalities with learnable weighting coefficients is proposed. Extensive experiments on three benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr-25K, demonstrating that AGADH outperforms state-of-the-art methods in both fully paired and incompletely paired scenarios, showing its robustness and effectiveness in cross-modal retrieval tasks. Shuang Zhang 0009, Lei Shi 0030, Huilong Jin, Feifei Kou, Pengfei Zhang 0010, Mingying Xu, Pengtao Lv |
AAAI | 7 |
| 2026 | MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative RecommendationabstractGenerative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods. Yuqiu Zhao, Lei Shi 0030, Yan Zhong 0001, Feifei Kou, Pengfei Zhang 0010, Jiwei Zhang 0007, Mingying Xu |
AAAI | 7 |
| 2026 | Dual-perspective hypergraph learning network for multimodal entity and relation extraction
Jie Liu 0022, Mingying Xu, Baowen Wu, Linqi Song, Yinqiao Li, Lei Shi 0030, Feifei Kou |
Expert Syst. Appl. | 3 |
| 2026 | A self-modified hypergraph neural network for multimodal relation extraction
Mingying Xu, Jie Liu 0022, Linqi Song, Yinqiao Li, Lei Shi 0030 |
Inf. Process. Manag. | 1 |
| 2026 | Beyond individual diagnosis: a graph learning framework with bidirectional distillation for group cognitive diagnosis
Xinhua Wang 0003, Zhenxi Sun, Mingying Xu, Peiyu Liu 0001, Lei Guo 0008 |
Knowl. Inf. Syst. | 4 |
| 2026 | Dual Graph Network Hashing for Cross-Modal Retrieval
Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Weiping Ding 0001, Mingying Xu, Muhammet Deveci |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2025 | Leveraging the Dual Capabilities of LLM: LLM-Enhanced Text Mapping Model for Personality DetectionabstractPersonality detection aims to deduce a user’s personality from their published posts. The goal of this task is to map posts to specific personality types. Existing methods encode post information to obtain user vectors, which are then mapped to personality labels. However, existing methods face two main issues: first, only using small models makes it hard to accurately extract semantic features from multiple long documents. Second, the relationship between user vectors and personality labels is not fully considered. To address the issue of poor user representation, we utilize the text embedding capabilities of LLM. To solve the problem of insufficient consideration of the relationship between user vectors and personality labels, we leverage the text generation capabilities of LLM. Therefore, we propose the LLM-Enhanced Text Mapping Model (ETM) for Personality Detection. The model applies LLM’s text embedding capability to enhance user vector representations. Additionally, it uses LLM’s text generation capability to create multi-perspective interpretations of the labels, which are then used within a contrastive learning framework to strengthen the mapping of these vectors to personality labels. Experimental results show that our model achieves state-of-the-art performance on benchmark datasets. Weihong Bi, Feifei Kou, Lei Shi 0030, Yawen Li 0001, Hai-Sheng Li 0002, Jinpeng Chen 0001, Mingying Xu |
AAAI | 7 |
| 2025 | EVICheck: Evidence-Driven Independent Reasoning and Combined Verification Method for Fact-CheckingabstractLarge Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have demonstrated significant potential in automated fact-checking. However, existing methods face limitations in insufficient evidence utilization and lack of explicit verification criteria. Specifically, these approaches aggregate evidence for collective reasoning without independently analyzing each piece, hindering their ability to leverage the available information thoroughly. Additionally, they rely on simple prompts or few-shot learning for verification, which makes truthfulness judgments less reliable, especially for complex claims. To address these limitations, we propose a novel method to enhance evidence utilization and introduce explicit verification criteria, named EVICheck. Our approach independently reasons each evidence piece and synthesizes the results to enable more thorough exploration and enhance interpretability. Additionally, by incorporating fine-grained truthfulness criteria, we make the model's verification process more structured and reliable, especially when handling complex claims. Experimental results on the public RAWFC dataset demonstrate that EVICheck achieves state-of-the-art performance across all evaluation metrics. Our method demonstrates strong potential in fake news verification, significantly improving the accuracy. Lei Shi 0030, Feifei Kou, Ligu Zhu, Chen Ma 0003, Pengfei Zhang 0010, Mingying Xu |
IJCAI | 7 |
| 2025 | Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal RetrievalabstractThe demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms. Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Meiyu Liang, Mingying Xu |
NeurIPS | 9 |
| 2025 | Making meta-learning solve cross-prompt automatic essay scoring
Jie Liu 0022, Mingying Xu, Liguang Yang, Jianshe Zhou |
Expert Syst. Appl. | 5 |
| 2025 | Fine-grained entity typing based on hyperbolic representation and label-context interaction
Mingying Xu, Jie Liu 0022, Weiping Ding 0001, Lei Shi 0030, Kaiyang Zhong |
Inf. Sci. | 1 |
| 2024 | Self-derived Knowledge Graph Contrastive Learning for RecommendationabstractKnowledge Graphs (KGs) serve as valuable auxiliary information to improve the accuracy of recommendation systems. Previous methods have leveraged the knowledge graph to enhance item representation and thus achieve excellent performance. However, these approaches heavily rely on high-quality knowledge graphs and learn enhanced representations with the assistance of carefully designed triplets. Furthermore, the emergence of knowledge graphs has led to models that ignore the inherent relationships between items and entities. To address these challenges, we propose a Self-Derived Knowledge Graph Contrastive Learning framework (CL-SDKG) to enhance recommendation systems. Specifically, we employ the variational graph reconstruction technique to estimate the Gaussian distribution of user-item nodes corresponding to the graph neural network aggregation layer. This process generates multiple KGs, referred to as self-derived KGs. The self-derived KG acquires more robust perceptual representations through the consistency of the estimated structure. Besides, the self-derived KG allows models to focus on user-item interactions and reduce the negative impact of miscellaneous dependencies introduced by conventional KGs. Finally, we apply contrastive learning to the self-derived KG to further improve the robustness of CL-SDKG through the traditional KG contrast-enhanced process. We conducted comprehensive experiments on three public datasets, and the results demonstrate that our CL-SDKG outperforms state-of-the-art baselines. Lei Shi 0030, Pengtao Lv, Feifei Kou, Jia Luo 0001, Mingying Xu |
ACM Multimedia | 7 |
| 2024 | An Entailment Tree Generation Approach for Multimodal Multi-Hop Question Answering with Mixture-of-Experts and Iterative Feedback MechanismabstractWith the rise of large-scale language models (LLMs), it is currently popular and effective to convert multimodal information into text descriptions for multimodal multi-hop question answering. However, we argue that the current methods of multi-modal multi-hop question answering still mainly face two challenges: 1) The retrieved evidence containing a large amount of redundant information, inevitably leads to a significant drop in performance due to irrelevant information misleading the prediction. 2) The reasoning process without interpretable reasoning steps makes the model difficult to discover the logical errors for handling complex questions. To solve these problems, we propose a unified LLMs-based approach but without heavily relying on them due to the LLM's potential errors, and innovatively treat multimodal multi-hop question answering as a joint entailment tree generation and question answering problem. Specifically, we design a multi-task learning framework with a focus on facilitating common knowledge sharing across interpretability and prediction tasks while preventing task-specific errors from interfering with each other via mixture of experts. Afterward, we design an iterative feedback mechanism to further enhance both tasks by feeding back the results of the joint training to the LLM for regenerating entailment trees, aiming to iteratively refine the potential answer. Notably, our method has won the first place in the official leaderboard of WebQA (since April 10, 2024), and achieves competitive results on MultimodalQA. Haocheng Lv, Jie Liu 0022, Jianyong Duan, Hao Wang 0018, Mingying Xu |
ACM Multimedia | 8 |
| 2022 | A scientific research topic trend prediction model based on multi-LSTM and graph convolutional networkabstractPredicting the development trend of future scientific research not only provides a reference for researchers to understand the development of the discipline, but also provides support for decision-making and fund allocation for decision-makers. The continuous growth of scientific publications has brought challenges to track the development trends of scientific research topics. The existing topic trend prediction methods have proved that the research topic trend of a publication is influenced by other peer publications. However, they ignore the fact that the research topics of different publications belong to different research topic space. Moreover, the existing topic prediction methods do not fully consider the interactive influence among publications that the research topic of one publication affects the topics of other publications, it is also influenced by the research topics of other publications. In line with this, this paper proposes a scientific research topic trend prediction model based on multi-long short-term memory (multi-LSTM) and Graph Convolutional Network. Specifically, multiple LSTMs are employed to map research topics of different publications into their respective topic space. Then, the graph convolutional neural network is applied to learn the scientific influence context of each publication, so that the research topic of each publication not only integrates the influence of neighbor nodes, but also considers the influence of the neighbors of the neighbor node on the research topic of the publication, so as to more accurately fuse scientific influence context of research topic of peer publications. Experiments results on the data set of scientific research papers in the field of artificial intelligence and data mining demonstrate that the model improves the prediction precision and achieves the state-of-the-art research topic trend prediction effect compared with the other baseline models. Mingying Xu, Junping Du 0001, Zhe Xue, Zeli Guan, Feifei Kou, Lei Shi 0030 |
Int. J. Intell. Syst. | 1 |
| 2021 | A semi-supervised semantic-enhanced framework for scientific literature retrieval
Mingying Xu, Junping Du 0001, Zhe Xue, Feifei Kou |
Neurocomputing | 1 |