VLDB 2026 Research / reviewers in the wild / expert
Mengjia Wu
dblp:271/4391
· DBLP profile ↗
6ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0003-3956-7808ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (1 first)Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agent-Enhanced Heterogeneous Graph RAG for Academic Question AnsweringabstractAcademic question answering requires reasoning over heterogeneous scholarly graphs, where queries range from simple attribute lookups to multi-hop inference across author--paper--venue structures. Existing retrieval-augmented generation (RAG) systems struggle in this setting due to three limitations: (1) fixed retrieval strategies that do not adapt to varying query complexity, (2) the absence of sufficiency evaluation leading to incomplete or misaligned evidence, and (3) a lack of structured verification against graph facts. To address these issues, we propose an agentic heterogeneous graph RAG method that transforms the three core stages of the RAG pipeline into explicit agentic decision steps. A query-aware retrieval agent analyzes query type and selects an appropriate graph traversal strategy; a sufficiency-aware reranking agent assesses evidence completeness and adaptively expands the retrieved subgraph; and a graph-grounded verification agent checks entity, relation, and attribute correctness before finalizing the answer. Experiments on heterogeneous graphs constructed from OpenAlex and DBLP suggest that our method consistently outperforms strong LLM, graph-augmented RAG, and agent-based baselines. Runsong Jia, Mengjia Wu, Ying Ding 0001, Jie Lu 0001, Yi Zhang 0095 |
WWW | 2 |
| 2026 | From Newborn to Impact: Bias-Aware Citation PredictionabstractAs a key to accessing research impact, citation dynamics underpins research evaluation, scholarly recommendation, and the study of knowledge diffusion. Citation prediction is particularly critical for newborn papers, where early assessment must be performed without citation signals and under highly long-tailed distributions. We identify two key research gaps: (i) insufficient modeling of implicit factors of scientific impact, leading to reliance on coarse proxies; and (ii) a lack of bias-aware learning that can deliver stable predictions on lowly cited papers. We address these gaps by proposing a Bias-Aware Citation Prediction Framework, which combines multi-agent feature extraction with robust graph representation learning. First, a multi-agent x graph co-learning module derives fine-grained, interpretable signals, such as reproducibility, collaboration network, and text quality, from metadata and external resources, and fuses them with heterogeneous-network embeddings to provide rich supervision even in the absence of early citation signals. Second, we incorporate a set of robust mechanisms: a two-stage forward process that routes explicit factors through an intermediate exposure estimate, GroupDRO to optimize worst-case group risk across environments, and a regularization head that performs what-if analyses on controllable factors under monotonicity and smoothness constraints. Comprehensive experiments on two real-world datasets demonstrate the effectiveness of our proposed model. Specifically, our model achieves around a 13% reduction in error metrics (MALE and RMSLE) and a notable 5.5% improvement in the ranking metric (NDCG) over the baseline methods. Mingfei Lu, Mengjia Wu, Jiawei Xu 0006, Weikai Li 0002, Feng Liu 0003, Ying Ding 0001, Yizhou Sun, Jie Lu 0001, Yi Zhang 0095 |
WWW | 2 |
| 2026 | Explainable prediction of knowledge recombination: A synergized method with heterogeneous hypergraph learning and large language modelsabstractDespite growing interest in graph-based models for knowledge recombination prediction using academic knowledge graphs, existing approaches suffer from significant limitations: they fail to learn informative and robust knowledge entity representations by neglecting high-order information, inadequately account for real-world dynamics, and crucially, cannot provide readable rationales for their predictions. We address these challenges with H2GLM, which reformulates traditional graph learning as heterogeneous hypergraph learning to capture high-order information, incorporating a variational autoencoder (VAE) mechanism to enhance informativeness and robustness. Our approach then integrates large language models (LLMs) with the learned graph contextual information through a step-wise methodology, enabling evidence-supported decisions with clear, readable rationales. Experimental results highlight that H2GLM outperforms previous strong graph-based and LLM-based baselines by 4% to 8% in accuracy, 3% to 9% in AUC and 5% to 8% in F1 on extensive academic knowledge graphs containing over 1,000,000 nodes, with a small amount of training data. Visualizations and case studies further illustrate our method’s substantial utility over existing approaches in real-world scenarios. Further explainability and efficiency analyses underscore the practical value of our method Mengjia Wu, Qian Liu 0012, Yi Zhang 0095 |
Inf. Process. Manag. | 2 |
| 2025 | Scaling research aim identification: Language models for classifying scientific and societal-oriented studiesabstractAbstract The classification of research according to its aims has been a longstanding focus in the fields of quantitative science studies and R&D statistics. Since 1963, the Organization for Economic Co‐operation and Development (OECD) has employed a classical distinction among basic, applied, and experimental research. Building on this framework, our previous work highlighted the utility of differentiating between scientific and societal progress as two primary research objectives. This distinction enabled the quantitative analysis of scientific publication abstracts and the development of an automated method for large‐scale classification. In the current study, we systematically evaluate text classification techniques, including traditional text mining models, classification tools, BERT‐based language models, and decoder‐only large language models (LLMs) such as ChatGPT. Our findings show that the fine‐tuned GPT‐4o‐mini model performs the best among single‐model approaches. However, traditional and BERT‐based models outperform in certain fine‐grained classification tasks. Leveraging majority voting strategies to incorporate their strengths yields performance comparable to closed‐source GPT models. A case study on 10 biomedical journals further validates the method, demonstrating strong alignment between journal scopes, model predictions, and outputs generated by the fine‐tuned GPT‐4o‐mini model. These results highlight the robustness and practical effectiveness of the proposed methodology for nuanced research aim classification. Mengjia Wu, Gunnar Sivertsen, Lin Zhang 0004, Fan Qi, Yi Zhang 0095 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2023 | Stepping beyond your comfort zone: Diffusion-based network analytics for knowledge trajectory recommendationabstractAbstract Predicting a researcher's knowledge trajectories beyond their current foci can leverage potential inter‐/cross‐/multi‐disciplinary interactions to achieve exploratory innovation. In this study, we present a method of diffusion‐based network analytics for knowledge trajectory recommendation. The method begins by constructing a heterogeneous bibliometric network consisting of a co‐topic layer and a co‐authorship layer. A novel link prediction approach with a diffusion strategy is then used to capture the interactions between social elements (e.g., collaboration) and knowledge elements (e.g., technological similarity) in the process of exploratory innovation. This diffusion strategy differentiates the interactions occurring among homogeneous and heterogeneous nodes in the heterogeneous bibliometric network and weights the strengths of these interactions. Two sets of experiments—one with a local dataset and the other with a global dataset—demonstrate that the proposed method is prior to 10 selected baselines in link prediction, recommender systems, and upstream graph representation learning. A case study recommending knowledge trajectories of information scientists with topical hierarchy and explainable mediators reveals the proposed method's reliability and potential practical uses in broad scenarios. Yi Zhang 0095, Mengjia Wu, Guangquan Zhang 0001, Jie Lu 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2021 | Unraveling the capabilities that enable digital transformation: A data-driven methodology and the case of artificial intelligence
Mengjia Wu, Dilek Cetindamar, Yi Zhang 0095 |
Adv. Eng. Informatics | 1 |