EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Xu 0006
dblp:79/8798-6
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-4386-5277ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 43% Language models and text generation · 43% Question answering and dialogue systems · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 62% Data mining · 38% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and social media mining › citation network analysis
citation prediction |
1.0 | 1 | 2026 | From Newborn to Impact: Bias-Aware Citation Prediction · WWW 2026 |
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models · EMNLP 2025 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering |
0.9 | 1 | 2025 | MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models · EMNLP 2025 |
Natural language and speech › Language models and text generation › prompting
prompt sensitivity |
0.9 | 1 | 2025 | Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models · AAAI 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration |
0.9 | 1 | 2025 | Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models · AAAI 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty decomposition |
0.9 | 1 | 2025 | Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models · AAAI 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models · AAAI 2025 |
Data mining › representation learning
graph representation learning |
0.3 | 1 | 2026 | From Newborn to Impact: Bias-Aware Citation Prediction · WWW 2026 |
Data mining › representation learning › graph representation learning
heterogeneous information network embedding |
0.3 | 1 | 2026 | From Newborn to Impact: Bias-Aware Citation Prediction · WWW 2026 |
Medical and health informatics
clinical decision support |
0.3 | 1 | 2025 | MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
domain-specific knowledge prompting · 1.7bidirectional entailment clustering · 1.7regularization · 1.0multi-agent feature extraction · 1.0graph representation learning · 1.0GroupDRO · 1.0semantic sampling · 0.9paraphrasing perturbation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Newborn to Impact: Bias-Aware Citation PredictionabstractAs a key to accessing research impact, citation dynamics underpins research evaluation, scholarly recommendation, and the study of knowledge diffusion. Citation prediction is particularly critical for newborn papers, where early assessment must be performed without citation signals and under highly long-tailed distributions. We identify two key research gaps: (i) insufficient modeling of implicit factors of scientific impact, leading to reliance on coarse proxies; and (ii) a lack of bias-aware learning that can deliver stable predictions on lowly cited papers. We address these gaps by proposing a Bias-Aware Citation Prediction Framework, which combines multi-agent feature extraction with robust graph representation learning. First, a multi-agent x graph co-learning module derives fine-grained, interpretable signals, such as reproducibility, collaboration network, and text quality, from metadata and external resources, and fuses them with heterogeneous-network embeddings to provide rich supervision even in the absence of early citation signals. Second, we incorporate a set of robust mechanisms: a two-stage forward process that routes explicit factors through an intermediate exposure estimate, GroupDRO to optimize worst-case group risk across environments, and a regularization head that performs what-if analyses on controllable factors under monotonicity and smoothness constraints. Comprehensive experiments on two real-world datasets demonstrate the effectiveness of our proposed model. Specifically, our model achieves around a 13% reduction in error metrics (MALE and RMSLE) and a notable 5.5% improvement in the ranking metric (NDCG) over the baseline methods. Mingfei Lu, Mengjia Wu, Jiawei Xu 0006, Weikai Li 0002, Feng Liu 0003, Ying Ding 0001, Yizhou Sun, Jie Lu 0001, Yi Zhang 0095 |
WWW | 3 |
| 2025 | Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsabstractAn interesting behavior in large language models (LLMs) is prompt sensitivity. When provided with different but semantically equivalent versions of the same prompt, models may produce very different distributions of answers. This suggests that the uncertainty reflected in a model's output distribution for one prompt may not reflect the model's uncertainty about the meaning of the prompt. We model prompt sensitivity as a type of generalization error, and show that sampling across the semantic concept space with paraphrasing perturbations improves uncertainty calibration without compromising accuracy. Additionally, we introduce a new metric for uncertainty decomposition in black-box LLMs that improves upon entropy-based decomposition by modeling semantic continuities in natural language generation. We show that this decomposition metric can be used to quantify how much LLM uncertainty is attributed to prompt sensitivity. Our work introduces a new way to improve uncertainty calibration in prompt-sensitive language models, and provides evidence that some LLMs fail to exhibit consistent general reasoning about the meanings of their inputs. Kyle Cox, Jiawei Xu 0006, Yikun Han, Chi-Yang Hsu, Tianlong Chen 0001, Walter Gerych, Ying Ding 0001 |
AAAI | 2 |
| 2025 | MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsabstractAdvancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliability.A critical challenge lies in hallucination, where models generate plausible yet factually incorrect outputs.In the medical domain, this poses serious risks to patient safety and clinical decision-making.To address this, we introduce MedHallu, one of the first benchmark specifically designed for medical hallucination detection.Med-Hallu comprises 10,000 high-quality questionanswer pairs derived from PubMedQA, with hallucinated answers systematically generated through a controlled pipeline.Our experiments show that state-of-the-art LLMs, including GPT-4o, Llama-3.1, and the medically fine-tuned UltraMedical, struggle with this binary hallucination detection task, with the best model achieving an F1 score as low as 0.625 for detecting "hard" category hallucinations.Using bidirectional entailment clustering, we show that harder-to-detect hallucinations are semantically closer to ground truth.Through experiments, we also show incorporating domainspecific knowledge and introducing a "not sure" category as one of the answer categories improves the precision and F1 scores by up to 38% relative to baselines. Shrey Pandit, Jiawei Xu 0006, Junyuan Hong, Zhangyang Wang, Tianlong Chen 0001, Kaidi Xu, Ying Ding 0001 |
EMNLP | 2 |
| 2025 | Knowledge integration and diffusion structures of interdisciplinary research: A large-scale analysis based on propensity score matchingabstractAbstract While facilitating science, interdisciplinary research (IDR) has a heavier cognitive burden for researchers compared to unidisciplinary research (UDR). Yet, little has been known about patterns of knowledge integration and diffusion structures of IDR. Here we adopt a causal inference strategy, namely propensity score matching, with all journal publications in 2005 in Microsoft Academic Graph to better understand the IDR effect in various research fields. We use the diversity of reference fields of one paper as the proxy of the paper's interdisciplinarity and estimate the effect of a research article being IDR on its knowledge integration and diffusion measured by its high‐order citation/reference cascade. We find that, in disciplines where IDR articles are less popular, such as mathematics, physics, and chemistry, IDR needs a more extensive knowledge base than UDR to gain a similar number of citations. In disciplines where IDR articles are more popular, for example, psychology, geology, biology, and economics, a small knowledge base is enough for a high‐impact IDR article. As to knowledge diffusion, no matter whether IDR or UDR, a more extensive knowledge base leads to stronger knowledge diffusion ability. Findings imply potential drawbacks of pure interdisciplinarity‐oriented research policy; rather, the establishment of policies may vary across disciplines. Jiawei Xu 0006, Zhihan Zheng, Win-bin Huang, Yi Bu 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |