VLDB 2026 Research / reviewers in the wild / expert
Junbo Huang
dblp:138/8642
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-3192-5896ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Variance to Invariance: Qualitative Content Analysis for Narrative Graph Annotation
Junbo Huang, Max Weinig, Ulrich Fritsche, Ricardo Usbeck |
LREC | 1 |
| 2025 | LLM Agents for Georelating - A New Task for Locating EventsabstractAccurately identifying disaster-affected areas is crucial for data-driven disaster resilience. In response, we introduce Georelating, a task that infers affected areas from textual reports containing complex locative expressions, moving beyond traditional geoparsing approaches that rely on explicit point locations. Georelating instead combines resolving unnamed regions and reasoning about spatial relations to represent event-affected areas within standardized Discrete Global Grid Systems (DGGSs). Kai Moltzen, Junbo Huang, Ricardo Usbeck |
SIGSPATIAL/GIS | 2 |
| 2025 | ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs
Mikhail Salnikov, Andrey Sakhovskiy, Irina Nikishina, Aida Usmanova, Angelie Kraft, Cedric Möller, Debayan Banerjee, Junbo Huang, Longquan Jiang 0001, Rana Abdullah, Xi Yan 0001, Elena Tutubalina, Ricardo Usbeck, Alexander Panchenko |
NLDB (1) | 8 |
| 2025 | Robust Polyp Detection and Diagnosis Through Compositional Prompt-Guided Diffusion ModelsabstractColorectal cancer (CRC) is a significant global health concern, and early detection through screening plays a critical role in reducing mortality. While deep learning models have shown promise in improving polyp detection, classification, and segmentation, their generalization across diverse clinical environments, particularly with out-of-distribution (OOD) data, remains a challenge. Multi-center datasets like PolypGen have been developed to address these issues, but their collection is costly and time-consuming. Traditional data augmentation techniques provide limited variability, failing to capture the complexity of medical images. Diffusion models have emerged as a promising solution for generating synthetic polyp images, but the image generation process in current models mainly relies on segmentation masks as the condition, limiting their ability to capture the full clinical context. To overcome these limitations, we propose a Progressive Spectrum Diffusion Model (PSDM) that integrates diverse clinical annotations-such as segmentation masks, bounding boxes, and colonoscopy reports-by transforming them into compositional prompts. These prompts are organized into coarse and fine components, allowing the model to capture both broad spatial structures and fine details, generating clinically accurate synthetic images. By augmenting training data with PSDM-generated samples, our model significantly improves polyp detection, classification, and segmentation. For instance, on the PolypGen dataset, PSDM increases the F1 score by 2.12% and the mean average precision by 3.09%, demonstrating superior performance in OOD scenarios and enhanced generalization. Peiyao Fu, Junbo Huang, Quanlin Li, Pinghong Zhou, Zhihua Wang 0008, Fei Wu 0001, Shuo Wang 0011, Xian Yang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Revisiting Supervised Contrastive Learning for Microblog ClassificationabstractMicroblog content (e.g., Tweets) is noisy due to its informal use of language and its lack of contextual information within each post.To tackle these challenges, state-of-the-art microblog classification models rely on pre-training language models (LMs).However, pre-training dedicated LMs is resource-intensive and not suitable for small labs.Supervised contrastive learning (SCL) has shown its effectiveness with small, available resources.In this work, we examine the effectiveness of fine-tuning transformer-based language models, regularized with a SCL loss for English microblog classification.Despite its simplicity, the evaluation on two English microblog classification benchmarks (TweetEval and Tweet Topic Classification) shows an improvement over baseline models.The result shows that, across all subtasks, our proposed method has a performance gain of up to 11.9 percentage points.All our models are open source. Junbo Huang, Ricardo Usbeck |
EMNLP | 1 |
| 2024 | SSNF: Optimizing Entity Alignment with a Novel Structural and Semantic Neighbor Filtering
Junbo Huang |
KSEM (2) | 1 |