EDBT 2026 Demo / reviewers in the wild / expert
Yuntong Hu
dblp:323/9826
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-3802-9039ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Network Tomography with Path-Centric Graph Neural NetworkabstractNetwork tomography is a crucial problem in network monitoring, where the observable path performance metric values are used to infer the unobserved ones, making it essential for tasks such as route selection, fault diagnosis, and traffic control. However, most existing methods either assume complete knowledge of network topology and metric formulas—an unrealistic expectation in many real-world scenarios with limited observability—or rely entirely on black-box end-to-end models. To tackle this, in this paper, we argue that a good network tomography requires synergizing the knowledge from both data and appropriate inductive bias from (partial) prior knowledge. To see this, we propose Deep Network Tomography (DeepNT), a novel framework that leverages a path-centric graph neural network to predict path performance metrics without relying on predefined hand-crafted metrics, assumptions, or the real network topology. The path-centric graph neural network learns the path embedding by inferring and aggregating the embeddings of the sequence of nodes that compose this path. Training path-centric graph neural networks requires learning the neural netowrk parameters and network topology under discrete constraints induced by the observed path performance metrics, which motivates us to design a learning objective that imposes connectivity and sparsity constraints on topology and path performance triangle inequality on path performance. Extensive experiments on real-world and synthetic datasets demonstrate the superiority of DeepNT in predicting performance metrics and inferring graph topology compared to state-of-the-art methods. Yuntong Hu, Liang Zhao 0002 |
WSDM | 1 |
| 2025 | GraphNarrator: Generating Textual Explanations for Graph Neural NetworksabstractBo Pan, Zhen Xiong, Guanchen Wu, Zheng Zhang, Yifei Zhang, Yuntong Hu, Liang Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bo Pan 0009, Zhen Xiong, Guanchen Wu, Zheng Zhang 0047, Yifei Zhang 0006, Yuntong Hu, Liang Zhao 0002 |
ACL (1) | 6 |
| 2025 | TAGA: Text-Attributed Graph Self-Supervised Learning by Synergizing Graph and Text Mutual TransformationsabstractText-Attributed Graphs (TAGs) enhance graph structures with natural language descriptions, enabling detailed representation of data and their relationships across a broad spectrum of real-world scenarios. Despite the potential for deeper insights, existing TAG representation learning primarily omit the semantic relationship among node texts, and mostly relies on supervised methods, necessitating extensive labeled data and limiting applicability across diverse contexts. This paper introduces a new self-supervised learning framework, Text-Attributed-Graph Multi-View Alignment (TAGA), which overcomes these constraints by integrating TAGs' structural and semantic dimensions. TAGA constructs two complementary views: Text-of-Graph view, which organizes node texts into structured documents based on graph topology, and the Graph-of-Text view, which converts textual nodes and connections into graph data. By aligning representations from both views, TAGA captures joint textual and structural information. In addition, a novel structure-preserving random walk algorithm is proposed for efficient training on large-sized TAGs. Our framework demonstrates strong performance in zero-shot and few-shot scenarios across eight real-world datasets. Zheng Zhang 0047, Yuntong Hu, Bo Pan 0009, Chen Ling 0003, Liang Zhao 0002 |
CIKM | 2 |
| 2025 | CG-RAG: Research Question Answering by Citation Graph Retrieval-Augmented LLMsabstractResearch question answering requires accurate retrieval and contextual understanding of scientific literature. However, current Retrieval-Augmented Generation (RAG) methods often struggle to balance complex document relationships with precise information retrieval. In this paper, we introduce Contextualized Graph Retrieval-Augmented Generation (CG-RAG), a novel framework that integrates sparse and dense retrieval signals within graph structures to enhance retrieval efficiency and subsequently improve generation quality for research question answering. First, we propose a contextual graph representation for citation graphs, effectively capturing both explicit and implicit connections within and across documents. Next, we introduce Lexical-Semantic Graph Retrieval (LeSeGR), which seamlessly integrates sparse and dense retrieval signals with graph encoding. It bridges the gap between lexical precision and semantic understanding in citation graph retrieval, demonstrating generalizability to existing graph retrieval and hybrid retrieval methods. Finally, we present a context-aware generation strategy that utilizes the retrieved graph-structured information to generate precise and contextually enriched responses using large language models (LLMs). Extensive experiments on research question answering benchmarks across multiple domains demonstrate that our CG-RAG framework significantly outperforms RAG methods combined with various state-of-the-art retrieval approaches, delivering superior retrieval accuracy and generation quality. Yuntong Hu, Zhihan Lei, Zhongjie Dai, Allen Zhang 0005, Abhinav Angirekula, Zheng Zhang 0047, Liang Zhao 0002 |
SIGIR | 1 |
| 2024 | Distilling Large Language Models for Text-Attributed Graph LearningabstractText-Attributed Graphs (TAGs) are graphs of connected textual documents. Graph models can efficiently learn TAGs, but their training heavily relies on human-annotated labels, which are scarce or even unavailable in many applications. Large language models (LLMs) have recently demonstrated remarkable capabilities in few-shot and zero-shot TAG learning, but they suffer from scalability, cost, and privacy issues. Therefore, in this work, we focus on synergizing LLMs and graph models with their complementary strengths by distilling the power of LLMs into a local graph model on TAG learning. To address the inherent gaps between LLMs (generative models for texts) and graph models (discriminative models for graphs), we propose first to let LLMs teach an interpreter with rich rationale and then let a student model mimic the interpreter's reasoning without LLMs' rationale. We convert LLM's textual rationales to multi-level graph rationales to train the interpreter model and align the student model with the interpreter model based on the features of TAGs. Extensive experiments validate the efficacy of our proposed framework. Bo Pan 0009, Zheng Zhang 0047, Yifei Zhang 0006, Yuntong Hu, Liang Zhao 0002 |
CIKM | 4 |
| 2024 | Transferable Unsupervised Outlier Detection Framework for Human Semantic TrajectoriesabstractSemantic trajectories, which enrich spatial-temporal data with textual information such as trip purposes or location activities, are key for identifying outlier behaviors critical to healthcare, social security, and urban planning. Traditional outlier detection relies on heuristic rules, which requires domain knowledge and limits its ability to identify unseen outliers. Besides, there lacks a comprehensive approach that can jointly consider multi-modal data across spatial, temporal, and textual dimensions. Addressing the need for a domain-agnostic model, we propose the Transferable Outlier Detection for Human Semantic Trajectories (TOD4Traj) framework. TOD4Traj first introduces a modality feature unification module to align diverse data feature representations, enabling the integration of multi-modal information and enhancing transferability across different datasets. A contrastive learning module is further proposed for identifying regular mobility patterns both temporally and across populations, allowing for a joint detection of outliers based on individual consistency and group majority patterns. Our experimental results have shown TOD4Traj's superior performance over existing models, demonstrating its effectiveness and adaptability in detecting human trajectory outliers across various datasets. Zheng Zhang 0047, Dazhou Yu, Yuntong Hu, Liang Zhao 0002, Andreas Züfle |
SIGSPATIAL/GIS | 4 |
| 2024 | PolygonGNN: Representation Learning for Polygonal Geometries with Heterogeneous Visibility GraphabstractPolygon representation learning is essential for diverse applications, encompassing tasks such as shape coding, building pattern classification, and geographic question answering. While recent years have seen considerable advancements in this field, much of the focus has been on single polygons, overlooking the intricate inner- and inter-polygonal relationships inherent in multipolygons. To address this gap, our study introduces a comprehensive framework specifically designed for learning representations of polygonal geometries, particularly multipolygons. Central to our approach is the incorporation of a heterogeneous visibility graph, which seamlessly integrates both inner- and inter-polygonal relationships. To enhance computational efficiency and minimize graph redundancy, we implement a heterogeneous spanning tree sampling method. Additionally, we devise a rotation-translation invariant geometric representation, ensuring broader applicability across diverse scenarios. Finally, we introduce Multipolygon-GNN, a novel model tailored to leverage the spatial and semantic heterogeneity inherent in the visibility graph. Experiments on five real-world and synthetic datasets demonstrate its ability to capture informative representations for polygonal geometries. Dazhou Yu, Yuntong Hu, Yun Li 0005, Liang Zhao 0002 |
KDD | 2 |
| 2024 | TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge GraphsabstractText-Attributed Graphs (TAGs) augment graph structures with natural language descriptions, facilitating detailed depictions of data and their interconnections across various real-world settings. However, existing TAG datasets predominantly feature textual information only at the nodes, with edges typically represented by mere binary or categorical attributes. This lack of rich textual edge annotations significantly limits the exploration of contextual relationships between entities, hindering deeper insights into graph-structured data. To address this gap, we introduce Textual-Edge Graphs Datasets and Benchmark (TEG-DB), a comprehensive and diverse collection of benchmark textual-edge datasets featuring rich textual descriptions on nodes and edges. The TEG-DB datasets are large-scale and encompass a wide range of domains, from citation networks to social networks. In addition, we conduct extensive benchmark experiments on TEG-DB to assess the extent to which current techniques, including pre-trained language models, graph neural networks, and their combinations, can utilize textual node and edge information. Our goal is to elicit advancements in textual-edge graph research, specifically in developing methodologies that exploit rich textual node and edge descriptions to enhance graph analysis and provide deeper insights into complex real-world networks. The entire TEG-DB project is publicly accessible as an open-source repository on Github, accessible at https://github.com/Zhuofeng-Li/TEG-Benchmark. Zhuofeng Li, Zixing Gou, Xiangnan Zhang, Zhongyuan Liu, Yuntong Hu, Chen Ling 0003, Zheng Zhang 0047, Liang Zhao 0002 |
NeurIPS | 6 |
| 2023 | Time-Series Forecasting Based on Fuzzy Cognitive Visibility Graph and Weighted Multisubgraph SimilarityabstractThis article aims to address the problem of time-series forecasting. Current state-of-the-art forecasting models lack the ability to mine the spatiotemporal dependence. How to mine more useable features of time series to make accuracy predictions is still an open issue. To address these challenges, from the perspective of fuzzy interaction between nodes, we propose a novel network constructing model called fuzzy cognitive visibility graph (FCVG) for time series to convert the time series into a pair of directed weighted graphs. To calculate the similarity between nodes in the FCVG, we develop the weighted multisubgraph similarity (WMSS). With these tools, we introduce the prediction based on fuzzy similarity distribution (PFSD), a novel forecasting method for time series, that can efficiently capture the spatiotemporal dependence. A time series is converted into a network by the FCVG, and the similarity scores between nodes are calculated through the WMSS. Based on the normalized similarity distribution, the predictions of time series are made. Extensive experiments on different datasets confirm the benefits of leveraging fuzzy interaction in time-series forecasting. Moreover, the construction cost index is predicted to show how to apply PFSD to forecast a specific time series. Yuntong Hu, Fuyuan Xiao 0001 |
IEEE Trans. Fuzzy Syst. | 1 |