EDBT 2026 Demo / reviewers in the wild / expert
Haodi Zhong
dblp:258/9015
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-6033-4959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias CorrectionabstractDynamic graph-level embedding aims to capture structural evolution in networks, which is essential for modeling real-world scenarios. However, existing methods face two critical yet under-explored issues: Structural Visit Bias, where random walk sampling disproportionately emphasizes high-degree nodes, leading to redundant and noisy structural representations; and Abrupt Evolution Blindness, the failure to effectively detect sudden structural changes due to rigid or overly simplistic temporal modeling strategies, resulting in inconsistent temporal embeddings. To overcome these challenges, we propose EvoFormer, an evolution-aware Transformer framework tailored for dynamic graph-level representation learning. To mitigate Structural Visit Bias, EvoFormer introduces a Structure-Aware Transformer Module that incorporates positional encoding based on node structural roles, allowing the model to globally differentiate and accurately represent node structures. To overcome Abrupt Evolution Blindness, EvoFormer employs an Evolution-Sensitive Temporal Module, which explicitly models temporal evolution through a sequential three-step strategy: (I) Random Walk Timestamp Classification, generating initial timestamp-aware graph-level embeddings; (II) Graph-Level Temporal Segmentation, partitioning the graph stream into segments reflecting structurally coherent periods; and (III) Segment-Aware Temporal Self-Attention combined with an Edge Evolution Prediction task, enabling the model to precisely capture segment boundaries and perceive structural evolution trends, effectively adapting to rapid temporal shifts. Extensive evaluations on five benchmark datasets confirm that EvoFormer achieves state-of-the-art performance in graph similarity ranking, temporal anomaly detection, and temporal segmentation tasks, validating its effectiveness in correcting structural and temporal biases. Code is available at https://github.com/zlx0823/EvoFormerCode. Haodi Zhong, Liuxin Zou, Di Wang 0011, Bo Wan 0002, Zhenxing Niu, Quan Wang 0006 |
CIKM | 1 |
| 2025 | Scene-enhanced multi-scale temporal aware network for video moment retrieval
Di Wang 0011, Yousheng Yu, Haodi Zhong, Lin Zhao 0003 |
Pattern Recognit. | 4 |
| 2024 | Leveraging Coarse-to-Fine Grained Representations in Contrastive Learning for Differential Medical Visual Question Answering
Di Wang 0011, Zhicheng Jiao, Haodi Zhong, Mengyu Yang, Quan Wang 0006 |
MICCAI (5) | 5 |
| 2024 | Fine-grained Semantics-aware Representation Learning for Text-based Person RetrievalabstractText-based person retrieval aims to search for target persons based on a given text description query. However, existing methods often have the following problems: (1) Ignoring local attribute information between different persons in feature learning, which results in the low distinguishability of similar people's feature representations. (2) Lacking fine-grained semantics alignment between visual images and text descriptions, which leads to inconsistency in person details between query and target. To address these issues, we propose a Fine-grained Semantics-aware Representation Learning (FSRL) method that establishing intra-modal local attribute correlations and inter-modal fine-grained semantic correlations. Specifically, we first design an identity self-distillation module, which explores soft identity labels that reflect local attribute similarities among different people. The soft identity labels assist the model in learning discriminative features associated with fine-grained attributes of persons. Secondly, we propose a visual-language relationship modeling module that enforces the model to proofread "error words" randomly changed in text during the cross-modal interaction process to establish fine-grained image-text semantic correlations. Extensive experiments show that the proposed method achieves new state-of-the-art results on three benchmark datasets and also performs well on the domain generalization task. Our code is available at https://github.com/y416f/FSRL. Di Wang 0011, Yifeng Wang 0004, Lin Zhao 0003, Haodi Zhong |
ICMR | 6 |
| 2024 | Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report GenerationabstractRadiology report generation aims to automatically generate clinical descriptions for radiology images, reducing the workload of radiologists. Compared to general image captioning tasks, the subtle differences in medical images and the specialized, complex nature of medical terminology limit the performance of data-driven radiology report generation. Previous research has attempted to leverage prior knowledge, such as organ-disease graphs, to enhance models' abilities to identify specific diseases and generate corresponding medical terminology. However, these methods cover only a limited number of disease types, focusing solely on disease terms mentioned in reports but ignoring their normal or abnormal attributes, which are critical to generating accurate reports. To address this issue, we propose a Divide-and-Conquer approach, named DCG, which separately constructs disease-free and disease-specific nodes within the knowledge graphs. Specifically, we extracted more comprehensive organ-disease entities from reports than previous methods and constructed disease-free and disease-specific nodes by rigorously distinguishing between normal conditions and specific diseases. This enables our model to consciously focus on abnormal information and mitigate the impact of excessively common diseases on report generation. Subsequently, the constructed graph is utilized to enhance the correlation between visual representations and disease terminology, thereby guiding the decoder in report generation. Extensive experiments conducted on benchmark datasets IU-Xray and MIMIC-CXR demonstrate the superiority of our proposed method. Code is available at https://github.com/ecoxial2007/DCG_Enhanced_distilGPT2. Yanlei Zhang, Di Wang 0011, Haodi Zhong, Ronghan Li, Quan Wang 0006 |
ACM Multimedia | 4 |
| 2024 | Candidate-Heuristic In-Context Learning: A new framework for enhancing medical visual question answering with LLMs
Di Wang 0011, Haodi Zhong, Quan Wang 0006, Ronghan Li, Rui Jia, Bo Wan 0002 |
Inf. Process. Manag. | 3 |
| 2024 | DiagSWin: A multi-scale vision transformer with diagonal-shaped windows for object detection and segmentation
Ke Li 0024, Di Wang 0011, Gang Liu 0006, Wenxuan Zhu, Haodi Zhong, Quan Wang 0006 |
Neural Networks | 5 |
| 2024 | Language-Guided Progressive Attention for Visual Grounding in Remote Sensing ImagesabstractVisual grounding in remote sensing (RSVG) images aims to detect specific objects associated with referring expressions in remote sensing images. Existing methods typically combine outputs of pretrained visual and linguistic backbones to locate referred objects. However, due to the lack of interaction with the language modality during the visual feature extraction process, the visual backbone may suffer from attention drift, limiting RSVG’s performance. To avoid this, we propose a novel RSVG framework, namely, language-guided progressive visual attention (LPVA), which achieves precise attention on referred objects by adjusting visual features with a progressive attention (PA) module and a multilevel feature enhancement (MFE) decoder. Specifically, the former can dynamically generate multiscale weights and biases, enabling the visual backbone to gradually focus on expression-related features at spatial and channel levels. The latter is designed to aggregate visual contextual information of the referred objects to enhance features’ distinctiveness while simultaneously suppressing information of irrelevant regions. To thoroughly examine the localization capability of RSVG models, we construct a new large-scale benchmark dataset, namely, OPT-RSVG, which poses challenges in comprehensive understanding among complex scenarios. Experimental results show that the proposed method pushes the accuracy score to 82.27% (6.29% absolute improvement) on the DIOR-RSVG dataset and 78.03% on the OPT-RSVG dataset, thus setting new records. The source codes of the proposed method and OPT-RSVG dataset are available athttps://github.com/like413/OPT-RSVG. Ke Li 0024, Di Wang 0011, Haodi Zhong, Cong Wang 0033 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Ego-Network Segmentation via (Weighted) Jaccard MedianabstractAn ego-network is a graph representing the interactions of a node (ego) with its neighbors and the interactions among those neighbors. A sequence of ego-networks having the same ego can thus model the evolution of these interactions over time. We introduce the problem of segmenting a sequence of ego-networks into$k$segments, for any given integer$k$. Each segment is represented by a summary network, and the goal is to minimize the total loss of representing$k$segments by$k$summaries. The problem allows partitioning the sequence into homogeneous segments with respect to the activities or properties of the ego (e.g., to identify time periods when a user acquired different circles of friends in a social network) and to compactly represent each segment with a summary. The main challenge is to construct a summary that represents a collection of ego-networks with minimum loss. To address this challenge, we employ Jaccard Median (JM), a well-known NP-hard problem for summarizing sets, for which, however, no effective and efficient algorithms are known. We develop a series of algorithms for JM offering different effectiveness/efficiency trade-offs: (I) an exact exponential-time algorithm, based on Mixed Integer Linear Programming; (II) exact and approximation polynomial-time algorithms for minimizing an upper bound of the objective function of JM; and (III) efficient heuristics for JM, which are based on an effective scoring scheme and one of them also on sketching. We also study a generalization of the segmentation problem, in which there may be multiple edges between a pair of nodes in an ego-network. To tackle this problem, we develop a series of algorithms, based on a more general problem than JM, called Weighted Jaccard Median WJM: (I) an exact exponential-time algorithm, based on Mixed Integer Linear Programming; (II) exact algorithms for minimizing an upper bound of the objective function of WJM; and (III) efficient heuristics, based on the percentiles of edge multiplicities and one of them also on divide-and-conquer. By building upon the above results, we design algorithms for segmenting a sequence of ego-networks. Experiments with 10 real datasets and with synthetic datasets show that our algorithms produce optimal or near-optimal solutions to JM or to WJM, and that they substantially outperform state-of-the-art methods which can be employed for ego-network segmentation. Haodi Zhong, Grigorios Loukides, Alessio Conte, Solon P. Pissis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Jaccard Median for Ego-network SegmentationabstractAn ego-network is a graph representing the interactions of a node (ego) with its neighbors and the interactions among those neighbors. A sequence of ego-networks having the same ego can thus model the evolution of these interactions over time. We introduce the problem of segmenting a sequence of ego-networks into k segments, for any given integer k. Each segment is represented by a summary network, and the goal is to minimize the total loss of representing k segments by k summaries. The problem allows partitioning the sequence into homogeneous segments with respect to the activities or properties of the ego (e.g., to identify time periods when a user acquired different circles of friends in a social network) and to compactly represent each segment with a summary. The main challenge is to construct a summary that represents a collection of ego-networks with minimum loss. To address this challenge, we employ Jaccard Median (JM), a well-known NP-hard problem for summarizing sets, for which, however, no effective and efficient algorithms are known. We develop a series of algorithms for JM offering different effectiveness/efficiency trade-offs: (I) an exact exponential-time algorithm, based on Mixed Integer Linear Programming and (II) exact and approximation polynomial-time algorithms for minimizing an upper bound of the objective function of JM. By building upon these results, we design two algorithms for segmenting a sequence of ego-networks that are effective, as shown experimentally. Haodi Zhong, Grigorios Loukides, Alessio Conte, Solon P. Pissis |
ICDM | 1 |
| 2022 | Clustering sequence graphsabstractIn application domains ranging from social networks to e-commerce, it is important to cluster users with respect to both their relationships (e.g., friendship or trust) and their actions (e.g., visited locations or rated products). Motivated by these applications, we introduce here the task of clustering the nodes of a sequence graph, i.e., a graph whose nodes are labeled with strings (e.g., sequences of users’ visited locations or rated products). Both string clustering algorithms and graph clustering algorithms are inappropriate to deal with this task, as they do not consider the structure of strings and graph simultaneously. Moreover, attributed graph clustering algorithms generally construct poor solutions because they need to represent a string as a vector of attributes, which inevitably loses information and may harm clustering quality. We thus introduce the problem of clustering a sequence graph. We first propose two pairwise distance measures for sequence graphs, one based on edit distance and shortest path distance and another one based on SimRank. We then formalize the problem under each measure, showing also that it is NP-hard. In addition, we design a polynomial-time 2-approximation algorithm, as well as a heuristic for the problem. Experiments using real datasets and a case study demonstrate the effectiveness and efficiency of our methods. Haodi Zhong, Grigorios Loukides, Solon P. Pissis |
Data Knowl. Eng. | 1 |
| 2022 | Clustering Demographics and Sequences of Diagnosis CodesabstractA Relational-Sequential dataset (or RS-dataset for short) contains records comprised of a patient's values in demographic attributes and their sequence of diagnosis codes. The task of clustering an RS-dataset is helpful for analyses ranging from pattern mining to classification. However, existing methods are not appropriate to perform this task. Thus, we initiate a study of how an RS-dataset can be clustered effectively and efficiently. We formalize the task of clustering an RS-dataset as an optimization problem. At the heart of the problem is a distance measure we design to quantify the pairwise similarity between records of an RS-dataset. Our measure uses a tree structure that encodes hierarchical relationships between records, based on their demographics, as well as an edit-distance-like measure that captures both the sequentiality and the semantic similarity of diagnosis codes. We also develop an algorithm which first identifies k representative records (centers), for a given k, and then constructs k clusters, each containing one center and the records that are closer to the center compared to other centers. Experiments using two Electronic Health Record datasets demonstrate that our algorithm constructs compact and well-separated clusters, which preserve meaningful relationships between demographics and sequences of diagnosis codes, while being efficient and scalable. Haodi Zhong, Grigorios Loukides, Solon P. Pissis |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Clustering datasets with demographics and diagnosis codes
Haodi Zhong, Grigorios Loukides, Robert Gwadera |
J. Biomed. Informatics | 1 |