VLDB 2026 Research / reviewers in the wild / expert
Yongting Hu
dblp:336/4705
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-3924-3547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease DiagnosisabstractMultimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness. Haoran Li 0024, Haoyu Cao 0002, Yongting Hu, Qihao Xu, Chengliang Liu 0003, Xiaoling Luo 0001, Zhihao Wu 0002, Yong Xu 0001, Wei Wang 0169 |
AAAI | 4 |
| 2026 | Towards Zero-Shot Diabetic Retinopathy Grading: Learning Generalized Knowledge via Prompt-Driven Matching and EmulatingabstractAs one of the primary causes of visual impairment, Diabetic Retinopathy (DR) requires accurate and robust grading to facilitate timely diagnosis and intervention. Different from conventional DR grading methods that utilize single-view images, recent clinical studies have revealed that multi-view fundus images can significantly enhance DR grading performance by expanding the field of view (FOV). However, there is a long-tailed distribution problem in fundus image analysis, i.e., a high prevalence of mild DR grades and a low prevalence of rare ones (e.g., cases of high severity), which presents a significant challenge to developing a unified model capable of detecting rare or unseen DR grades not encountered during training. In this paper, we propose ProME-DR, a Prompt-driven zero-shot DR grading framework, which leverages prompt Matching and Emulating to recognize the unseen DR categories and views beyond the training set. ProME-DR disentangles the training process into two stages to learn generalized knowledge for novel DR disease grading. Initially, ProME-DR leverages two sets of prompt units to capture semantic and inter-view consistency knowledge via a split-and-mask manner, gathering instance-level DR visual clues. Subsequently, it constructs a concept-aware emulator to generate context prompt units, linking extensible knowledge learned from the previously seen DR attributes for zero-shot DR grading. Extensive experiments conducted on eight datasets and various scenarios confirm the superiority of ProME-DR. Haoran Li 0024, Huaming Chen, Jun Yan 0005, Jiahua Shi, Qihao Xu, Yongting Hu, Yong Xu 0001, Jun Shen 0001 |
AAAI | 9 |
| 2026 | Vision-Language Models Guided Graph Concept Reasoning for Interpretable Diabetic Retinopathy DiagnosisabstractDeep neural networks (DNNs) have significantly advanced diabetic retinopathy (DR) diagnosis, yet their black-box nature limits clinical acceptance due to a lack of interpretability. Concept bottleneck model (CBM) offers a promising solution by enabling concept-level reasoning and test-time intervention, with recent DR studies modeling lesions as concepts and grades as outcomes. However, current methods often ignore relationships between lesion concepts across different DR grades and struggle when fine-grained lesion concepts are unavailable, limiting their interpretability and real-world applicability. To bridge these gaps, we propose VLM-GCR, a vision-language model guided graph concept reasoning framework for interpretable DR diagnosis. VLM-GCR emulates the diagnostic process of ophthalmologists by constructing a grading-aware lesion concept graph that explicitly models the interactions among lesions and their relationships to disease grades. In concept-free clinical scenarios, our method introduces a vision-language guided dynamic concept pseudo-labeling mechanism to mitigate the challenges of existing concept-based models in fine-grained lesion recognition. Additionally, we introduce a multi-level intervention method that supports error correction, enabling transparent and robust human-AI collaboration. Experiments on two public DR benchmarks show that VLM-GCR achieves strong performance in both lesion and grading tasks, while delivering clear and clinically meaningful reasoning steps. Qihao Xu, Xiaoling Luo 0001, Chengliang Liu 0003, Yongting Hu, Xinheng Lyu, Yong Xu 0001 |
AAAI | 5 |
| 2026 | A CNN-injected transformer network with lesion reconstruction for multi-view diabetic retinopathy grading
Yongting Hu, Haoran Li 0024, Qihao Xu, Jiahua Shi, Jun Shen 0001, Yong Xu 0001, Xiaoyan Dou |
Medical Image Anal. | 1 |
| 2025 | Wavelet-based Global-Local Interaction Network with Cross-Attention for Multi-View Diabetic Retinopathy DetectionabstractMulti-view diabetic retinopathy (DR) detection has recently emerged as a promising method to address the issue of incomplete lesions faced by single-view DR. However, it is still challenging due to the variable sizes and scattered locations of lesions. Furthermore, existing multi-view DR methods typically merge multiple views without considering the correlations and redundancies of lesion information across them. Therefore, we propose a novel method to overcome the challenges of difficult lesion information learning and inadequate multi-view fusion. Specifically, we introduce a two-branch network to obtain both local lesion features and their global dependencies. The high-frequency component of the wavelet transform is used to exploit lesion edge information, which is then enhanced by global semantic to facilitate difficult lesion learning. Additionally, we present a cross-view fusion module to improve multi-view fusion and reduce redundancy. Experimental results on large public datasets demonstrate the effectiveness of our method. The code is open sourced on https://github.com/HuYongting/WGLIN. Yongting Hu, Chengliang Liu 0003, Xiaoling Luo 0001, Xiaoyan Dou, Qihao Xu, Yong Xu 0001 |
ICME | 1 |
| 2025 | Learning multiscale residual prototypes and global-local correspondence for video anomaly detection
Yongting Hu, Yuanhong Zhong |
Comput. Vis. Image Underst. | 1 |
| 2025 | A Two-Stage Framework With Memory for Anomaly Detection via Video Decomposition and Bidirectional ConsistencyabstractExisting unsupervised video anomaly detection methods based on prediction typically employ a memory module to limit the generalization ability of the network so that normal frames can be accurately reconstructed or predicted while abnormal frames cannot. These memory-based methods usually utilize memory to record the fusion prototypes of appearance and motion. However, the motion part of the fusion prototypes is obtained from implicit motion representation, which is incomplete and constrain the ability for abnormal detection. To tackle the above issue, we proposed a Two-Stage framework with Memory via video Decomposition and bidirectional Consistency (TSMDC), which employs explicit motion data to learn comprehensive motion prototypes and use video decomposition and bidirectional consistency to learn fine granularity and advanced prototypes. We first decompose the video clip into three components: motion, scene, and object. In the first stage, the features of the motion are extracted to obtain a comprehensive motion representation, and its prototypes are stored in the motion memory. For further use of the bidirectional information of video, we present a cascaded frame prediction network that is utilized to learn and record the advanced spatio-temporal prototypes in the second stage. Specifically, the fine-granularity features of the scene and object are extracted and fused to predict the future frame. Then the initial frame of the video clip is predicted based on bidirectional consistency and motion prototypes enhancement. And the advanced spatio-temporal prototypes of video are recorded in this process. Anomalies are evaluated using the combination anomaly score of the predicted future and initial frame.Extensive experimental results on three public datasets indicate the effectiveness of the proposed method. Code will be available at https://github.com/yangugu/TSMDC. Yuanhong Zhong, Yongting Hu, Ruyue Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Associative Memory With Spatio-Temporal Enhancement for Video Anomaly DetectionabstractMemory network has been extensively used to record prototypical normal patterns to prevent overgeneralization of the network to reconstruct anomalies for video anomaly detection. However, existing memory-based methods only record the lossy representation of normal item prototypes, without recording the rich relationships between them. In this work, we propose an Associative Memory with Spatio-Temporal Enhancement (AMSTE) which introduces the global context information constraint of motion to enhance the appearance features and learn the normal item prototypes and their relationship. Specifically, we utilize two encoders to extract spatio-temporal features with the Spatio-Temporal Enhancement Module (STEM) to enhance appearance features with global motion constraints. Then, the prototypical patterns of normal data and their relationships are recorded in the item memory and relational memory, respectively. Finally, we retrieve features from the memory pools and reconstruct the video frame through the decoder. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our approach. The code will be released athttps://github.com/HuYongting/AMSTE. Yuanhong Zhong, Yongting Hu, Panliang Tang, Heng Wang 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Bidirectional Spatio-Temporal Feature Learning With Multiscale Evaluation for Video Anomaly DetectionabstractVideo anomaly detection aims to detect the segments containing abnormal events from video sequence, which is a current research hotspot due to the importance in maintaining social security. Recent detection methods tend to build frame reconstruction or frame prediction model based on deep learning to learn features of events. The reconstruction-based methods reproduce the input frame one-to-one, inevitably losing some temporal features. The prediction-based methods predict frames according to the natural time order, but ignore the reverse time information, causing the deviation in information learning. Besides, anomaly evaluation methods based on patch-level error neglect the diversity of object sizes in complex scenes, and it is difficult to determine the optimal size of the error patch accurately. For these issues, we propose a bidirectional spatio-temporal feature learning framework with multi-scale anomaly evaluation strategy. A video sequence is input to a double-encoder double-decoder network, and bidirectional spatio-temporal features are obtained for bidirectional prediction by fusing forward and backward features extracted from the two encoders. The multi-scale anomaly evaluation method is implemented based on error pyramid and mean pooling, which effectively detects target objects with different sizes. Experiments on several publicly video datasets show that our method outperforms most of existing methods. Yuanhong Zhong, Yongting Hu, Panliang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |