EDBT 2026 Demo / reviewers in the wild / expert
Yufan Hu
dblp:186/9205
· DBLP profile ↗
20ranked-venue papers
11as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Group Orthogonal Low-Rank Adaptation for RGB-T TrackingabstractParameter-efficient fine-tuning has emerged as a promising paradigm in RGB-T tracking, enabling downstream task adaptation by freezing pretrained parameters and fine-tuning only a small set of parameters. This set forms a rank space made up of multiple individual ranks, whose expressiveness directly shapes the model's adaptability. However, quantitative analysis reveals low-rank adaptation exhibits significant redundancy in the rank space, with many ranks contributing almost no practical information. This hinders the model's ability to learn more diverse knowledge to address the various challenges in RGB-T tracking. To address this issue, we propose the Group Orthogonal Low-Rank Adaptation (GOLA) framework for RGB-T tracking, which effectively leverages the rank space through structured parameter learning. Specifically, we adopt a rank decomposition partitioning strategy utilizing singular value decomposition to quantify rank importance, freeze crucial ranks to preserve the pretrained priors, and cluster the redundant ranks into groups to prepare for subsequent orthogonal constraints. We further design an inter-group orthogonal constraint strategy. This constraint enforces orthogonality between rank groups, compelling them to learn complementary features that target diverse challenges, thereby alleviating information redundancy. Experimental results demonstrate that GOLA effectively reduces parameter redundancy and enhances feature representation capabilities, significantly outperforming state-of-the-art methods across four benchmark datasets and validating its effectiveness in RGB-T tracking tasks. Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001 |
AAAI | 2 |
| 2026 | Later is More: Trading Tolerable Latency to Meet Stringent Jitter Requirement in Time-Sensitive Networking
Shaodong Huang, Jiawei Huang 0001, Qichen Su, Yufan Hu, Xiaojuan Lu |
IWQoS | 8 |
| 2026 | Intra-inter modality learning for copper electrolysis short circuit detection
Yufan Hu |
Adv. Eng. Informatics | 5 |
| 2026 | InMoE: Interaction-aware graph mixture of experts for trajectory prediction
Hong Liu 0002, Yufan Hu |
Pattern Recognit. | 2 |
| 2026 | History-Guided Prompt Generation for Vision-and-Language NavigationabstractVision-and-language navigation (VLN) has garnered extensive attention in the field of embodied artificial intelligence. VLN involves time series information, where historical observations contain rich contextual knowledge and play a crucial role in navigation. However, current methods do not explicitly excavate the connection between rich contextual information in history and the current environment, and ignore adaptive learning of clues related to the current environment. Therefore, we explore a Prompt Learning-based strategy which adaptively mines information in history that is highly relevant to the current environment to enhance the agent's perception of the current environment and propose a history-guided prompt generation (HGPG) framework. Specifically, HGPG includes two parts, one is an entropy-based history acquisition module that assesses the uncertainty of the action probability distribution from the preceding step to determine whether historical information should be used at the current time step. The other part is the prompt generation module that transforms historical context into prompt vectors by sampling from an end-to-end learned token library. These prompt tokens serve as discrete, knowledge-rich representations that encode semantic cues from historical observations in a compact form, making them easier for the decision network to understand and utilize. In addition, we share the token library across various navigation tasks, mining common features between different tasks to improve generalization to unknown environments. Extensive experimental results on four mainstream VLN benchmarks (R2R, REVERIE, SOON, R2R-CE) demonstrate the effectiveness of our proposed method. Code is available at https://github.com/Wzmshdong/HGPG. Wen Guo 0003, Zongmeng Wang, Yufan Hu, Junyu Gao 0002 |
IEEE Trans. Cybern. | 3 |
| 2025 | Unity in Diversity: Video Editing via Gradient-Latent PurificationabstractRecently, text-driven video editing methods that optimize target latent representations have garnered significant attention and demonstrated promising results. However, these methods rely on self-supervised objectives to compute the gradients needed for updating latent representations, which inevitably introduces gradient noise, compromising content generation quality. Additionally, it is challenging to determine the optimal stopping point for the editing process, making it difficult to achieve an optimal solution for the latent representation. To address these issues, we propose a unified gradient-latent purification framework that collects gradient and latent information across different stages to identify effective and concordant update directions. We design a local coordinate system construction method based on feature decomposition, enabling short-term gradients and final-stage latents to be reprojected onto new axes. Then, we employ tailored coefficient regularization terms to effectively aggregate the decomposed information. Additionally, a temporal smoothing axis extension strategy is developed to enhance the temporal coherence of the generated content. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art methods across various editing tasks, delivering superior editing performance. Project page is available in https://unityin-diversity-editing.github.io. Junyu Gao 0001, Yufan Hu |
CVPR | 4 |
| 2025 | PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T TrackingabstractMaintaining stable tracking of objects in domain shift scenarios is crucial for RGB-T tracking, prompting us to explore the use of unlabeled test sample information for effective online model adaptation. However, current Test-Time Adaptation (TTA) methods in RGB-T tracking dramatically change the model’s internal parameters during long-term adaptation. At the same time, the gradient computations involved in the optimization process impose a significant computational burden. To address these challenges, we propose a Parameter Update-Recovery Adaptation (PURA) framework based on parameter decomposition. Firstly, our fast parameter update strategy adjusts model parameters using statistical information from test samples without requiring gradient calculations, ensuring consistency between the model and test data distribution. Secondly, our parameter decomposition recovery employs orthogonal decomposition to identify the principal update direction and recover parameters in this direction, aiding in the retention of critical knowledge. Finally, we leverage the information obtained from decomposition to provide feedback on the momentum during the update phase, ensuring a stable updating process. Experimental results demonstrate that PURA outperforms current state-of-the-art methods across multiple datasets, validating its effectiveness. The project page is at https://melantech.github.io/PURA. Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001 |
CVPR | 2 |
| 2025 | Synthesizing Realistic fMRI: A Physiological Dynamics-Driven Hierarchical Diffusion Model for Efficient fMRI AcquisitionabstractFunctional magnetic resonance imaging (fMRI) is essential for mapping brain activity but faces challenges like lengthy acquisition time and sensitivity to patient movement, limiting its clinical and machine learning applications. While generative models such as diffusion models can synthesize fMRI signals to alleviate these issues, they often underperform due to neglecting the brain's complex structural and dynamic properties.
To address these limitations, we propose the Physiological Dynamics-Driven Hierarchical Diffusion Model, a novel framework integrating two key brain physiological properties into the diffusion process: brain hierarchical regional interactions and multifractal dynamics.
To model complex interactions among brain regions, we construct hypergraphs based on the prior knowledge of brain functional parcellation reflected by resting-state functional connectivity (rsFC). This enables the aggregation of fMRI signals across multiple scales and generates hierarchical signals.
Additionally, by incorporating the prediction of two key dynamics properties of fMRI—the multifractal spectrum and generalized Hurst exponent—our framework effectively guides the diffusion process, ensuring the preservation of the scale-invariant characteristics inherent in real fMRI data.
Our framework employs progressive diffusion generation, with signals representing broader brain region information conditioning those that capture localized details, and unifies multiple inputs during denoising for balanced integration.
Experiments demonstrate that our model generates physiologically realistic fMRI signals, potentially reducing acquisition time and enhancing data quality, benefiting clinical diagnostics and machine learning in neuroscience. Yufan Hu, Yu Jiang 0013, Wuyang Li, Yixuan Yuan |
ICLR | 1 |
| 2025 | Learning Evidential Delta Denoising Scores for Video Editing
Yufan Hu, Junyu Gao 0002, Bin Fan 0001, Hongmin Liu 0001 |
ACM Multimedia | 1 |
| 2025 | Learning semantic-unified cross-modal representations for open-vocabulary video scene graph generation
Yufan Hu, Junling Gao |
Multim. Syst. | 1 |
| 2025 | Long-tailed video recognition via majority-guided diffusion model
Yufan Hu |
Multim. Syst. | 1 |
| 2025 | Learning Boundary Continuity-Aware Gaussian Encoder for Oriented Object DetectionabstractOriented object detection has been crucial for rotation-sensitive tasks and has garnered significant attention. Most existing methods generate angles as detector output vectors, but this strategy can abnormally magnify visually similar differences between two boxes in certain circumstances, termed boundary discontinuity issue. To overcome this limitation, we propose a boundary continuity-aware Gaussian encoder (BCGE). Specifically, BCGE directly predicts target Gaussian distributions for proposals and learns an oriented bounding box as an integrated 2-D matrix, effectively addressing boundary discontinuity issues. We also propose a transformation from Gaussian representation back to boxes and extend this transformation theory to the complex domain to adapt to the learning characteristics of neural networks. Furthermore, BCGE serves as a versatile plug-and-play architectural encoder, directly replacing the standard coding process in various oriented detectors with adaptability. Experimental results on five popular datasets, i.e., DOTA, UCAS-AOD, HRSC2016, SSDD, and HRSID, consistently show the effectiveness of our approach. Hongmin Liu 0001, Chengyi Zhao, Bin Fan 0001, Yufan Hu |
IEEE Trans. Cybern. | 5 |
| 2025 | Dual-Level Modality De-Biasing for RGB-T TrackingabstractRGB-T tracking aims to effectively leverage the complement ability of visual (RGB) and infrared (TIR) modalities to achieve robust tracking performance in various scenarios. Existing RGB-T tracking methods typically adopt backbone networks pre-trained on large-scale RGB datasets, which can lead to a predisposition toward RGB image patterns. RGB and TIR modalities also exhibit inconsistent responses to regions with diverse properties, resulting in imbalances in tracking decisions. We refer to these issues as feature-level and decision-level biases in the TIR modality. In this paper, we propose a novel dual-level modality de-biasing framework for RGB-T tracking to eliminate the inherent feature and decision-level biases. Specifically, we propose a joint infrared-fusion adapter, comprising an infrared-aware adapter and a cross-fusion adapter, designed to adaptively mitigate feature-level biases and utilize complementary information between the two modalities. In addition to implicit feature-level adjustment, we propose a response-decoupled distillation strategy to explicitly alleviate decision-level biases, aiming to achieve consistently accurate decision-making between the RGB and TIR modalities. Extensive experiments on several popular RGB-T tracking benchmarks validate the effectiveness of our proposed method. Yufan Hu, Zekai Shao 0002, Bin Fan 0001, Hongmin Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Unsupervised face image deblurring via disentangled representation learning
Yufan Hu, Junyong Xia |
Pattern Recognit. Lett. | 1 |
| 2024 | Learning Proposal-Aware Re-Ranking for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) aims to localize and classify action instances in untrimmed videos with only video-level labels available. Despite the remarkable success of existing methods, whose generated proposals are commonly far more than the ground-truth action instances, it still makes sense to improve the ranking accuracy of the generated proposals since users in real-world scenarios usually prioritize the action proposals with the highest confidence scores. The inaccuracy of the proposal ranking mainly comes from two aspects: For one thing, the traditional proposal generation manner entirely relies on snippet-level perception, resulting in a significant yet unnoticed gap with the target of proposal-level localization. For another, existing methods commonly employ a hand-crafted proposal generation manner, a post-process that does not participate in model optimization. To address the above issues, we propose an end-to-end trained two-stage method, termed as Learning Proposal-aware Re-ranking (LPR) for WTAL. In the first stage, we design a proposal-aware feature learning module to inject the proposal-aware contextual information into each snippet, and then the enhanced features are utilized for predicting initial proposals. Furthermore, to perform effective and efficient proposal re-ranking, in the second stage, we contrast the proposals attached with high confidence scores with our constructed multi-scale foreground/background prototypes for further optimization. Evaluated by both the vanilla and Top-$k$mAP metrics, results of extensive experiments on two popular benchmarks demonstrate the effectiveness of our proposed method. Yufan Hu, Jie Fu 0004, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Exploring Rich Semantics for Open-Set Action RecognitionabstractOpen-set action recognition (OSAR) aims to learn a recognition framework capable of both classifying known classes and identifying unknown actions in open-set scenarios. Existing OSAR methods typically reside in a data-driven paradigm, which ignore the rich semantics in both known and unknown categories. In fact, we humans have the capability of leveraging the captured semantic information, i.e., knowledge and experience, to incisively distinguish samples from known and unknown classes. Motivated by this observation, in this paper, we propose a Unified Semantic Exploration (USE) framework for recognizing actions in open-set scenarios. Specifically, we explore the explicit knowledge semantics by simulating the unknown classes with knowledge-guided virtual classes based on an external knowledge graph, which enables the model to simulate open-set perception during model training. Besides, we propose to learn the implicit data semantics by transferring the knowledge structure of action categories to the visual prototype space for semantic structure preservation. Extensive experiments on several action recognition benchmarks validate the effectiveness of our proposed method. Yufan Hu, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Learning Multi-Expert Distribution Calibration for Long-Tailed Video ClassificationabstractMost existing state-of-the-art video classification methods assume that the training data obey a uniform distribution. However, video data in the real world typically exhibit an imbalanced long-tailed class distribution, resulting in a model bias towards head class and relatively low performance on tail class. While the current long-tailed classification methods usually focus on image classification, adapting them to video data is not a trivial extension. We propose an end-to-end multi-expert distribution calibration method to address these challenges based on two-level distribution information. The method jointly considers the distribution of samples in each class (intra-class distribution) and the overall distribution of diverse data (inter-class distribution) to solve the issue of imbalanced data under long-tailed distribution. By modeling the two-level distribution information, the model can jointly consider the head classes and the tail classes and significantly transfer the knowledge from the head classes to improve the performance of the tail classes. Extensive experiments verify that our method achieves state-of-the-art performance on the long-tailed video classification task. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2023 | Learning Scene-Aware Spatio-Temporal GNNs for Few-Shot Early Action PredictionabstractWe aim to address a new task named few-shot early action prediction (FS-EAP) that learns classifiers for novel actions from only a few partially observed videos. We argue that the task is extremely challenging since the partially observed videos do not contain enough action information in a few-shot environment. To tackle this task, in this paper, we propose a scene-aware spatio-temporal graph neural network (SA-STGNN) by leveraging the fine-grained spatio-temporal interactions in the video scenes. Specifically, we first generate a spatio-temporal graph corresponding to the partially observed video to capture comprehensive spatio-temporal correlations. Then we utilize the spatio-temporal graph as the input of our SA-STGNN and predict the augmented video features corresponding to the complete video. The architecture uses several scene-aware learning blocks, which are a combination of edge fusion graph neural layers and temporal gated convolutional layers to jointly model spatial and temporal dependencies. Finally, we employ an early action predictor to exploit the learned video features for predicting actions in the few-shot setting. Extensive experimental results on two widely adopted video datasets demonstrate the effectiveness of our approach and its superior performance over the state-of-the-art approaches. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2022 | Enhancing BERT for Short Text Classification with Latent Information
Ailing Tang, Yufan Hu |
ICONIP (3) | 2 |
| 2021 | Learning Dual-Pooling Graph Neural Networks for Few-Shot Video ClassificationabstractWe address the problem of few-shot video classification that learns classifiers for novel concepts from only a few examples. Most current methods ignore to explicitly consider the relations in both intra-video and inter-video domains, thus cannot take full advantage of the structural information in few-shot learning. In this paper, we propose to exploit the comprehensive intra-video and inter-video relations via Graph Neural Networks (GNNs). To improve the discriminative ability for accurately selecting the representative video content and refining video relations, a Dual-Pooling GNN (DPGNN) is constructed, which stacks customized graph pooling layers in a hierarchical fashion. Specifically, to select the most representative frames in a video, we build intra-video graphs and utilize a node pooling module to extract robust video-level features. We construct an inter-video graph by taking the video-level features as nodes. By designing an edge pooling module, the proposed method can adaptively eliminate the negative relations in the inter-video graph. Extensive experimental results show that our method consistently outperforms the state-of-the-art on two benchmarks. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |