EDBT 2026 Demo / reviewers in the wild / expert
Hao Fang 0010
dblp:06/2484-10
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-8846-8294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empowering DINO Representations for Underwater Instance Segmentation via Aligner and PrompterabstractUnderwater Instance Segmentation (UIS), integrating pixel-level understanding and instance-level discrimination, is a pivotal technology in marine resource exploration and ecological protection. In recent years, large-scale pretrained visual foundation models, exemplified by DINO, have advanced rapidly and demonstrated remarkable performance on complex downstream tasks. In this paper, we demonstrate that DINO can serve as an effective feature learner for UIS, and we introduce DiveSeg, a novel framework built upon two insightful components: (1) The AquaStyle Aligner, designed to embed underwater color style features into the DINO fine-tuning process, facilitating better adaptation to the underwater domain. (2) The ObjectPrior Prompter, which incorporates binary segmentation-based prompts to deliver object-level priors, provides essential guidance for instance segmentation task that requires both object- and instance-level reasoning. We conduct thorough experiments on the popular UIIS and USIS10K datasets, and the results show that DiveSeg achieves the state-of-the-art performance. Chen Zhang 0013, Hao Fang 0010, Runmin Cong |
AAAI | 3 |
| 2026 | OVFormer+: Improved Open-Vocabulary Video Instance Segmentation via Text-Guided Unified Embedding Alignment
Hao Fang 0010, Xiankai Lu, Henghui Ding, Yunchao Wei, Yawei Li 0001, Runmin Cong |
Int. J. Comput. Vis. | 1 |
| 2026 | Breaking Barriers, Localizing Saliency: A Large-Scale Benchmark and Baseline for Condition-Constrained Salient Object DetectionabstractSalient Object Detection (SOD) aims to identify and segment the most prominent objects in an image. In real open environments, intelligent systems often encounter complex and challenging scenes, such as low-light, rain, snow, etc., which we call constrained conditions. These real situations pose more severe challenges to existing SOD models. However, there is no comprehensive and in-depth exploration of this field at both the data and model levels, and most of them focus on ideal situations or a single condition. To bridge this gap, we launch a new task, Condition-Constrained Salient Object Detection (CSOD), aimed at robustly and accurately locating salient objects in constrained environments. On the one hand, to compensate for the lack of datasets, we construct the first large-scale condition-constrained salient object detection dataset CSOD10 K, comprising 10,000 pixel-level annotated images and over 100 categories of salient objects. This dataset is oriented towards the real environment and includes 8 real-world constrained scenes under 3 main constraint types, making it extremely challenging. On the other hand, we abandon the paradigm of "restoration before detection" and instead introduce a unified end-to-end framework CSSAM that fully explores scene attributes, eliminating the need for additional ground-truth restored images and reducing computational overhead. Specifically, we design a Scene Prior-Guided Adapter (SPGA), which injects scene priors to enable the foundation model to better adapt to downstream constrained scenes. To automatically decode salient objects, we propose a Hybrid Prompt Decoding Strategy (HPDS), which can effectively integrate multiple types of prompts to achieve adaptation to the SOD task. Extensive experiments show that our model significantly outperforms state-of-the-art methods on both the CSOD10 K dataset and existing standard SOD benchmarks. Runmin Cong, Hao Fang 0010, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | G2HFNet: GeoGran-Aware Hierarchical Feature Fusion Network for Salient Object Detection in Optical Remote Sensing ImagesabstractRemote sensing images captured from aerial perspectives often exhibit significant scale variations and complex backgrounds, posing challenges for salient object detection (SOD). Existing methods typically extract multi-level features at a single scale using uniform attention mechanisms, leading to suboptimal representations and incomplete detection results. To address these issues, we propose a GeoGran-Aware Hierarchical Feature Fusion Network (G2HFNet) that fully exploits geometric and granular cues in optical remote sensing images. Specifically, G2HFNet adopts Swin Transformer as the backbone to extract multi-level features and integrates three key modules: the multi-scale detail enhancement (MDE) module to handle object scale variations and enrich fine details, the dual-branch geo-gran complementary (DGC) module to jointly capture fine-grained details and positional information in mid-level features, and the deep semantic perception (DSP) module to refine high-level positional cues via self-attention. Additionally, a local-global guidance fusion (LGF) module is introduced to replace traditional convolutions for effective multi-level feature integration. Extensive experiments demonstrate that G2HFNet achieves high-quality saliency maps and significantly improves detection performance in challenging remote sensing scenarios. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | RSONet: Region-Guided Selective Optimization Network for RGB-T Salient Object DetectionabstractThis paper focuses on the inconsistency in salient regions between RGB and thermal images. To address this issue, we propose the Region-guided Selective Optimization Network for RGB-T Salient Object Detection, which consists of the region guidance stage and saliency generation stage. In the region guidance stage, three parallel branches with same encoder-decoder structure equipped with the context interaction (CI) module and spatial-aware fusion (SF) module are designed to generate the guidance maps which are leveraged to calculate similarity scores. Then, in the saliency generation stage, the selective optimization (SO) module fuses RGB and thermal features based on the previously obtained similarity values to mitigate the impact of inconsistent distribution of salient targets between the two modalities. After that, to generate high-quality detection result, the dense detail enhancement (DDE) module which adopts the multiple dense connections and visual state space blocks is applied to low-level features for optimizing the detail information. In addition, the mutual interaction semantic (MIS) module is placed in the high-level features to dig the location cues by the mutual fusion strategy. We conduct extensive experiments on the RGB-T dataset, and the results demonstrate that the proposed RSONet achieves competitive performance against 27 state-of-the-art SOD methods. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Decoupled Motion Expression Video SegmentationabstractMotion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is more challenging. Previous works achieved it by simply injecting text information into the video instance segmentation (VIS) model. However, this requires retraining the entire model and optimization is difficult. In this work, we propose DMVS, a simple framework constructed on the existing query-based VIS model, emphasizing decoupling the task into video instance segmentation and motion expression understanding. Firstly, we use a frozen video instance segmenter to extract object-specific contexts and convert them into frame-level and video-level queries. Secondly, we interact two levels of queries with static and motion cues, respectively, to further encode visually enhanced motion expressions. Furthermore, we propose a novel query initialization strategy that uses video queries guided by classification priors to initialize motion queries, greatly reducing the difficulty of optimization. Without bells and whistles, DMVS achieves state-of-the-art performance on the MeViS dataset at a lower training cost. Extensive experiments verify the effectiveness and efficiency of our framework. Hao Fang 0010, Runmin Cong, Xiankai Lu, Xiaofei Zhou 0003, Sam Kwong, Wei Zhang 0021 |
CVPR | 1 |
| 2025 | Semantic and Sequential Alignment for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) seeks to segment the objects within a video referred by linguistic expressions. Existing RVOS solutions follow a "fuse then select" paradigm: establishing semantic correlation between visual and linguistic feature, and performing frame-level query interaction to select the instance mask per frame with instance segmentation module. This paradigm overlooks the challenge of semantic gap between the linguistic descriptor and the video object as well as the underlying clutters in the video. This paper proposes a novel Semantic and Sequential Alignment (SSA) paradigm to handle these challenges. We first insert a lightweight adapter after the vision language model (VLM) to perform the semantic alignment. Then, prior to selecting mask per frame, we exploit the trajectory-to-instance enhancement for each frame via sequential alignment. This paradigm leverages the visual-language alignment inherent in VLM during adaptation and tries to capture global information by ensembling trajectories. This helps understand videos and the corresponding descriptors by mitigating the discrepancy with intricate activity semantics, particularly when facing occlusion or similar interference. SSA demonstrates competitive performance while maintaining fewer learnable parameters. Feiyu Pan, Hao Fang 0010, Fangkai Li, Yawei Li 0001, Luca Benini, Xiankai Lu |
CVPR | 2 |
| 2025 | A Conditional Probability Framework for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of known objects and attributes by leveraging knowledge from previously seen compositions. Traditional approaches primarily focus on disentangling attributes and objects, treating them as independent entities during learning. However, this assumption overlooks the semantic constraints and contextual dependencies inside a composition. For example, certain attributes naturally pair with specific objects (e.g., "striped" applies to "zebra" or "shirts" but not "sky" or "water"), while the same attribute can manifest differently depending on context (e.g., "young" in "young tree" vs. "young dog"). Thus, capturing attribute-object interdependence remains a fundamental yet long-ignored challenge in CZSL. In this paper, we adopt a Conditional Probability Framework (CPF) to explicitly model attribute-object dependencies. We decompose the probability of a composition into two components: the likelihood of an object and the conditional likelihood of its attribute. To enhance object feature learning, we incorporate textual descriptors to highlight semantically relevant image regions. These enhanced object features then guide attribute learning through a cross-attention mechanism, ensuring better contextual alignment. By jointly optimizing object likelihood and conditional attribute likelihood, our method effectively captures compositional dependencies and generalizes well to unseen compositions. Extensive experiments on multiple CZSL benchmarks demonstrate the superiority of our approach. Code is available at here. Peng Wu 0014, Qiuxia Lai, Hao Fang 0010, Guosen Xie, Yilong Yin, Xiankai Lu, Wenguan Wang |
ICCV | 3 |
| 2025 | UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
Runmin Cong, Zongji Yu, Hao Fang 0010, Haoyan Sun, Sam Kwong |
ACM Multimedia | 3 |
| 2025 | Hierarchical spatiotemporal Feature Interaction Network for video saliency prediction
Yingjie Jin, Xiaofei Zhou 0003, Hao Fang 0010, Xiaobin Xu 0002 |
Image Vis. Comput. | 4 |
| 2025 | Shape Embedding and Knowledge Mining Network for Generalized Few-Shot Remote Sensing SegmentationabstractIn recent years, generalized few-shot segmentation (GFSS) has received widespread attention from scholars by virtue of its superiority in low-data regimes. Most of the existing research focuses on natural image processing, and few studies have been devoted to the practical but challenging topic of remote sensing image (RSI) understanding. In this paper, we propose a Shape Embedding and Knowledge Mining Network (SKNet) for generalized few-shot RSI segmentation. Specifically, the framework is divided into two key stages: (a) In the base class learning stage, shape representation embedding is introduced to enhance the network’s ability to perceive remote sensing objects. Simultaneously, we introduce the self-reconstruction constraint to prevent new unseen classes from merging, thereby improving the representation uniqueness of these classes. (b) In the novel class learning stage, a base class knowledge mining mechanism is designed to update the prototypes of the novel class by using the prototype representation of the base class, so as to enhance the discrimination ability of the network. We validated our methods on the adapted version of OpenEarthMap and iSAID datasets. In comparison with existing GFSS methods, the proposed approach demonstrates an advancement. Zifeng Qiu, Hongyu Liu 0003, Chengliang Di, Hao Fang 0010, Runmin Cong |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Learning Better Video Query With SAM for Video Instance SegmentationabstractRecently, Transformer-based offline video instance segmentation (VIS) solutions have made significant progress by decomposing the whole task into global segmentation map generation and instance discrimination. We argue that the quality of video queries that represent all instances in a video clip is crucial for offline VIS methods. Existing methods typically interact video queries with dense spatio-temporal features, resulting in significant computational complexity and redundant information. Thus, we propose a novel video instance segmentation framework, LBVQ, dedicated to learning better video queries. Specifically, we first obtain the frame queries for each frame independently without any complex inter-frame spatial-temporal association operations. Secondly, we propose an adaptive query initialization module (AQI), which adaptively integrates frame queries to initialize video queries instead of traditional random initialization strategies. This initialization method preserves rich instance clues and accelerates the optimization of the whole model. Finally, to enhance the quality of video queries, we propose a query propagation module (QPM) that captures relevant instance information in frame queries frame by frame, greatly improving the model’s understanding of long videos. By learning higher quality video queries, LBVQ achieves the state-of-the-art on VIS benchmarks with a ResNet-50 backbone: 52.2 AP, 44.8 AP on YouTube-VIS 2019 & 2021. Moreover, LBVQ achieves 39.7 AP on YouTube-VIS 2022 and 22.2 AP on OVIS, demonstrating superior potential for long videos. To further improve the quality of segmentation masks, a large-scale pretrained SAM is employed to refine the segmentation results. Code is available at https://github.com/fanghaook/LBVQ. Hao Fang 0010, Xiaofei Zhou 0003, Xinxin Zhang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
Hao Fang 0010, Peng Wu 0014, Yawei Li 0001, Xinxin Zhang 0004, Xiankai Lu |
ECCV (70) | 1 |
| 2024 | Structural Transformer with Region Strip Attention for Video Object Segmentation
Qingfeng Guan 0002, Hao Fang 0010, Chenchen Han, Zhicheng Wang 0017, Ruiheng Zhang 0001, Xiankai Lu |
Neurocomputing | 2 |
| 2024 | Gradformer: A Framework for Multi-Aspect Multi-Granularity Pronunciation AssessmentabstractAutomatic pronunciation assessment is an indispensable technology in computer-assisted pronunciation training systems. To further evaluate the quality of pronunciation, multi-task learning with simultaneous output of multi-granularity and multi-aspect has become a mainstream solution. Existing methods either predict scores at all granularity levels simultaneously through a parallel structure, or predict individual granularity scores layer by layer through a hierarchical structure. However, these methods do not fully understand and take advantage of the correlation between the three granularity levels of phoneme, word, and utterance. To address this issue, we propose a novel method, Granularity-decoupled Transformer (Gradformer), which is able to model the relationships between multiple granularity levels. Specifically, we first use a convolution-augmented transformer encoder to encode acoustic features, where the convolution module helps the model better capture local information. The model outputs both phoneme- and word-level granularity scores with high correlation by the encoder. Then, we use utterance queries to interact with the output of the encoder through the transformer decoder, ultimately obtaining the utterance scores. Through unique encoder and decoder architecture, we achieve decoupling at three granularity levels, and handling the relationship between each granularity. Experiments on the speachocean762 dataset show that our model has advantages over state-of-the-art methods in various metrics, especially in key metrics such as phoneme accuracy, word accuracy, and total score. Hao-Chen Pei, Hao Fang 0010, Xin Luo 0006, Xin-Shun Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Learning Feature Semantic Matching for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks. Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 2 |