EDBT 2026 Demo / reviewers in the wild / expert
Jiannan Ge
dblp:293/9559
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-2580-9055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityabstractMultimodal Large Language Models (MLLMs) have shown remarkable progress in temporal or spatial localization tasks, but struggle with joint spatio-temporal video grounding (STVG). We identify two key bottlenecks hindering this capability: (1) the sheer number of visual tokens makes long-range and fine-grained visual modeling challenging; (2) generating a long sequence of bounding boxes in text makes it hard to accurately align each box with its specific video frame. Distinct from prior efforts that rely on attaching complex modules, we argue for a more elegant paradigm that unlocks the inherent potential of MLLMs and leverages their strengths. To this end, we propose \textbf{\textit{SpaceVLLM}}, a MLLM equipped with spatio-temporal video grounding capabilities. Specifically, we propose Spatio-Temporal Aware Queries, interleaved with video frames, to guide the MLLM in capturing both static appearance and dynamic motion features. We further present a lightweight Query-Guided Space Head that maps queries to precise spatial coordinates, bypassing the need for direct textual coordinate generation and enabling the MLLM to focus on video understanding. To further facilitate research in this area, we propose an automated data synthesis pipeline to construct \textbf{V-STG} dataset, comprising 110K STVG instances. Extensive experiments show that \textit{SpaceVLLM} achieves the state-of-the-art performance on STVG benchmarks and maintains strong performance on various video understanding tasks, validating our approach's effectiveness. Jiankang Wang, Jiannan Ge, Hongtao Xie 0001, Yongdong Zhang 0001 |
AAAI | 5 |
| 2025 | CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Sun-Ao Liu, Xiaopeng Zhang 0008, Qi Tian 0001, Yongdong Zhang 0001 |
ICCV | 1 |
| 2025 | ReferSAM: Unleashing Segment Anything Model for Referring Image SegmentationabstractThe Segment Anything Model (SAM) has demonstrated remarkable capability as a general segmentation model given visual prompts such as points or boxes. While SAM is conceptually compatible with text prompts, it merely employs linguistic features from vision-language models as prompt embeddings and lacks fine-grained cross-modal interaction. This deficiency limits its application in referring image segmentation (RIS), where the targets are specified by free-form natural language expressions. In this paper, we introduce ReferSAM, a novel SAM-based framework that enhances cross-modal interaction and reformulates prompt encoding, thereby unleashing SAM’s segmentation capability for RIS. Specifically, ReferSAM incorporates the Vision-Language Interactor (VLI) to integrate linguistic features with visual features during the image encoding stage of SAM. This interactor introduces fine-grained alignment between linguistic features and multi-scale visual representations without altering the architecture of pre-trained models. Additionally, we present the Vision-Language Prompter (VLP) to generate dense and sparse prompt embeddings by aggregating the aligned linguistic and visual features. Consequently, the generated embeddings sufficiently prompt SAM’s mask decoder to provide precise segmentation results. Extensive experiments on five public benchmarks demonstrate that ReferSAM achieves state-of-the-art performance on both classic and generalized RIS tasks. The code and models are available at https://github.com/lsa1997/ReferSAM. Sun'ao Liu, Hongtao Xie 0001, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Denoised and Dynamic Alignment Enhancement for Zero-Shot LearningabstractZero-shot learning (ZSL) focuses on recognizing unseen categories by aligning visual features with semantic information. Recent advancements have shown that aligning each attribute with its corresponding visual region significantly improves zero-shot learning performance. However, the crude semantic proxies used in these methods fail to capture the varied appearances of each attribute, and are also easily confused by the presence of semantically redundant backgrounds, leading to suboptimal alignment. To combat these issues, we introduce a novel Alignment-Enhanced Network (AENet), designed to denoise the visual features and dynamically perceive semantic information, thus enhancing visual-semantic alignment. Our approach comprises two key innovations. (1) A visual denoising encoder, employing a class-agnostic mask to filter out semantically redundant visual information, thus producing refined visual features adaptable to unseen classes. (2) A dynamic semantic generator that crafts content-aware semantic proxies adaptively, steered by visual features, enabling AENet to discriminate fine-grained variations in visual contents. Additionally, we integrate a cross-fusion module to ensure comprehensive interaction between the denoised visual features and the generated dynamic semantic proxies, further facilitating visual-semantic alignment. Through extensive experiments across three datasets, the proposed method demonstrates that it narrows down the visual-semantic gap and sets a new benchmark in this setting. Jiannan Ge, Pandeng Li, Lingxi Xie, Yongdong Zhang 0001, Qi Tian 0001, Hongtao Xie 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalabstractVideo Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic richness inherent in videos far exceeds that of a given limited-length sentence. Therefore, in pursuit of better alignment, a natural idea is enhancing the video modality to filter out query-irrelevant semantics, and enhancing the text modality to capture more segment-relevant knowledge. In this paper, we introduce Modal-Enhanced Semantic Modeling (MESM), a novel framework for more balanced alignment through enhancing features at two levels. First, we enhance the video modality at the frame-word level through word reconstruction. This strategy emphasizes the portions associated with query words in frame-level features while suppressing irrelevant parts. Therefore, the enhanced video contains less redundant semantics and is more balanced with the textual modality. Second, we enhance the textual modality at the segment-sentence level by learning complementary knowledge from context sentences and ground-truth segments. With the knowledge added to the query, the textual modality thus maintains more meaningful semantics and is more balanced with the video modality. By implementing two levels of MESM, the semantic information from both modalities is more balanced to align, thereby bridging the modality gap. Experiments on three widely used benchmarks, including the out-of-distribution settings, show that the proposed framework achieves a new start-of-the-art performance with notable generalization ability (e.g., 4.42% and 7.69% average gains of [email protected] on Charades-STA and Charades-CG). The code will be available at https://github.com/lntzm/MESM. Hongtao Xie 0001, Pandeng Li, Jiannan Ge, Sun'ao Liu, Guoqing Jin |
AAAI | 5 |
| 2024 | AlignZeg: Mitigating Objective Misalignment for Zero-Shot Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Xiaopeng Zhang 0008, Yongdong Zhang 0001, Qi Tian 0001 |
ECCV (43) | 1 |
| 2024 | Towards Discriminative Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize both seen and unseen categories by establishing visual and semantic relations. Recently, generation-based methods that focus on synthesizing fictitious visual features from corresponding attributes have gained significant attention. However, these generated features often lack discriminative capabilities due to inadequate training of the generative model. To address this issue, we propose a novel Discriminative Enhanced Network (DENet) to harness the potential of the generative model by adapting the training features and imposing constraints on the generated features. Our approach incorporates three pivotal modules: (1) Before the generative network training, we implement a Pre-Tuning Module (PTM) to eliminate irrelevant background noise in the raw features extracted from a fixed CNN backbone. Therefore, PTM can provide tuned training features without redundant noise for generative model. (2) During the generative network training, we propose an Asymmetry Cross-authenticity Contrastive (AC2) loss to group visual features of the same category while repel features from different categories by optimizing a large number of sample pairs. Additionally, we incorporate intra-class and relation-specific inter-class boundaries within the AC2 loss to enrich sample diversity and preserve valid semantic information. (3) Also within the generative network training, a Dual-semantic Alignment Module (DAM) is designed to align visual features with both attributes and label embeddings, enabling the model to learn attribute-related information and discriminative extended semantics. Experiments on four standard benchmarks demonstrate that our approach learns more discriminative features and surpasses the existing methods. Jiannan Ge, Hongtao Xie 0001, Pandeng Li, Lingxi Xie, Shaobo Min, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Balanced Classification: A Unified Framework for Long-Tailed Object DetectionabstractConventional detectors suffer from performance degradation when dealing with long-tailed data due to a classification bias towards the majority head categories. In this article, we contend that the learning bias originates from two factors: 1) the unequal competition arising from the imbalanced distribution of foreground categories, and 2) the lack of sample diversity in tail categories. To tackle these issues, we introduce a unified framework calledBAlancedCLassification (BACL), which enables adaptive rectification of inequalities caused by disparities in category distribution and dynamic intensification of sample diversities in a synchronized manner. Specifically, a novel foreground classification balance loss (FCBL) is developed to ameliorate the domination of head categories and shift attention to difficult-to-differentiate categories by introducing pairwise class-aware margins and auto-adjusted weight terms, respectively. This loss prevents the over-suppression of tail categories in the context of unequal competition. Moreover, we propose a dynamic feature hallucination module (FHM), which enhances the representation of tail categories in the feature space by synthesizing hallucinated samples to introduce additional data variances. In this divide-and-conquer approach, BACL sets a new state-of-the-art on the challenging LVIS benchmark with a decoupled training pipeline, surpassing vanilla Faster R-CNN with ResNet-50-FPN by 5.8% AP and 16.1% AP for overall and tail categories. Extensive experiments demonstrate that BACL consistently achieves performance improvements across various datasets with different backbones and architectures. Tianhao Qi, Hongtao Xie 0001, Pandeng Li, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalabstractThe performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text description only corresponds to a spatio-temporal part of videos. Hence, the matching process should consider both fine-grained spatial content and various temporal semantic events. To this end, we propose a text-video learning framework with progressive spatio-temporal prototype matching. Specifically, the matching process is decomposed into two complementary phases: object-phrase prototype matching and event-sentence prototype matching. In the object-phrase prototype matching phase, the spatial prototype generation mechanism predicts key patches or words, which are aggregated into object or phrase prototypes. Importantly, optimizing the local alignment between object-phrase prototypes helps the model perceive spatial details. In the event-sentence prototype matching phase, we design a temporal prototype generation mechanism to associate intra-frame objects and interact inter-frame temporal relations. Such progressively generated event prototypes can reveal semantic diversity in videos for dynamic matching. Validated by comprehensive experiments, our method consistently outperforms the state-of-the-art methods on four video retrieval benchmark.1 Pandeng Li, Chen-Wei Xie, Hongtao Xie 0001, Jiannan Ge, Deli Zhao, Yongdong Zhang 0001 |
ICCV | 5 |
| 2023 | Frequency-based Zero-Shot Learning with Phase AugmentationabstractZero-Shot Learning (ZSL) aims to recognize images from seen and unseen classes by aligning visual and semantic knowledge (e.g., attribute descriptions). However, the fine-grained attributes in the RGB domain can be easily affected by background noise (e.g., the grey bird tail blending with the ground), making it difficult to effectively distinguish them. Analyzing the features in the frequency domain assists in better distinguishing the attributes since their patterns remain consistent across different images, unlike noise which may be more variable. Nevertheless, existing ZSL methods typically learn visual features directly from the RGB domain, which can impede the recognition of certain attributes. To overcome this limitation, we propose a novel ZSL method named Frequency-based Phase Augmentation (FPA) network, which learns an effective representation of the attributes in the frequency domain. Specifically, we introduce a Hybrid Phase Augmentation (HPA) module to transform visual features into the frequency domain and augment the phase component for better retention of semantic information of the attributes. The use of phase-augmented features enables FPA to capture more semantic knowledge that can be challenging to distinguish in the RGB domain, suppress noise, and highlight significant attributes. Our extensive experiments show that FPA achieves state-of-the-art performance across four standard datasets. Wanting Yin, Hongtao Xie 0001, Lei Zhang 0119, Jiannan Ge, Pandeng Li, Chuanbin Liu 0001, Yongdong Zhang 0001 |
ACM Multimedia | 4 |
| 2023 | Neighborhood-Adaptive Multi-Cluster Ranking for Deep Metric LearningabstractDeep metric learning methods generally concentrate on designing distance-based losses to learn sample embeddings, which tacitly presuppose the neighborhood structure around each sample (e.g., hypersphere for Euclidean distance). However, this supposition is overly optimistic: 1) visual data is often located on low-dimensional manifolds curved in high-dimensional space, and all regions of the manifold may hardly share the same local structures in the input space; 2) it is unlikely that the local structure in the output embedding space is as homogeneous as assumed due to the non-linearity of neural networks. Hence, simply characterizing sample embeddings while ignoring the respective neighborhood structures leads to limitations. To address this problem, this paper presents a Neighborhood-Adaptive Multi-cluster Ranking (NAMR) framework by leveraging the heterogeneity of local structures. Specifically, considering that indexing algorithms are usually required in large-scale retrieval, NAMR characterizes an image from two kinds of embeddings (i.e., sample embedding and structure embedding). The sample embedding can be trained using any distance-based loss, while the structure embedding representing the neighborhood structure can be jointly learned with the sample embedding in a self-supervised multi-cluster ranking manner. In this way, existing indexing algorithms can seamlessly support large-scale retrieval employing NAMR embeddings without any modifications. We evaluate the proposed model on five standard benchmarks, consistently and explicitly improving four baselines (especially the simplest triplet loss) and achieving state-of-the-art performance. Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval
Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Lei Zhang 0119, Shaobo Min, Yongdong Zhang 0001 |
ECCV (14) | 3 |
| 2022 | Dual Part Discovery Network for Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to recognize unseen classes by transferring knowledge from seen classes. Recent methods focus on learning a common semantic space to align visual and attribute information. However, they always over-relied on provided attributes and ignored the category discriminative information that contributes to accurate unseen class recognition, resulting in weak transferability. To this end, we propose a novel Dual Part Discovery Network (DPDN) that considers both attribute and category discriminative information by discovering attribute-guided parts and category-guided parts simultaneously to improve knowledge transfer. Specifically, for attribute-guided parts discovery, DPDN can localize the regions with specific attribute information and significantly bridge the gap between visual and semantic information guided by the given attributes. For category-guided parts discovery, the local parts are explored to discover other important regions that bring latent crucial details ignored by attributes, with the guidance of adaptive category prototypes. To better mine the transferable knowledge, we impose class correlations constraints to regularize the category prototypes. Finally, attribute- and category-guided parts complement each other and provide adequate discriminative subtle information for more accurate unseen class recognition. Extensive experimental results demonstrate that DPDN can discover discriminative parts and outperform state-of-the-art methods on three standard benchmarks. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Pandeng Li, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2022 | Deep Fourier Ranking Quantization for Semi-Supervised Image RetrievalabstractTo reduce the extreme label dependence of supervised product quantization methods, the semi-supervised paradigm usually employs massive unlabeled data to assist in regularizing deep networks, thereby improving model performance. However, the existing method focuses on the overall distribution consistency between unlabeled data and class prototypes, while ignoring subtle individual variances between unlabeled instances. Therefore, the local neighborhood structure is not fully explored, which will cause the model to easily overfit in the training set. In this paper, we introduce a new Fourier perspective to alleviate this issue by exploring the semantic relations between unlabeled instances in a self-supervised manner. Specifically, based on Fourier Transform, we first design a Phase Mixing (PM) strategy, which can manipulate the mixing area and values of the phase component between two images to control the proportion of semantic information. In this way, we can construct multi-level similarity neighbors naturally for unlabeled data. Then, a ranking quantization loss is formulated to perceive multi-level semantic variances in neighbor instances, which improves the robustness and generalization of the model. Extensive experiments in three different semi-supervised settings show that our method outperforms existing state-of-the-art methods by averaged 3.95% improvement on four datasets. Pandeng Li, Hongtao Xie 0001, Shaobo Min, Jiannan Ge, Xun Chen 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot LearningabstractGeneralized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively bridge the semantic-visual gap. However, these are all based on one-off visual localizers, lacking of interpretability and flexibility. In this paper, we propose a novel Semantic-guided Reinforced Region Embedding (SR2E) network that can localize important objects in the long-term interests to construct semantic-visual embedding space. SR2E consists of Reinforced Region Module (R2M) and Semantic Alignment Module (SAM). First, without the annotated bounding box as supervision, R2M encodes the semantic category guidance into the reward and punishment criteria to teach the localizer serialized region searching. Besides, R2M explores different action spaces during the serialized searching path to avoid local optimal localization, which thereby generates discriminative visual features with less redundancy. Second, SAM preserves the semantic relationship into visual features via semantic-visual alignment and designs a domain detector to alleviate the domain confusion. Experiments on four public benchmarks demonstrate that the proposed SR2E is an effective GZSL method with reinforced embedding space, which obtains averaged 6.1% improvements. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Yongdong Zhang 0001 |
AAAI | 1 |