Jie Hu 0018

dblp:90/5064-18 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-3187-1656ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 9 since 2021
YearPublicationVenuePosition
2025 U-SAM: Upgrade Segment Anything Model With Semantic-Aware and Memory-Efficient
abstract
Segment Anything Model (SAM) has achieved remarkable success in the field of class-agnostic image segmentation by utilizing points or boxes as prompts. However, we identify two significant limitations when compared to traditional image segmentation models: (1) Trained in a category-agnostic interactive segmentation manner, SAM lacks the ability to discern object granularity and semantics, rendering it ineffective for traditional instance, semantic, and panoptic segmentation tasks. (2) SAM’s inefficient use of instance-independent visual features and tokens necessitates maintaining unique features and tokens for each instance, leading to excessive GPU memory consumption and diminished segmentation efficiency. To address these issues, we propose the Universal Segment Anything Model (U-SAM), a semantic-aware and memory-efficient segmentation model designed to perform both promptable and traditional segmentation tasks within a compact and unified framework. Specifically, U-SAM enhances SAM by integrating the Multi-Scale Semantic-Aware Image Encoder (S2IE), thus providing multi-scale semantic features for achieving traditional image segmentation tasks. Additionally, U-SAM is equipped with a Twin Token Mask Decoder (T2MD) which reduces GPU memory overhead by substituting replicated visual features with replicated tokens. Extensive experiments across interactive, instance, semantic, and panoptic segmentation demonstrate U-SAM’s promising results. Notably, U-SAM is 9× smaller and 10× faster than SAM, showing strong performance in zero-shot segmentation. Moreover, U-SAM surpasses the SOTA object-prompter-based model, RSPrompter, by achieving a 6.2% increase in PQ, operating 14× faster, and cutting training memory usage by 61%.
Xiaofeng Jin, Jie Hu 0018, Jianghang Lin, Shengchuan Zhang, Liujuan Cao
ICASSP2
2025 DuPI: Dual-resolution Pseudo-label Integration for Semi-supervised Instance Segmentation
abstract
The role of high-quality pseudo-labels is pivotal in semi-supervised instance segmentation (SSIS). However, existing SSIS frameworks predominantly produce pseudo-labels at a single resolution, which can introduce noise that adversely affects the quality of learning at both the pixel level and in terms of class discrimination. This paper introduces the Dual-Resolution Pseudo-Label Integration for Semi-Supervised Instance Segmentation (DuPI), a novel framework designed to enhance learning by integrating pseudo-labels derived from dual-resolution inputs. The DuPI framework incorporates a Dual-Resolution Pseudo-Label Correction (DPC) module, which refines pseudo-labels through a process of cross-resolution rectification and fusion. Furthermore, the framework introduces an Area-Adaptive Learning (AAL) strategy aimed at enhancing the quality of pseudo-labels sourced from extra-resolution inputs. The AAL strategy addresses the training challenges associated with small objects at lower resolutions by re-weighting pseudo-labels corresponding to tiny mask areas using Intersection over Union (IoU) metrics from the assignments. Experiments on the COCO and BDD100K datasets demonstrate that DuPI achieves state-of-the-art SSIS performance under various semi-supervised settings.
Yue Ma 0030, Jie Hu 0018, Chen Chen 0001, Shengchuan Zhang, Xianming Lin, Liujuan Cao
ICASSP2
2025 WildSeg3D: Segment Any 3D Objects in the Wild from 2D Images
Yansong Guo, Jie Hu 0018, Yansong Qu, Liujuan Cao
ICCV2
2025 Universal Image Segmentation With Efficiency
abstract
In this paper, we present UISE, a unified image segmentation framework that achieves efficient performance across various segmentation tasks, eliminating the need for multiple specialized pipelines. UISE employs dynamic convolutions between universal segmentation kernels and image feature maps, enabling a single pipeline for different tasks such as panoptic, instance, semantic, and video instance segmentation. To address computational requirements, we introduce a feature pyramid aggregator for image feature extraction and a separable dynamic decoder for generating segmentation kernels. The aggregator re-parameterizes interpolation-first modules in a convolution-first manner, resulting in a significant acceleration of the pipeline without incurring additional costs. The decoder incorporates multi-head cross-attention through separable dynamic convolution, enhancing both efficiency and accuracy. Extensive experiments are conducted to validate UISE's performance across different segmentation tasks. To the best of our knowledge, UISE is the first universal segmentation framework that delivers competitive performance in terms of both speed and accuracy when compared to current state-of-the-art models.
Jie Hu 0018, Liujuan Cao, Xiaofeng Jin, Shengchuan Zhang, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 CamoTeacher: Dual-Rotation Consistency Learning for Semi-supervised Camouflaged Object Detection
Xunfa Lai, Jie Hu 0018, Shengchuan Zhang, Liujuan Cao, Guannan Jiang, Songan Zhang, Rongrong Ji
ECCV (45)3
2024 Prompting to Adapt Foundational Segmentation Models
abstract
Foundational segmentation models, predominantly trained on scenes typical of natural environments, struggle to generalize across varied image domains. Traditional "training-to-adapt'' methods rely heavily on extensive data retraining and model architectures modifications. This significantly limits the models' generalization capabilities and efficiency in deployment. In this study, we propose a novel adaptation paradigm, termed "prompting-to-adapt'', to tackle the above issue by introducing an innovative image prompter. This prompter generates domain-specific prompts through few-shot image-mask pairs, incorporating diverse image processing techniques to enhance adaptability. To tackle the inherent non-differentiability of image prompts, we further devise an information-estimation-based gradient descent strategy that leverages the information entropy of image processing combinations to optimize the prompter, ensuring effective adaptation. Through extensive experiments across nine datasets spanning seven image domains (i.e., depth, thermal, camouflage, endoscopic, ultrasound, grayscale, and natural) and four scenarios (i.e., common scenes, camouflage objects, medical images, and industrial data), we demonstrate that our approach significant improves the foundational models' adaptation capabilities. Moreover, the interpretability of the generated prompts provides insightful revelations into their image processing mechanisms. Source code is available at: \urlgithub.com/yuema1303/Prompting-to-Adapt-FSM.
Jie Hu 0018, Jie Li 0052, Yue Ma 0030, Liujuan Cao, Songan Zhang, Wei Zhang 0217, Guannan Jiang, Rongrong Ji
ACM Multimedia1
2024 Adaptive Selection based Referring Image Segmentation
abstract
Referring image segmentation (RIS) aims to segment a particular region based on a specific expression. Existing one-stage methods have explored various fusion strategies, yet they encounter two significant issues. Primarily, most methods rely on manually selected visual features from the visual encoder layers. Moreover, the direct fusion of word-level features into coarse aligned features disrupts the established vision-language alignment. In this paper, we introduce an innovative framework for RIS that seeks to overcome these challenges with adaptive alignment of vision and language features, termed the Adaptive Selection with Dual Alignment (ASDA). ASDA innovates in two aspects. Firstly, we design an Adaptive Feature Selection and Fusion (AFSF) module to dynamically select visual features focusing on different regions related to various descriptions. AFSF is equipped with scale-wise feature aggregator to provide hierarchically coarse features that preserve crucial low-level details. Secondly, a Word Guided Dual-Branch Aligner (WGDA) is leveraged to integrate coarse features with linguistic cues by word-guided attention, which effectively addresses the common issue of vision-language misalignment. Extensive experimental results demonstrate that our ASDA framework surpasses state-of-the-art methods on RefCOCO, RefCOCO+ and G-Ref benchmark.
Pengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu 0018, Hongwei Niu, Haixin Ding, Yan Zhang 0109, Guannan Jiang, Liujuan Cao, Rongrong Ji
ACM Multimedia4
2024 ISTR: Mask-Embedding-Based Instance Segmentation Transformer
abstract
Transformer-based instance-level recognition has attracted increasing research attention recently due to the superior performance. However, although attempts have been made to encode masks as embeddings into Transformer-based frameworks, how to combine mask embeddings and spatial information for a transformer-based approach is still not fully explored. In this paper, we revisit the design of mask-embedding-based pipelines and propose an Instance Segmentation TRansformer (ISTR) with Mask Meta-Embeddings (MME), leveraging the strengths of transformer models in encoding embedding information and incorporating spatial information from mask embeddings. ISTR incorporates a recurrent refining head that consists of a Dynamic Box Predictor (DBP), a Mask Information Generator (MIG), and a Mask Meta-Decoder (MMD). To improve the quality of mask embeddings, MME interprets the mask encoding-decoding processes as a mutual information maximization problem, which unifies the objective functions of different decoding schemes such as Principal Component Analysis (PCA) and Discrete Cosine Transform (DCT) with a meta-formulation. Under the meta-formulation, a learnable Spatial Mask Tuner (SMT) is further proposed, which fuses the spatial and embedding information produced from MIG and can significantly boost the segmentation performance. The resulting varieties, i.e., ISTR-PCA, ISTR-DCT, and ISTR-SMT, demonstrate the effectiveness and efficiency of incorporating mask embeddings with the query-based instance segmentation pipelines. On the COCO dataset, ISTR surpasses all predominant mask-embedding-based models by a large margin, and achieves competitive performance compared to concurrent state-of-the-art models. On the Cityscapes dataset, ISTR also outperforms several strong baselines. Our code has been made available at: https://github.com/hujiecpp/ISTR.
Jie Hu 0018, Yao Lu 0034, Shengchuan Zhang, Liujuan Cao
IEEE Trans. Image Process.1
2024 Bilateral Knowledge Interaction Network for Referring Image Segmentation
abstract
Referring image segmentation aims to segment objects that are described by natural language expressions. Although remarkable advancements have been made to align natural language expressions with visual representations for better performance, the interaction between image-level and text-level information is still not formulated properly. Most of the previous works focus on building correlations between vision and language, ignoring the variety of objects. The target objects with unique appearances may not be correctly located or completely segmented. In this article, we propose a novel Bilateral Knowledge Interaction Network, termed BKINet, which reformulates the image-text interaction in a bilateral manner to adapt concrete knowledge of the target object in the image. BKINet contains two key components: a knowledge learning module (KLM) and a knowledge applying module (KAM). In the KLM, the abstract knowledge from text features is replenished with concrete knowledge from visual features to adapt to the target objects in the input images, which generates the knowledge interaction kernels (KI kernels) containing abundant referring information. With the referring information of KI kernels, the KAM is designed to highlight the most relevant visual features for predicting the accurate segmentation mask. Extensive experiments on three widely-used datasets,i.e.RefCOCO, RefCOCO+, and G-ref, demonstrate the superiority of BKINet over the state-of-the-art.
Haixin Ding, Shengchuan Zhang, Qiong Wu 0012, Songlin Yu, Jie Hu 0018, Liujuan Cao, Rongrong Ji
IEEE Trans. Multim.5
2023 You Only Segment Once: Towards Real-Time Panoptic Segmentation
abstract
In this paper, we propose YOSO, a real-time panoptic segmentation framework. YOSO predicts masks via dynamic convolutions between panoptic kernels and image feature maps, in which you only need to segment once for both instance and semantic segmentation tasks. To reduce the computational overhead, we design a feature pyramid aggregator for the feature map extraction, and a separable dynamic decoder for the panoptic kernel generation. The aggregator re-parameterizes interpolation-first modules in a convolution-first way, which significantly speeds up the pipeline without any additional costs. The decoder performs multi-head cross-attention via separable dynamic convolution for better efficiency and accuracy. To the best of our knowledge, YOSO is the first real-time panoptic segmentation framework that delivers competitive performance compared to state-of-the-art models. Specifically, YOSO achieves 46.4 PQ, 45.6 FPS on COCO; 52.5 PQ, 22.6 FPS on Cityscapes; 38.0 PQ, 35.4 FPS on ADE20K; and 34.1 PQ, 7.1 FPS on Mapillary Vistas. Code is available at https://github.com/hujiecpp/YOSO.
Jie Hu 0018, Linyan Huang, Tianhe Ren, Shengchuan Zhang, Rongrong Ji, Liujuan Cao
CVPR1
2023 DistilPose: Tokenized Pose Regression with Heatmap Distillation
abstract
In the field of human pose estimation, regression-based methods have been dominated in terms of speed, while heatmap-based methods are far ahead in terms of performance. How to take advantage of both schemes remains a challenging problem. In this paper, we propose a novel human pose estimation framework termed DistilPose, which bridges the gaps between heatmap-based and regression-based methods. Specifically, DistilPose maximizes the transfer of knowledge from the teacher model (heatmap-based) to the student model (regression-based) through Token-distilling Encoder (TDE) and Simulated Heatmaps. TDE aligns the feature spaces of heatmap-based and regression-based models by introducing tokenization, while Simulated Heatmaps transfer explicit guidance (distribution and confidence) from teacher heatmaps into student models. Extensive experiments show that the proposed DistilPose can significantly improve the performance of the regression-based models while maintaining efficiency. Specifically, on the MSCOCO validation dataset, DistilPose-S obtains 71.6% mAP with 5.36M parameters, 2.38 GFLOPs, and 40.2 FPS, which saves 12.95×, 7.16× computational cost and is 4.9× faster than its teacher model with only 0.9 points performance drop. Furthermore, DistilPose-L obtains 74.4% mAP on MSCOCO validation dataset, achieving a new state-of-the-art among predominant regression-based models. Code will be available at https://github.com/yshMars/DistilPose.
Suhang Ye, Jie Hu 0018, Liujuan Cao, Shengchuan Zhang, Jun Wang 0006, Shouhong Ding, Rongrong Ji
CVPR3
2023 CANDY: Category-Kernelized Dynamic Convolution for Instance Segmentation
abstract
Instance segmentation has been dominated by the paradigm that predicts masks using local RoI features and simplicity frameworks based on global mask prediction. Despite the comparable performance between local-based and global-based approaches, the AP results of objects on different scales vary significantly. In this paper, we first point out that the key factor to bridging such a gap lies in the utilization of local RoI information for global mask prediction. Then, we observe a ’class-agnostic segmentation’ problem exists in the nearby region of interesting objects after implementing the above combination. To overcome this issue, we further propose a CAtegory-kerNelized DYnamic (CANDY) convolution. Benefiting from it, the discriminative ability of the resulting instance segmentation framework, i.e., CANDY-Mask, on foreground objects is significantly enhanced. Extensive experiments on the MS-COCO dataset are conducted to verify the performance of CANDY-Mask. Our proposed CANDY-Mask obtains 48.1% boxes AP and 40.7% masks AP on the MS-COCO test-dev set with ResNet50 backbone, achieving state-of-the-art among various models.
Yao Lu 0034, Jie Hu 0018, Liujuan Cao, Shengchuan Zhang
ICASSP4
2023 Pseudo-label Alignment for Semi-supervised Instance Segmentation
abstract
Pseudo-labeling is significant for semi-supervised instance segmentation, which generates instance masks and classes from unannotated images for subsequent training. However, in existing pipelines, pseudo-labels that contain valuable information may be directly filtered out due to mismatches in class and mask quality. To address this issue, we propose a novel framework, called pseudo-label aligning instance segmentation (PAIS), in this paper. In PAIS, we devise a dynamic aligning loss (DALoss) that adjusts the weights of semi-supervised loss terms with varying class and mask score pairs. Through extensive experiments conducted on the COCO and Cityscapes datasets, we demonstrate that PAIS is a promising framework for semi-supervised instance segmentation, particularly in cases where labeled data is severely limited. Notably, with just 1% labeled data, PAIS achieves 21.2 mAP (based on MaskRCNN) and 19.9 mAP (based on K-Net) on the COCO dataset, outperforming the current state-of-the-art model, i.e., NoisyBoundary with 7.7 mAP, by a margin of over 12 points. Code is available at: https://github.com/hujiecpp/PAIS.
Jie Hu 0018, Chen Chen 0001, Liujuan Cao, Shengchuan Zhang, Annan Shu, Guannan Jiang, Rongrong Ji
ICCV1
2023 CAM R-CNN: End-to-End Object Detection with Class Activation Maps
Shengchuan Zhang, Songlin Yu, Haixin Ding, Jie Hu 0018, Liujuan Cao
Neural Process. Lett.4
2021 Image-to-Image Translation via Hierarchical Style Disentanglement
abstract
Recently, image-to-image translation has made significant progress in achieving both multi-label (i.e., translation conditioned on different labels) and multi-style (i.e., generation with diverse styles) tasks. However, due to the unexplored independence and exclusiveness in the labels, existing endeavors are defeated by involving uncontrolled manipulations to the translation results. In this paper, we propose Hierarchical Style Disentanglement (HiSD) to address this issue. Specifically, we organize the labels into a hierarchical tree structure, in which independent tags, exclusive attributes, and disentangled styles are allocated from top to bottom. Correspondingly, a new translation process is designed to adapt the above structure, in which the styles are identified for controllable translations. Both qualitative and quantitative results on the CelebA-HQ dataset verify the ability of the proposed HiSD. The code has been released at https://github.com/imlixinyang/HiSD.
Shengchuan Zhang, Jie Hu 0018, Liujuan Cao, Xiaopeng Hong, Xudong Mao, Feiyue Huang, Yongjian Wu 0001, Rongrong Ji
CVPR3
2021 Architecture Disentanglement for Deep Neural Networks
abstract
Understanding the inner workings of deep neural networks (DNNs) is essential to provide trustworthy artificial intelligence techniques for practical applications. Existing studies typically involve linking semantic concepts to units or layers of DNNs, but fail to explain the inference process. In this paper, we introduce neural architecture disentanglement (NAD) to fill the gap. Specifically, NAD learns to disentangle a pre-trained DNN into sub-architectures according to independent tasks, forming information flows that describe the inference processes. We investigate whether, where, and how the disentanglement occurs through experiments conducted with handcrafted and automatically-searched network architectures, on both object-based and scene-based datasets. Based on the experimental results, we present three new findings that provide fresh insights into the inner logic of DNNs. First, DNNs can be divided into sub-architectures for independent tasks. Second, deeper layers do not always correspond to higher semantics. Third, the connection type in a DNN affects how the information flows across layers, leading to different disentanglement behaviors. With NAD, we further explain why DNNs sometimes give wrong predictions. Experimental results show that misclassified images have a high probability of being assigned to task sub-architectures similar to the correct ones. Our code is available at https://github.com/hujiecpp/NAD.
Jie Hu 0018, Liujuan Cao, Qixiang Ye, Shengchuan Zhang, Ke Li 0015, Feiyue Huang, Ling Shao 0001, Rongrong Ji
ICCV1
2019 Towards Visual Feature Translation
abstract
Most existing visual search systems are deployed based upon fixed kinds of visual features, which prohibits the feature reusing across different systems or when upgrading systems with a new type of feature. Such a setting is obviously inflexible and time/memory consuming, which is indeed mendable if visual features can be ``translated" across systems. In this paper, we make the first attempt towards visual feature translation to break through the barrier of using features across different visual search systems. To this end, we propose a Hybrid Auto-Encoder (HAE) to translate visual features, which learns a mapping by minimizing the translation and reconstruction errors. Based upon HAE, an Undirected Affinity Measurement (UAM) is further designed to quantify the affinity among different types of visual features. Extensive experiments have been conducted on several public datasets with sixteen different types of widely-used features in visual search systems. Quantitative results show the encouraging possibilities of feature translation. For the first time, the affinity among widely-used features like SIFT and DELF is reported.
Jie Hu 0018, Rongrong Ji, Hong Liu 0009, Shengchuan Zhang, Cheng Deng 0002, Qi Tian 0001
CVPR1
2019 Information Competing Process for Learning Diversified Representations
abstract
Learning representations with diversified information remains as an open problem. Towards learning diversified representations, a new approach, termed Information Competing Process (ICP), is proposed in this paper. Aiming to enrich the information carried by feature representations, ICP separates a representation into two parts with different mutual information constraints. The separated parts are forced to accomplish the downstream task independently in a competitive environment which prevents the two parts from learning what each other learned for the downstream task. Such competing parts are then combined synergistically to complete the task. By fusing representation parts learned competitively under different conditions, ICP facilitates obtaining diversified representations which contain rich information. Experiments on image classification and image reconstruction tasks demonstrate the great potential of ICP to learn discriminative and disentangled representations in both supervised and self-supervised learning settings.
Jie Hu 0018, Rongrong Ji, Shengchuan Zhang, Xiaoshuai Sun, Qixiang Ye, Chia-Wen Lin, Qi Tian 0001
NeurIPS1
2019 Face Sketch Synthesis by Multidomain Adversarial Learning
abstract
Given a training set of face photo-sketch pairs, face sketch synthesis targets at learning a mapping from the photo domain to the sketch domain. Despite the exciting progresses made in the literature, it retains as an open problem to synthesize high-quality sketches against blurs and deformations. Recent advances in generative adversarial training provide a new insight into face sketch synthesis, from which perspective the existing synthesis pipelines can be fundamentally revisited. In this paper, we present a novel face sketch synthesis method by multidomain adversarial learning (termed MDAL), which overcomes the defects of blurs and deformations toward high-quality synthesis. The principle of our scheme relies on the concept of "interpretation through synthesis." In particular, we first interpret face photographs in the photodomain and face sketches in the sketch domain by reconstructing themselves respectively via adversarial learning. We define the intermediate products in the reconstruction process as latent variables, which form a latent domain. Second, via adversarial learning, we make the distributions of latent variables being indistinguishable between the reconstruction process of the face photograph and that of the face sketch. Finally, given an input face photograph, the latent variable obtained by reconstructing this face photograph is applied for synthesizing the corresponding sketch. Quantitative comparisons to the state-of-the-art methods demonstrate the superiority of the proposed MDAL method.
Shengchuan Zhang, Rongrong Ji, Jie Hu 0018, Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Robust Face Sketch Synthesis via Generative Adversarial Fusion of Priors and Parametric Sigmoid
abstract
Despite the extensive progress in face sketch synthesis, existing methods are mostly workable under constrained conditions, such as fixed illumination, pose, background and ethnic origin that are hardly to control in real-world scenarios. The key issue lies in the difficulty to use data under fixed conditions to train a model against imaging variations. In this paper, we propose a novel generative adversarial network termed pGAN, which can generate face sketches efficiently using training data under fixed conditions and handle the aforementioned uncontrolled conditions. In pGAN, we embed key photo priors into the process of synthesis and design a parametric sigmoid activation function for compensating illumination variations. Compared to the existing methods, we quantitatively demonstrate that the proposed method can work well on face photos in the wild.
Shengchuan Zhang, Rongrong Ji, Jie Hu 0018, Yue Gao 0002, Chia-Wen Lin
IJCAI3