Jin Xie 0005

dblp:80/1949-5 · DBLP profile ↗
← Back
44ranked-venue papers
6as first author
39since 2021 · last 2026
0000-0001-6978-8834ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sparse gain adaptation with dual-domain fusion network for multimodal object detection
Xichuan Zhou, Boya Wei, Cong Mao, Lihui Chen 0002, Haijun Liu 0001, Jin Xie 0005, Jing Nie 0001
Neurocomputing8
2026 Multispectral remote sensing object detection via selective cross-modal interaction and aggregation
Minghao Cui, Jing Nie 0001, Hanqing Sun 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang, Xuelong Li 0001
Neural Networks4
2026 iSeg: An Iterative Refinement-Based Framework for Training-Free Segmentation
abstract
Stable Diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers have explored employing Stable Diffusion for training-free segmentation. Most existing approaches refine cross-attention map by self-attention map once, demonstrating that self-attention map contains useful semantic information to improve segmentation. To fully utilize self-attention map, we present a deep experimental analysis on iteratively refining cross-attention map with self-attention map, and propose an effective iterative refinement framework for training-free segmentation, named iSeg. Our iSeg introduces an entropy-reduced self-attention module that utilizes a gradient descent scheme to reduce the entropy of self-attention map, thereby suppressing the weak responses corresponding to irrelevant global information. Leveraging the entropy-reduced self-attention module, our iSeg stably improves cross-attention map with iterative refinement. Further, we design a category-enhanced cross-attention module to generate accurate cross-attention map, providing a better initial input for iterative refinement. Extensive experiments across different datasets and diverse segmentation tasks (weakly-supervised semantic segmentation, open-vocabulary semantic segmentation, unsupervised segmentation, and mask generation on synthetic dataset) reveal the merits of proposed contributions, leading to promising performance. For unsupervised semantic segmentation on Cityscapes, our iSeg achieves an absolute gain of 3.8% in terms of mIoU compared to the best existing training-free approach in literature. Moreover, our proposed iSeg can support segmentation with different kinds of images and interactions, and also be used as a post-processing, or in different frameworks, to improve training-free segmentation.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 SED++: A Simple Encoder-Decoder for Improved Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to partition an image into distinct semantic regions based on an open set of categories. Existing approaches primarily rely on image-level pre-trained vision-language models to perform this pixel-level task. In this paper, we propose SED, a simple yet effective encoder-decoder architecture for open-vocabulary semantic segmentation leveraging pre-trained vision-language models. SED consists of a hierarchical image encoder, a text encoder, and a gradual fusion decoder. The hierarchical image encoder and text encoder collaboratively generate a cost volume, which is progressively decoded by the gradual fusion decoder to produce segmentation results. In contrast to a plain encoder, the hierarchical encoder better captures image detail information while maintaining linear computational complexity with respect to input size. The gradual fusion decoder adopts a top-down structure to progressively integrate high-resolution features with the cost volume. Furthermore, a category early rejection strategy is introduced in gradual fusion decoder to filter out non-existent categories at different layers, significantly improving inference efficiency. Based on SED, we further introduce two modules, including non-label text embedding and additional category early rejection in the encoder. Moreover, we extend our method with minimal decoder modification for open-vocabulary video semantic segmentation. Extensive experiments on multiple datasets validate the effectiveness and efficiency of our proposed method. With ConvNeXt-B, our method achieves an mIoU of 34.9% on the ADE20 K with 150 classes (i.e., A-150) at an inference speed of 69 ms per image on a single A6000 GPU, and has an mIoU score of 40.2% on video segmentation dataset VSPW.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 SNNSIR: A fully Spiking Neural Network for Stereo Image Restoration
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang
Pattern Recognit.2
2026 HG-LMM: Unleashing High-Quality Pixel Grounding Capabilities in Frozen Large Multimodal Models
abstract
Large Multimodal Models (LMMs) have demonstrated remarkable capabilities in multimodal understanding and conversation. Recently, some researchers have explored fine-tuning LMM for pixel grounding, leading to catastrophic loss of their inherent conversational capabilities. To preserve the conversational capability, some researchers explore to freeze LMM, but employ heavy segmenter SAM for high-quality grounding. In this paper, we propose a novel approach, named HG-LMM, to fully exploit the inherent features of frozen LMM for high-quality pixel grounding. Our HG-LMM introduces two main modules: an LLM-guided instance-aware feature generation (LIFG) and a layer-wise detail and semantic injection (LDSI). The LIFG module employs the output text embeddings of LMM belonging to grounded instances to generate multi-level instance-aware feature maps from the image encoder. Afterwards, we employ the LDSI module to inject more detail and semantic information into these instance-aware feature maps. With these instance-aware feature maps, we employ a simple top-down fusion to predict the segmentation masks of different instances. We perform experiments on various tasks, including referring expression segmentation, panoramic narrative grounding, reasoning segmentation, grounded conversation generation, and visual chain-of-thought reasoning. When using DeepSeekVL-1.3B, our HG-LMM is 12.6% better than F-LMM without SAM in terms of segmentation accuracy on the all set of PNG dataset. Compared to F-LMM with SAM, our HG-LMM achieves comparable segmentation accuracy while being 2.7 times faster. We release our source code and models at https://github.com/WenjieLi2008/HG-LMM.
Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang
IEEE Trans. Image Process.3
2026 Frequency-Decomposed Interaction Network for Stereo Image Restoration
abstract
Stereo image restoration in adverse environments, such as low-light conditions, rain, and low resolution, requires effective exploitation of cross-view complementary information to recover degraded visual content. In monocular image restoration, frequency decomposition has proven effective, where high-frequency components aid in recovering fine textures and reducing blur, while low-frequency components facilitate noise suppression and illumination correction. However, existing stereo restoration methods have yet to explore cross-view interactions by frequency decomposition, which is a promising direction for enhancing restoration quality. To address this, we propose a frequency-aware framework comprising a Frequency Decomposition Module (FDM), Detail Interaction Module (DIM), Structural Interaction Module (SIM), and Adaptive Fusion Module (AFM). FDM employs learnable filters to decompose the image into high- and low-frequency components. DIM enhances the high-frequency branch by capturing local detail cues through deformable convolution. SIM processes the low-frequency branch by modeling global structural correlations via a cross-view row-wise attention mechanism. Finally, AFM adaptively fuses the complementary frequency-specific information to generate high-quality restored images. Extensive experiments demonstrate the efficacy and generalizability of our framework across three diverse stereo restoration tasks, where it achieves state-of-the-art performance in low-light enhancement, rain removal, alongside highly competitive results in super-resolution. Our code is available at https://github.com/C2022J/FDIN.
Xianmin Tian, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.2
2025 SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
abstract
Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually aggregate multimodal features at a single stage. However, leveraging multi-stage cross-modal features is crucial for detecting objects of various scales. Therefore, these methods often struggle to integrate features across different scales and modalities effectively, thereby restricting the accuracy of detection. Additionally, the time-consuming Query-Key-Value-based (QKV-based) cross-attention operations often utilized in existing methods aid in reasoning the location and existence of objects by capturing non-local contexts. However, this approach tends to increase computational complexity. To address these challenges, we present SSLFusion, a novel Scale & Space Aligned Latent Fusion Model, consisting of a scale-aligned fusion strategy (SAF), a 3D-to-2D space alignment module (SAM), and a latent cross-modal fusion module (LFM). SAF mitigates scale misalignment between modalities by aggregating features from both images and point clouds across multiple levels. SAM is designed to reduce the inter-modal gap between features from images and point clouds by incorporating 3D coordinate information into 2D image features. Additionally, LFM captures cross-modal non-local contexts in the latent space without utilizing the QKV-based attention operations, thus mitigating computational complexity. Experiments on the KITTI and DENSE datasets demonstrate that our SSLFusion outperforms state-of-the-art methods. Our approach obtains an absolute gain of 2.15% in 3D AP, compared with the state-of-art method GraphAlign on the moderate level of the KITTI test set.
Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao
AAAI2
2025 CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
abstract
Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level open-vocabulary semantic segmentation without additional training. The key is to improve spatial representation of image-level CLIP, such as replacing self-attention map at last layer with self-self attention map or vision foundation model based attention map. In this paper, we present a novel hierarchical framework, named CLIPer, that hierarchically improves spatial representation of CLIP. The proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module. We observe that, the embeddings and attention maps at early layers can preserve spatial structural information. Inspired by this, we design the early-layer fusion module to generate segmentation map with better spatial coherence. Afterwards, we employ a fine-grained compensation module to compensate the local details using the self-attention maps of diffusion model. We conduct the experiments on seven segmentation datasets. Our proposed CLIPer achieves the state-of-the-art performance on these datasets. For instance, using ViT-L, CLIPer has the mIoU of 69.8% and 43.3% on VOC and COCO Object, outperforming ProxyCLIP by 9.2% and 4.1% respectively.
Jiale Cao, Jin Xie 0005, Xiaoheng Jiang, Yanwei Pang
ICCV3
2025 DANet: spatial gene expression prediction from H&E histology images through dynamic alignment
abstract
Predicting spatial gene expression from Hematoxylin and Eosin histology images offers a promising approach to significantly reduce the time and cost associated with gene expression sequencing, thereby facilitating a deeper understanding of tissue architecture and disease mechanisms. Achieving accurate gene expression prediction requires the extraction of highly refined features from pathological images; however, existing methods often struggle to effectively capture fine-grained local details and model gene-gene correlations. Moreover, in bimodal contrastive learning, dynamically and efficiently aligning heterogeneous modalities remains a critical challenge. To address these issues, we propose a novel method for predicting gene expression. First, we introduce a dense connective structure that enables efficient feature reuse, thereby enhancing the capturing and mining of local refinement features. Second, we leverage the state space models to uncover underlying patterns and capture dependencies within 1D gene expression data, enabling more accurate modeling of gene-gene correlations. Furthermore, we design the Residual Kolmogorov-Arnold Network (RKAN) that uses a learnable activation function to dynamically adjust bimodal mappings based on input characteristics. Through continuous parameter updates during contrastive training, RKAN progressively refines the alignment between modalities. Extensive experiments conducted on two publicly available datasets, GSE240429 and HER2+, demonstrate the effectiveness of our approach and its significant improvements over existing methods. Source codes are available at https://github.com/202324131016T/DANet.
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yuansong Zeng
Briefings Bioinform.2
2025 Hybrid cross-modality fusion network for medical image segmentation with contrastive learning
Xichuan Zhou, Jing Nie 0001, Haijun Liu 0001, Fu Liang, Lihui Chen 0002, Jin Xie 0005
Eng. Appl. Artif. Intell.8
2025 E-TransConvNet: An enhanced transformer and convolutional network for medical image segmentation from ultrasound and CT images
Chukwuemeka Clinton Atabansi, Jing Nie 0001, Jiachen Huang, Haijun Liu 0001, Jin Xie 0005, Xichuan Zhou
Expert Syst. Appl.6
2025 CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
abstract
Open-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Language-Image Pre-training (CLIP) has shown robust zero-shot classification ability in image-level open-vocabulary tasks. In this paper, we propose a simple encoder-decoder network, called CLIP-VIS, to adapt CLIP for open-vocabulary video instance segmentation. Our CLIP-VIS adopts frozen CLIP and introduces three modules, including class-agnostic mask generation, temporal topK-enhanced matching, and weighted open-vocabulary classification. Given a set of initial queries, class-agnostic mask generation introduces a pixel decoder and a transformer decoder on CLIP pre-trained image encoder to predict query masks and corresponding object scores and mask IoU scores. Then, temporal topK-enhanced matching performs query matching across frames using the K mostly matched frames. Finally, weighted open-vocabulary classification first employs mask pooling to generate query visual features from CLIP pre-trained image encoder, and second performs weighted classification using object scores and mask IoU scores. Our CLIP-VIS does not require the annotations of instance categories and identities. The experiments are performed on various video instance segmentation datasets, which demonstrate the effectiveness of our proposed method, especially for novel categories. When using ConvNeXt-B as backbone, our CLIP-VIS achieves the AP and APn scores of 32.2% and 40.2% on the validation set of LV-VIS dataset, which outperforms OV2Seg by 11.1% and 23.9% respectively. We will release the source code and models athttps://github.com/zwq456/CLIP-VIS.git.
Jiale Cao, Jin Xie 0005, Shuangming Yang, Yanwei Pang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Video Instance Segmentation Without Using Mask and Identity Supervision
abstract
Video instance segmentation (VIS) is a challenging vision problem in which the task is to simultaneously detect, segment, and track all the object instances in a video. Most existing VIS approaches rely on pixel-level mask supervision within a frame as well as instance-level identity annotation across frames. However, obtaining these ‘mask and identity’ annotations is time-consuming and expensive. We propose the first mask-identity-free VIS framework that neither utilizes mask annotations nor requires identity supervision. Accordingly, we introduce a query contrast and exchange network (QCEN) comprising instance query contrast and query-exchanged mask learning. The instance query contrast first performs cross-frame instance matching and then conducts query feature contrastive learning. The query-exchanged mask learning exploits both intra-video and inter-video query exchange properties: exchanging queries of an identical instance from different frames within a video results in consistent instance masks, whereas exchanging queries across videos results in all-zero background masks. Extensive experiments on three benchmarks (YouTube-VIS 2019, YouTube-VIS 2021, and OVIS) reveal the merits of the proposed approach, which significantly reduces the performance gap between the identify-free baseline and our mask-identify-free VIS method. On the YouTube-VIS 2019 validation set, our mask-identity-free approach achieves 91.4% of the stronger-supervision-based baseline performance when utilizing the same ImageNet pre-trained model.
Jiale Cao, Hanqing Sun 0001, Rao Muhammad Anwer, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Multim.5
2025 Implicit and Explicit Language Guidance for Diffusion-Based Visual Perception
abstract
Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich textures and reasonable structures under different text prompts. However, adapting pre-trained diffusion models for visual perception is an open problem. In this paper, we propose an implicit and explicit language guidance framework for diffusion-based visual perception, named IEDP. Our IEDP comprises an implicit language guidance branch and an explicit language guidance branch. The implicit branch employs a frozen CLIP image encoder to directly generate implicit text embeddings that are fed to the diffusion model without explicit text prompts. The explicit branch uses the ground-truth labels of corresponding images as text prompts to condition feature extraction in diffusion model. During training, we jointly train the diffusion model by sharing the model weights of these two branches. As a result, the implicit and explicit branches can jointly guide feature learning. During inference, we employ only implicit branch for final prediction, which does not require any ground-truth labels. Experiments are performed on two typical perception tasks, including semantic segmentation and depth estimation. Our IEDP achieves promising performance on both tasks. For semantic segmentation, our IEDP has the mIoU$^\text{ss}$score of 55.9% on ADE20K validation set, which outperforms the baseline method VPD by 2.2%. For depth estimation, our IEDP outperforms the baseline method VPD with a relative gain of 11.0%.
Hefeng Wang, Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang
IEEE Trans. Multim.3
2024 MSAnomaly: Time Series Anomaly Detection with Multi-scale Augmentation and Fusion
Shikang Hou, Lijiao Zheng, Jin Xie 0005, Meng Yan 0001
ADMA (4)7
2024 Dual Interaction and Kernel-Diverse Network for Accurate Drug-Target Binding Affinity Prediction
abstract
Drug-target binding affinities measure the binding strength between drugs and targets, serving as the foundation for computer-aided drug design. The existing methods based on neural network rarely incorporate knowledge from biochemistry, which limits their performance. In this paper, we propose a dual interaction and kernel-diverse network for drug-target binding affinity prediction, termed DTANet, drawing inspiration from the biochemical properties of drugs and targets. Specifically, since functional groups and peptide chains are pivotal in determining the properties of drugs and their targets, and their lengths vary significantly. Drawing inspiration from this, we propose a Kernel-diverse Two-stream Network (KTN) and a Cross-Scale Interaction Module (CSIM) to extract the features of drugs and targets. The binding site of drug and target is crucial for precise affinity prediction. Inspired by this, we introduce the Drug-Target Interaction Module (DTIM) to explore the relationships between drugs and targets to focus on binding sites. Experiments have been conducted on two popular benchmarks: KIBA and Davis. Our DTANet obtains superior performance compared to existing methods on both datasets. Source codes are available at https://github.com/202324131016T/DTANet.
Jin Xie 0005, Jing Nie 0001, Yuansong Zeng
BIBM2
2024 Mamba-DTA: Drug-Target Binding Affinity Prediction with State Space Model
abstract
Biochemical methods for measuring drug-target binding are costly and slow, while deep learning offers a crucial solution. Deep learning methods for predicting drug-target binding affinity have achieved remarkable success, but existing approaches face challenges: adjacent elements in a three-dimensional structure may be far apart in a one-dimensional representation, and long sequences of drug and target data often contain lots of noise. In this paper, we introduce MambaDTA, a novel architecture for drug-target affinity prediction based on the State Space Model (SSM). Mamba-DTA utilizes SSM to model the drug molecules and target molecules and extract more discriminative spatial structural features efficiently and stably. Additionally, we design Interaction-based Selective Filtering (ISF) module to model drug-target interactions and filter out redundant information. The experimental results on two publicly available datasets, namely Davis and KIBA, demonstrate the effectiveness and superiority of our Mamba-DTA. Specifically, Mamba-DTA achieves a relative gain of 13.3% in terms of MAE on the Davis dataset. Source codes are available at https://github.com/202324131016T/Mamba-DTA.
Jin Xie 0005, Jing Nie 0001, Xiaohong Zhang 0002, Yuansong Zeng
BIBM2
2024 SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models, in which the key is to adapt the image-level model for pixel-level segmentation task. In this paper, we propose a simple encoder-decoder, named SED, for open-vocabulary semantic segmentation, which comprises a hierarchical encoder-based cost map generation and a gradual fusion decoder with category early rejection. The hierarchical encoder-based cost map generation employs hierarchical backbone, instead of plain transformer, to predict pixel-level image-text cost map. Compared to plain transformer, hierarchical backbone better captures local spatial information and has linear computational complexity with respect to input size. Our gradual fusion decoder employs a top-down structure to combine cost map and the feature maps of different backbone levels for segmentation. To accelerate inference speed, we introduce a category early rejection scheme in the decoder that rejects many no-existing categories at the early layer of decoder, resulting in at most 4.7 times acceleration without accuracy degradation. Experiments are performed on multiple open-vocabulary semantic segmentation datasets, which demonstrates the efficacy of our SED method. When using ConvNeXt-B, our SED method achieves mIoU score of 31.6% on ADE20K with 150 categories at 82 millisecond (ms) per image on a single A6000. Our source code is available at https://github.com/xb534/SED.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
CVPR3
2024 Localization-aware logit mimicking for object detection in adverse weather conditions
Peiyun Luo, Jing Nie 0001, Jin Xie 0005, Jiale Cao, Xiaohong Zhang 0002
Image Vis. Comput.3
2024 C2BG-Net: Cross-modality and cross-scale balance network with global semantics for multi-modal 3D object detection
Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao
Neural Networks2
2024 Multi-query and multi-level enhanced network for semantic segmentation
Jiale Cao, Rao Muhammad Anwer, Jin Xie 0005, Jing Nie 0001, Ai-Ping Yang, Yanwei Pang
Pattern Recognit.4
2024 Mirrored EAST: An Efficient Detector for Automatic Vehicle Identification Number Detection in the Wild
abstract
Vehicle identification number (VIN) is a unique serial number used to identify individual vehicles across various applications. The first crucial step in automatically collecting VINs is to accurately localize the VIN area. In this article, we present a novel VIN detection approach called Mirrored EAST (MEAST) based on an efficient and accurate scene text (EAST) Detection framework. MEAST learns to exploit the spatial consistency between an image and its mirrored version to improve localization performance, and employs a lighter but more discriminative backbone network to improve its applicability in mobile scenarios. To evaluate the VIN detection performance, we constructed a large-scale VIN image dataset named CQU-VD20 K, consisting of 20 000 VIN images in real scenarios. Based on this dataset, we have conducted a comprehensive empirical study of VIN detection. The results demonstrate the superiority of MEAST over other methods in VIN detection. Additionally, we also conducted extended experiments on a license plate dataset named CCPD-Rotate, which confirms the effectiveness of our approach in other industrial inspection tasks.
Guowei Yin, Sheng Huang 0001, Jin Xie 0005, Dan Yang 0001
IEEE Trans. Ind. Informatics4
2024 Transformer-Based Stereo-Aware 3D Object Detection From Binocular Images
abstract
Transformers have shown promising progress in various visual object detection tasks, including monocular 2D/3D detection and surround-view 3D detection. More importantly, the attention mechanism in the Transformer model and the 3D information extraction in binocular stereo are both similarity-based. However, directly applying existing Transformer-based detectors to binocular stereo 3D object detection leads to slow convergence and significant precision drops. We argue that a key cause of that defect is that existing Transformers ignore the binocular-stereo-specific image correspondence information. In this paper, we explore the model design of Transformers in binocular 3D object detection, focusing particularly on extracting and encoding task-specific image correspondence information. To achieve this goal, we present TS3D, a Transformer-based Stereo-aware 3D object detector. In the TS3D, a Disparity-Aware Positional Encoding (DAPE) module is proposed to embed the image correspondence information into stereo features. The correspondence is encoded as normalized sub-pixel-level disparity and is used in conjunction with sinusoidal 2D positional encoding to provide the 3D location information of the scene. To enrich multi-scale stereo features, we propose a Stereo Preserving Feature Pyramid Network (SPFPN). The SPFPN is designed to preserve the correspondence information while fusing intra-scale and aggregating cross-scale stereo features. Our proposed TS3D achieves a 41.29% Moderate Car detection average precision on the KITTI test set and takes 88 ms to detect objects from each binocular image pair. It is competitive with advanced counterparts in terms of both precision and inference speed.
Hanqing Sun 0001, Yanwei Pang, Jiale Cao, Jin Xie 0005, Xuelong Li 0001
IEEE Trans. Intell. Transp. Syst.4
2024 Binocular Image Dehazing via a Plain Network Without Disparity Estimation
abstract
Heavy haze leads to severely degraded visual quality for images, and thus the performance of high level image-based tasks such as object detection and semantic segmentation is deteriorated. It is necessary and important to design an effective dehazing method for the computer vision system. It is well known that image haze is a function of depth and binocular images can predict the depth. Existing binocular dehazing methods conduct disparity estimation and dehazing jointly to enhance each other. However, a small error in disparity gives rise to a large variation in depth and in the estimation of haze-free images. To alleviate the problem, we propose a plain binocular image dehazing network in this paper, called BidNet, to dehaze both the left and right images simultaneously. BidNet does not explicitly perform disparity estimation that is time-consuming and well-known to be challenging. Instead, we design a stereo transformation module to mine the relationship and correlation between binocular images, making the best of varying information of cross views. Additionally, we design a Stereo Foggy Cityscapes dataset extended from the Foggy Cityscapes dataset for training the proposed BidNet. Extensive experimental results demonstrate that BidNet significantly outperforms the SOTA dehazing methods on the synthetic stereo foggy datasets as well as in real stereo foggy scenes. Experimental results show that jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. Furthermore, when applying BidNet to preprocess foggy inputs, large improvements are obtained in the performance of object detection, instance segmentation, semantic segmentation, and stereo-based 3D object detection.
Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han, Xuelong Li 0001
IEEE Trans. Multim.3
2023 Object Detection in Foggy Images with Transmission Map Guidance
Jin Xie 0005, Jing Nie 0001
ICANN (7)2
2023 C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection
abstract
Multi-modal 3D object detection that classifies and locates objects in 3D space by combining point-clouds captured by lidars and RGB images captured by cameras, serves as the basis for autonomous driving. Most of the existing methods aggregate features from point-clouds and images by plain element-wise additions or multiplications. Although these methods improve detection accuracy, such simple operations have difficulties in balancing both modalities. Further, the multi-level features from images also suffer from imbalance problems in receptive fields. To address the above problems, we propose two novel networks: cross-modality balance network (CMN) and cross-scale balance network (CSN). CMN utilizes cross-modality attention mechanisms to balance the importance and receptive field of two modalities. CSN employs cross-scale attention mechanisms to reduce the imbalance in multi-level features. Experiments are performed on the challenging benchmark: KITTI. The experimental results show consistent improvements in different 3D object detection frameworks, which verifies the effectiveness and generality of our proposed networks.
Bonan Ding, Jin Xie 0005, Jing Nie 0001
ICASSP2
2023 DDNet: Density and depth-aware network for object detection in foggy scenes
abstract
Fog causes serious degradation in image quality that in turn can degrade the performance of object detection. The main reason can be concluded that (i) the degraded images make object localization difficult, (ii) the difficulty in extracting robust features for accurate detection results in various fog densities. To address the above two problems, in this paper, we propose a simple yet efficient network named density and depth-aware network (DDNet), which consists of a density-aware attention network (DAANet) and a depth-aware non-local contextual network (DNCNet). The DNCNet captures long-range dependencies guided by depth information to improve object localization. DAANet employs an attention mechanism guided by predicted fog densities to ensure the robustness of features under different fog densities. Experiments are performed on the FoggyDriving dataset. Our approach achieves the state-of-the-art performance.
Boyi Xiao, Jin Xie 0005, Jing Nie 0001
IJCNN2
2023 Attentive Alignment Network for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection is of great importance in various around-the-clock applications, i.e., self-driving and video surveillance. Fusing the features from RGB images and thermal infrared (TIR) images to explore the complementary information between different modalities is one of the most effective manners to improve multispectral pedestrian detection performance. However, the misalignment between different modalities in spatial dimension and modality reliability would introduce harmful information during feature fusion, limiting the performance of multispectral pedestrian detection. To address the above issues, we propose an attentive alignment network, consisting of an attentive position alignment (APA) module and an attentive modality alignment (AMA) module. Our APA module emphasizes pedestrian regions while aligning the pedestrian regions between different modalities. Our AMA module utilizes a channel-wise attention mechanism with illumination guidance to eliminate the imbalance between different modalities. The experiments are conducted on two widely used multispectral detection datasets, KASIT and CVC-14. Our approach surpasses the current state-of-the-art performance on both datasets.
Nuo Chen 0003, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang
ACM Multimedia2
2023 Context and detail interaction network for stereo rain streak and raindrop removal
Jing Nie 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang
Neural Networks2
2023 Latent Feature Pyramid Network for Object Detection
abstract
Object detection methods based on Convolution Neural Networks (CNN) usually utilize feature pyramid networks to detect objects with various scales. The state-of-the-art feature pyramid networks improve detection accuracy by enhancing multi-level feature representations. Fusing multi-level features is the most effective manner to enhance the feature representations. However, the existing feature pyramid networks usually fuse multi-level features by element-wise operations. It leads to the lack of long-range dependencies in the feature fusion. To address the problem, we propose a simple yet efficient feature pyramid network named latent feature pyramid network (LFPN). LFPN can enhance the feature representations by modeling inner-scale and cross-scale long-range dependencies through conducting inner-scale and cross-scale feature fusion in the latent space. Comprehensive experiments are performed on two challenge object detection datasets: MS COCO and Pascal VOC. The experimental results show consistent improvements on various feature pyramid networks, backbones, and object detectors, which demonstrates the effectiveness and generality of our LFPN.
Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han
IEEE Trans. Multim.1
2023 Complementary Feature Pyramid Network for Object Detection
abstract
The way of constructing a robust feature pyramid is crucial for object detection. However, existing feature pyramid methods, which aggregate multi-level features by using element-wise sum or concatenation, are inefficient to construct a robust feature pyramid. The reason is that these methods cannot be effective in discriminating the relevant semantics of objects. In this article, we propose a Complementary Feature Pyramid Network (CFPN) to aggregate multi-level features selectively and efficiently by exploring complementary information between multi-level features. Specifically, a Spatial Complementary Module (SCM) and a Channel Complementary Module (CCM) are designed and embedded in CFPN to enhance useful information and suppress irrelevant information during feature fusions along spatial and channel dimensions, respectively. CFPN is a generic feature extractor, as evidenced by its seamless integration into single-stage, two-stage, and end-to-end object detectors. Experiments conducted on the COCO and Pascal VOC datasets demonstrate that integrating our CFPN into RetinaNet, Faster RCNN, Cascade RCNN, and Sparse RCNN obtains consistent performance improvements with negligible overheads. Code and models are available at: https://github.com/VIPLab-CQU/CFPN .
Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han
ACM Trans. Multim. Comput. Commun. Appl.1
2022 PSTR: End-to-End One-Step Person Search With Transformers
abstract
We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along with a discriminative re-id decoder for person re-id. The discriminative re-id decoder utilizes a multi-level supervision scheme with a shared decoder for discriminative re-id feature learning and also comprises a part attention block to encode relationship between different parts of a person. We further introduce a simple multi-scale scheme to support re-id across person instances at different scales. PSTR jointly achieves the diverse objectives of object-level recognition (detection) and instance-level matching (re-id). To the best of our knowledge, we are the first to propose an end-to-end one-step transformer-based person search framework. Experiments are performed on two popular benchmarks: CUHK-SYSU and PRW. Our extensive ablations reveal the merits of the proposed contributions. Further, the proposed PSTR sets a new state-of-the-art on both benchmarks. On the challenging PRW benchmark, PSTR achieves a mean average precision (mAP) score of 56.5%. The source code is available at https://github.com/JialeCao001/PSTR.
Jiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal, Jin Xie 0005, Mubarak Shah, Fahad Shahbaz Khan
CVPR5
2022 Learning a Dynamic Cross-Modal Network for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection that enables continuous (day and night) localization of pedestrians has numerous applications. Existing approaches typically aggregate multispectral features by a simple element-wise operation. However, such a local feature aggregation scheme ignores the rich non-local contextual information. Further, we argue that a local tight correspondence across modalities is desired for multi-modal feature aggregation. To address these issues, we introduce a multispectral pedestrian detection framework that comprises a novel dynamic cross-modal network (DCMNet), which strives to adaptively utilize the local and non-local complementary information between multi-modal features. The proposed DCMNet consists of a local and a non-local feature aggregation module. The local module employs dynamically learned convolutions to capture local relevant information across modalities. On the other hand, the non-local module captures non-local cross-modal information by first projecting features from both modalities into the latent space and then obtaining dynamic latent feature nodes for feature aggregation. Comprehensive experiments are performed on two challenging benchmarks: KAIST and LLVIP. Experiments reveal the benefits of the proposed DCMNet, leading to consistently improved detection performance on diverse detection paradigms and backbones. When using the same backbone, our proposed detector achieves absolute gains of 1.74% and 1.90% over the baseline Cascade RCNN on the KAIST and LLVIP datasets.
Jin Xie 0005, Rao Muhammad Anwer, Hisham Cholakkal, Jing Nie 0001, Jiale Cao, Jorma Laaksonen, Fahad Shahbaz Khan
ACM Multimedia1
2022 From Handcrafted to Deep Features for Pedestrian Detection: A Survey
abstract
Pedestrian detection is an important but challenging problem in computer vision, especially in human-centric tasks. Over the past decade, significant improvement has been witnessed with the help of handcrafted features and deep features. Here we present a comprehensive survey on recent advances in pedestrian detection. First, we provide a detailed review of single-spectral pedestrian detection that includes handcrafted features based methods and deep features based approaches. For handcrafted features based methods, we present an extensive review of approaches and find that handcrafted features with large freedom degrees in shape and space have better performance. In the case of deep features based approaches, we split them into pure CNN based methods and those employing both handcrafted and CNN based features. We give the statistical analysis and tendency of these methods, where feature enhanced, part-aware, and post-processing methods have attracted main attention. In addition to single-spectral pedestrian detection, we also review multi-spectral pedestrian detection, which provides more robust features for illumination variance. Furthermore, we introduce some related datasets and evaluation metrics, and a deep experimental analysis. We conclude this survey by emphasizing open problems that need to be addressed and highlighting various future directions. Researchers can track an up-to-date list at https://github.com/JialeCao001/PedSurvey.
Jiale Cao, Yanwei Pang, Jin Xie 0005, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Stereo Refinement Dehazing Network
abstract
The performance of stereo vision tasks degrades when haze exists in the input stereo image pair. Independently applying single image dehazing algorithm on left and right images is not optimal. To overcome the problem, we propose an effective framework, called SRDNet, for simultaneously dehazing stereo images. The main idea of SRDNet is to make full use of the stereo information from cross views improving dehazing performance. It does not explicitly employ the disparity estimation and the correlation matrix. SRDNet comprises two parts: a weight-sharing coarse dehazing network (WSCDN) and a guided separated refinement network (GSRN). The WSCDN is utilized to predict a coarse dehazed image pair. Then the GSRN is introduced to predict the residues for different views by extracting the fused information of cross views and separating the features of different views with a guided channel and spatial refinement module. The residues are added to the coarse dehazed pair so as to make refinement and remove the remained haze. Experimental results demonstrate that our proposed SRDNet surpasses previous image dehazing methods by a significant margin both quantitatively and qualitatively. Moreover, our SRDNet could be a preprocessing step of the stereo-based 3D object detection and boosts the 3D detection accuracy in hazy scenes.
Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.3
2021 PSC-Net: learning part spatial co-occurrence for occluded pedestrian detection
Jin Xie 0005, Yanwei Pang, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
Sci. China Inf. Sci.1
2021 TJU-DHD: A Diverse High-Resolution Dataset for Object Detection
abstract
Vehicles, pedestrians, and riders are the most important and interesting objects for the perception modules of self-driving vehicles and video surveillance. However, the state-of-the-art performance of detecting such important objects (esp. small objects) is far from satisfying the demand of practical systems. Large-scale, rich-diversity, and high-resolution datasets play an important role in developing better object detection methods to satisfy the demand. Existing public large-scale datasets such as MS COCO collected from websites do not focus on the specific scenarios. Moreover, the popular datasets (e.g., KITTI and Citypersons) collected from the specific scenarios are limited in the number of images and instances, the resolution, and the diversity. To attempt to solve the problem, we build a diverse high-resolution dataset (called TJU-DHD). The dataset contains 115354 high-resolution images (52% images have a resolution of 1624×1200 pixels and 48% images have a resolution of at least 2, 560×1.440 pixels) and 709 330 labeled objects in total with a large variance in scale and appearance. Meanwhile, the dataset has a rich diversity in season variance, illumination variance, and weather variance. In addition, a new diverse pedestrian dataset is further built. With the four different detectors (i.e., the one-stage RetinaNet, anchor-free FCOS, two-stage FPN, and Cascade R-CNN), experiments about object detection and pedestrian detection are conducted. We hope that the newly built dataset can help promote the research on object detection and pedestrian detection in these two scenes. The dataset is available at https://github.com/tjubiit/TJU-DHD.
Yanwei Pang, Jiale Cao, Yazhao Li, Jin Xie 0005, Hanqing Sun 0001, Jinfeng Gong
IEEE Trans. Image Process.4
2021 Mask-Guided Attention Network and Occlusion-Sensitive Hard Example Mining for Occluded Pedestrian Detection
abstract
Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving other pedestrians and inter-class occlusions caused by other objects, such as cars and bicycles. These result in a multitude of occlusion patterns. We propose an approach for occluded pedestrian detection with the following contributions. First, we introduce a novel mask-guided attention network that fits naturally into popular pedestrian detection pipelines. Our attention network emphasizes on visible pedestrian regions while suppressing the occluded ones by modulating full body features. Second, we propose the occlusion-sensitive hard example mining method and occlusion-sensitive loss that mines hard samples according to the occlusion level and assigns higher weights to the detection errors occurring at highly occluded pedestrians. Third, we empirically demonstrate that weak box-based segmentation annotations provide reasonable approximation to their dense pixel-wise counterparts. Experiments are performed on CityPersons, Caltech and ETH datasets. Our approach sets a new state-of-the-art on all three datasets. Our approach obtains an absolute gain of 10.3% in log-average miss rate, compared with the best reported results on the heavily occluded HO pedestrian set of the CityPersons test set. Code and models are available at: https://github.com/Leotju/MGAN.
Jin Xie 0005, Yanwei Pang, Muhammad Haris Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Image Process.1
2020 BidNet: Binocular Image Dehazing Without Explicit Disparity Estimation
abstract
Heavy haze results in severe image degradation and thus hampers the performance of visual perception, object detection, etc. On the assumption that dehazed binocular images are superior to the hazy ones for stereo vision tasks such as 3D object detection and according to the fact that image haze is a function of depth, this paper proposes a Binocular image dehazing Network (BidNet) aiming at dehazing both the left and right images of binocular images within the deep learning framework. Existing binocular dehazing methods rely on simultaneously dehazing and estimating disparity, whereas BidNet does not need to explicitly perform time-consuming and well-known challenging disparity estimation. Note that a small error in disparity gives rise to a large variation in depth and in estimation of haze-free image. The relationship and correlation between binocular images are explored and encoded by the proposed Stereo Transformation Module (STM). Jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. We extend the Foggy Cityscapes dataset to a Stereo Foggy Cityscapes dataset with binocular foggy image pairs. Experimental results demonstrate that BidNet significantly outperforms state-of-the-art dehazing methods in both subjective and objective assessments.
Yanwei Pang, Jing Nie 0001, Jin Xie 0005, Jungong Han, Xuelong Li 0001
CVPR3
2020 Count- and Similarity-Aware R-CNN for Pedestrian Detection
Jin Xie 0005, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001, Mubarak Shah
ECCV (17)1
2020 Spectral Clustering by Joint Spectral Embedding and Spectral Rotation
abstract
Spectral clustering is an important clustering method widely used for pattern recognition and image segmentation. Classical spectral clustering algorithms consist of two separate stages: 1) solving a relaxed continuous optimization problem to obtain a real matrix followed by 2) applying K -means or spectral rotation to round the real matrix (i.e., continuous clustering result) into a binary matrix called the cluster indicator matrix. Such a separate scheme is not guaranteed to achieve jointly optimal result because of the loss of useful information. To obtain a better clustering result, in this paper, we propose a joint model to simultaneously compute the optimal real matrix and binary matrix. The existing joint model adopts an orthonormal real matrix to approximate the orthogonal but nonorthonormal cluster indicator matrix. It is noted that only in a very special case (i.e., all clusters have the same number of samples), the cluster indicator matrix is an orthonormal matrix multiplied by a real number. The error of approximating a nonorthonormal matrix is inevitably large. To overcome the drawback, we propose replacing the nonorthonormal cluster indicator matrix with a scaled cluster indicator matrix which is an orthonormal matrix. Our method is capable of obtaining better performance because it is easy to minimize the difference between two orthonormal matrices. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed method (called JSESR).
Yanwei Pang, Jin Xie 0005, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Cybern.2
2019 Mask-Guided Attention Network for Occluded Pedestrian Detection
abstract
Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving other pedestrians and inter-class occlusions caused by other objects, such as cars and bicycles. These results in a multitude of occlusion patterns. We propose an approach for occluded pedestrian detection with the following contributions. First, we introduce a novel mask-guided attention network that fits naturally into popular pedestrian detection pipelines. Our attention network emphasizes on visible pedestrian regions while suppressing the occluded ones by modulating full body features. Second, we empirically demonstrate that coarse-level segmentation annotations provide reasonable approximation to their dense pixel-wise counterparts. Experiments are performed on CityPersons and Caltech datasets. Our approach sets a new state-of-the-art on both datasets. Our approach obtains an absolute gain of 9.5% in log-average miss rate, compared to the best reported results [32] on the heavily occluded HO pedestrian set of CityPersons test set. Further, on the HO pedestrian set of Caltech dataset, our method achieves an absolute gain of 5.0% in log-average miss rate, compared to the best reported results [13]. Code and models are available at: https://github.com/Leotju/MGAN.
Yanwei Pang, Jin Xie 0005, Muhammad Haris Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
ICCV2
2019 Visual Haze Removal by a Unified Generative Adversarial Network
abstract
Existence of haze significantly degrades visual quality and hence negatively affects the performance of visual surveillance, video analysis, and human–machine interaction. To remove haze from a visual signal, in this paper, we propose a generative adversarial network for visual haze removal called HRGAN. HRGAN consists of a generator network and a discriminator network. A unified network jointly estimating transmission maps, atmospheric light, and haze-free images (called UNTA) is proposed as the generator network of HRGAN. Instead of being optimized by minimizing the pixel-wise loss, HRGAN is optimized by minimizing a novel loss function consisting of pixel-wise loss, perceptual loss, and adversarial loss produced by a discriminator network. Classical model-based image dehazing algorithms consist of three separate stages: 1) estimating transmission map; 2) estimating atmospheric light; and 3) restoring haze-free image by using an atmospheric scattering model to process the transmission map and atmospheric light. Such a separate scheme is not guaranteed to achieve optimal results. On the contrary, UNTA performs transmission map estimation and atmospheric light estimation simultaneously to obtain joint optimal solutions. The experimental results on both synthetic and real-world image databases demonstrate that HRGAN outperforms the state-of-the-art algorithms in terms of both effectiveness and efficiency.
Yanwei Pang, Jin Xie 0005, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2