VLDB 2026 Research / reviewers in the wild / expert
Jing Nie 0001
dblp:39/5493-1
· DBLP profile ↗
39ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0001-9872-9286ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remote sensing optical image matching through neighborhood-aware global propagation in graph neural networks
Yanchun Liu, Gemine Vivone, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou, Lihui Chen 0002 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | PM-adapter: MoE based dynamic denoising fine-tuning for thermal infrared object detection
Haijun Liu 0001, Boya Wei, Jing Nie 0001, Suju Li, Xichuan Zhou |
Neurocomputing | 5 |
| 2026 | Sparse gain adaptation with dual-domain fusion network for multimodal object detection
Xichuan Zhou, Boya Wei, Cong Mao, Lihui Chen 0002, Haijun Liu 0001, Jin Xie 0005, Jing Nie 0001 |
Neurocomputing | 9 |
| 2026 | RA-PTQ: Reparameterization-Aware Post-Training Quantization for accurate vision transformers in low-bit scenarios
Rui Ding 0009, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Knowl. Based Syst. | 4 |
| 2026 | Multispectral remote sensing object detection via selective cross-modal interaction and aggregation
Minghao Cui, Jing Nie 0001, Hanqing Sun 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
Neural Networks | 2 |
| 2026 | SNNSIR: A fully Spiking Neural Network for Stereo Image Restoration
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang |
Pattern Recognit. | 3 |
| 2026 | Frequency-Decomposed Interaction Network for Stereo Image RestorationabstractStereo image restoration in adverse environments, such as low-light conditions, rain, and low resolution, requires effective exploitation of cross-view complementary information to recover degraded visual content. In monocular image restoration, frequency decomposition has proven effective, where high-frequency components aid in recovering fine textures and reducing blur, while low-frequency components facilitate noise suppression and illumination correction. However, existing stereo restoration methods have yet to explore cross-view interactions by frequency decomposition, which is a promising direction for enhancing restoration quality. To address this, we propose a frequency-aware framework comprising a Frequency Decomposition Module (FDM), Detail Interaction Module (DIM), Structural Interaction Module (SIM), and Adaptive Fusion Module (AFM). FDM employs learnable filters to decompose the image into high- and low-frequency components. DIM enhances the high-frequency branch by capturing local detail cues through deformable convolution. SIM processes the low-frequency branch by modeling global structural correlations via a cross-view row-wise attention mechanism. Finally, AFM adaptively fuses the complementary frequency-specific information to generate high-quality restored images. Extensive experiments demonstrate the efficacy and generalizability of our framework across three diverse stereo restoration tasks, where it achieves state-of-the-art performance in low-light enhancement, rain removal, alongside highly competitive results in super-resolution. Our code is available at https://github.com/C2022J/FDIN. Xianmin Tian, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Multimodality Image Registration With Modality DistillationabstractMultimodal image registration aims to spatially align images from different modalities at the pixel level. However, due to the nonlinear relationship of radiation intensities caused by different imaging modalities, achieving high accuracy in multimodal image registration presents a significant challenge. Additionally, the presence of both global transformations (i.e., large-scale rigid affine transformations) and local distortions (i.e., small-scale nonrigid deformations) between paired images further complicates the registration process. This article addressed the challenge resulting from modality differences through modality distillation. Specifically, a teacher (i.e., a homomodal image registration model) is trained to guide the student (i.e., a multimodal image registration model). Besides, this article simultaneously aligned large-scale rigid and small-scale nonrigid deformations by predicting deformation flow from both global and local features, thereby achieving high-precision registration. Furthermore, this proposed method incorporated a deformation mask during training to mitigate the negative impact of black edges in the obtained registration results on model performance. Experimental results demonstrate that the proposed method delivers state-of-the-art registration accuracy across various multimodal datasets, with ablation studies confirming the effectiveness of each component. The codes will be available at https://github.com/2351056918/Multimodality-Image-Registration-with-Modailty-Distillation. Xichuan Zhou, Jicheng Zhao, Lihui Chen 0002, Gemine Vivone, Yanchun Liu, Jing Nie 0001, Haijun Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | CAPilot: A High-Performance and High-Reliability Communication Middleware for Autonomous DrivingabstractWith the swift advancement of artificial intelligence technology, autonomous driving has increasingly emerged as a pivotal technology in the future of transportation. Real-time data exchange and processing across modules in autonomous driving systems necessitate efficient and reliable communication middleware. However, existing communication methods suffer from delay, congestion and packet loss when dealing with high-frequency and large-data-volume transmission tasks, significantly impairing system performance and security. To reduce communication latency and CPU overhead, a multi-mode adaptive high-performance and high-reliability communication middleware CAPilot is proposed. Firstly, a novel shared-memory communication architecture is proposed, comprising a Data Pool, an Event Notification Index Pool, and a Cycle Index Pool. The Data Pool employs a lock-free mechanism to avert deadlock and starvation issues, while addressing frame-skipping using a real-time maintenance and discriminative approach. Event-triggered and period-triggered data acquisition strategies proficiently circumvent data security concerns and performance limitations inherent in conventional shared memory connectivity. Then, to mitigate the overhead associated with dynamic broadcasts within the constrained embedded resources of the network, an adaptive communication scheme is proposed. This scheme incorporates a profile-based static communication encoding that automatically determines the optimal communication method based on the environments of the communicating entities. Finally, the intra-process pointer passing method is optimised by introducing a dual adaptive buffered ring queue, which facilitates bulk data retrieval without using locks. Experimental results show that CAPilot outperforms existing communication middlewares such as ROS2, CyberRT, and DDS in terms of communication latency, message throughput, message frame loss rate, and resource utilisation. These advancements suggest that CAPilot is well-suited for extensive deployment in diverse autonomous driving applications. Pinzhong Qin, Changquan Xue, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou |
ACM Trans. Internet Techn. | 4 |
| 2025 | SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object DetectionabstractMultimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually aggregate multimodal features at a single stage. However, leveraging multi-stage cross-modal features is crucial for detecting objects of various scales. Therefore, these methods often struggle to integrate features across different scales and modalities effectively, thereby restricting the accuracy of detection. Additionally, the time-consuming Query-Key-Value-based (QKV-based) cross-attention operations often utilized in existing methods aid in reasoning the location and existence of objects by capturing non-local contexts. However, this approach tends to increase computational complexity. To address these challenges, we present SSLFusion, a novel Scale & Space Aligned Latent Fusion Model, consisting of a scale-aligned fusion strategy (SAF), a 3D-to-2D space alignment module (SAM), and a latent cross-modal fusion module (LFM). SAF mitigates scale misalignment between modalities by aggregating features from both images and point clouds across multiple levels. SAM is designed to reduce the inter-modal gap between features from images and point clouds by incorporating 3D coordinate information into 2D image features. Additionally, LFM captures cross-modal non-local contexts in the latent space without utilizing the QKV-based attention operations, thus mitigating computational complexity. Experiments on the KITTI and DENSE datasets demonstrate that our SSLFusion outperforms state-of-the-art methods. Our approach obtains an absolute gain of 2.15% in 3D AP, compared with the state-of-art method GraphAlign on the moderate level of the KITTI test set. Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao |
AAAI | 3 |
| 2025 | EigenSR: Eigenimage-Bridged Pre-Trained RGB Learners for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution (single-HSI-SR) aims to improve the resolution of a single input low-resolution HSI. Due to the bottleneck of data scarcity, the development of single-HSI-SR lags far behind that of RGB natural images. In recent years, research on RGB SR has shown that models pre-trained on large-scale benchmark datasets can greatly improve performance on unseen data, which may stand as a remedy for HSI. But how can we transfer the pre-trained RGB model to HSI, to overcome the data-scarcity bottleneck? Because of the significant difference in the channels between the pre-trained RGB model and the HSI, the model cannot focus on the correlation along the spectral dimension, thus limiting its ability to utilize on HSI. Inspired by the HSI spatial-spectral decoupling, we propose a new framework that first fine-tunes the pre-trained model with the spatial components (known as eigenimages), and then infers on unseen HSI using an iterative spectral regularization (ISR) to maintain the spectral correlation. The advantages of our method lie in: 1) we effectively inject the spatial texture processing capabilities of the pre-trained RGB model into HSI while keeping spectral fidelity, 2) learning in the spectral-decorrelated domain can improve the generalizability to spectral-agnostic data, and 3) our inference in the eigenimage domain naturally exploits the spectral low-rank property of HSI, thereby reducing the complexity. This work bridges the gap between pre-trained RGB models and HSI via eigenimages, addressing the issue of limited HSI training data, hence the name EigenSR. Extensive experiments show that EigenSR outperforms the state-of-the-art (SOTA) methods in both spatial and spectral metrics. Xi Su, Xiangfei Shen, Mingyang Wan, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
AAAI | 4 |
| 2025 | DANet: spatial gene expression prediction from H&E histology images through dynamic alignmentabstractPredicting spatial gene expression from Hematoxylin and Eosin histology images offers a promising approach to significantly reduce the time and cost associated with gene expression sequencing, thereby facilitating a deeper understanding of tissue architecture and disease mechanisms. Achieving accurate gene expression prediction requires the extraction of highly refined features from pathological images; however, existing methods often struggle to effectively capture fine-grained local details and model gene-gene correlations. Moreover, in bimodal contrastive learning, dynamically and efficiently aligning heterogeneous modalities remains a critical challenge. To address these issues, we propose a novel method for predicting gene expression. First, we introduce a dense connective structure that enables efficient feature reuse, thereby enhancing the capturing and mining of local refinement features. Second, we leverage the state space models to uncover underlying patterns and capture dependencies within 1D gene expression data, enabling more accurate modeling of gene-gene correlations. Furthermore, we design the Residual Kolmogorov-Arnold Network (RKAN) that uses a learnable activation function to dynamically adjust bimodal mappings based on input characteristics. Through continuous parameter updates during contrastive training, RKAN progressively refines the alignment between modalities. Extensive experiments conducted on two publicly available datasets, GSE240429 and HER2+, demonstrate the effectiveness of our approach and its significant improvements over existing methods. Source codes are available at https://github.com/202324131016T/DANet. Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yuansong Zeng |
Briefings Bioinform. | 3 |
| 2025 | Hybrid cross-modality fusion network for medical image segmentation with contrastive learning
Xichuan Zhou, Jing Nie 0001, Haijun Liu 0001, Fu Liang, Lihui Chen 0002, Jin Xie 0005 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | E-TransConvNet: An enhanced transformer and convolutional network for medical image segmentation from ultrasound and CT images
Chukwuemeka Clinton Atabansi, Jing Nie 0001, Jiachen Huang, Haijun Liu 0001, Jin Xie 0005, Xichuan Zhou |
Expert Syst. Appl. | 2 |
| 2025 | Progressive fine-to-coarse reconstruction for accurate low-bit post-training quantization in vision transformers
Rui Ding 0009, Liang Yong, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Neural Networks | 4 |
| 2025 | MIT-SAM: Medical Image-Text SAM With Mutually Enhanced Heterogeneous Features Fusion for Medical Image SegmentationabstractIn recent times, leveraging lesion text as supplementary data to enhance the performance of medical image segmentation models has garnered attention. Previous approaches only used attention mechanisms to integrate image and text features, while not effectively utilizing the highly condensed textual semantic information in improving the fused features, resulting in inaccurate lesion segmentation. This paper introduces a novel approach, the Medical Image-Text Segment Anything Model (MIT-SAM), for text-assisted medical image segmentation. Specifically, we introduce the SAM-enhanced image encoder and a Bert-based text encoder to extract heterogeneous features. To better leverage the highly condensed textual semantic information for heterogeneous feature fusion, such as crucial details like position and quantity, we propose the image-text interactive fusion (ITIF) block and self-supervised text reconstruction (SSTR) method. The ITIF block facilitates the mutual enhancement of homogeneous information among heterogeneous features and the SSTR method empowers the model to capture crucial details concerning lesion text, including location, quantity, and other key aspects. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on the QaTa-COV19 and MosMedData+ datasets. Xichuan Zhou, Lingfeng Yan, Rui Ding 0009, Chukwuemeka Clinton Atabansi, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Dual Interaction and Kernel-Diverse Network for Accurate Drug-Target Binding Affinity PredictionabstractDrug-target binding affinities measure the binding strength between drugs and targets, serving as the foundation for computer-aided drug design. The existing methods based on neural network rarely incorporate knowledge from biochemistry, which limits their performance. In this paper, we propose a dual interaction and kernel-diverse network for drug-target binding affinity prediction, termed DTANet, drawing inspiration from the biochemical properties of drugs and targets. Specifically, since functional groups and peptide chains are pivotal in determining the properties of drugs and their targets, and their lengths vary significantly. Drawing inspiration from this, we propose a Kernel-diverse Two-stream Network (KTN) and a Cross-Scale Interaction Module (CSIM) to extract the features of drugs and targets. The binding site of drug and target is crucial for precise affinity prediction. Inspired by this, we introduce the Drug-Target Interaction Module (DTIM) to explore the relationships between drugs and targets to focus on binding sites. Experiments have been conducted on two popular benchmarks: KIBA and Davis. Our DTANet obtains superior performance compared to existing methods on both datasets. Source codes are available at https://github.com/202324131016T/DTANet. Jin Xie 0005, Jing Nie 0001, Yuansong Zeng |
BIBM | 3 |
| 2024 | Mamba-DTA: Drug-Target Binding Affinity Prediction with State Space ModelabstractBiochemical methods for measuring drug-target binding are costly and slow, while deep learning offers a crucial solution. Deep learning methods for predicting drug-target binding affinity have achieved remarkable success, but existing approaches face challenges: adjacent elements in a three-dimensional structure may be far apart in a one-dimensional representation, and long sequences of drug and target data often contain lots of noise. In this paper, we introduce MambaDTA, a novel architecture for drug-target affinity prediction based on the State Space Model (SSM). Mamba-DTA utilizes SSM to model the drug molecules and target molecules and extract more discriminative spatial structural features efficiently and stably. Additionally, we design Interaction-based Selective Filtering (ISF) module to model drug-target interactions and filter out redundant information. The experimental results on two publicly available datasets, namely Davis and KIBA, demonstrate the effectiveness and superiority of our Mamba-DTA. Specifically, Mamba-DTA achieves a relative gain of 13.3% in terms of MAE on the Davis dataset. Source codes are available at https://github.com/202324131016T/Mamba-DTA. Jin Xie 0005, Jing Nie 0001, Xiaohong Zhang 0002, Yuansong Zeng |
BIBM | 3 |
| 2024 | Localization-aware logit mimicking for object detection in adverse weather conditions
Peiyun Luo, Jing Nie 0001, Jin Xie 0005, Jiale Cao, Xiaohong Zhang 0002 |
Image Vis. Comput. | 2 |
| 2024 | C2BG-Net: Cross-modality and cross-scale balance network with global semantics for multi-modal 3D object detection
Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao |
Neural Networks | 3 |
| 2024 | Multi-query and multi-level enhanced network for semantic segmentation
Jiale Cao, Rao Muhammad Anwer, Jin Xie 0005, Jing Nie 0001, Ai-Ping Yang, Yanwei Pang |
Pattern Recognit. | 5 |
| 2024 | ESGN: Efficient Stereo Geometry Network for Fast 3D Object DetectionabstractFast stereo based 3D object detectors have made great progress recently. However, they suffer from the inferior accuracy. We argue that the main reason is due to the poor geometry-aware feature representation in 3D space. To solve this problem, we propose an efficient stereo geometry network (ESGN). The key in our ESGN is an efficient geometry-aware feature generation (EGFG) module. Our EGFG module first uses a stereo correlation and reprojection module to construct multi-scale stereo volumes in camera frustum space, second employs a multi-scale bird’s eye view (BEV) projection and fusion module to generate multiple geometry-aware features. In these two steps, we adopt deep multi-scale information fusion for discriminative geometry-aware feature generation, without any complex aggregation networks. In addition, we introduce a deep geometry-aware feature distillation scheme to guide stereo feature learning with a LiDAR-based detector. The experiments are performed on the classical KITTI dataset. On KITTI test set, our ESGN outperforms the fast state-of-art-art detector YOLOStereo3D by 5.14% on mAP3d at$62ms$. To the best of our knowledge, our ESGN achieves a best trade-off between accuracy and speed. We hope that our efficient stereo geometry network can provide more possible directions for fast 3D object detection. Aqi Gao, Yanwei Pang, Jing Nie 0001, Jiale Cao, Yishun Guo, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Binocular Image Dehazing via a Plain Network Without Disparity EstimationabstractHeavy haze leads to severely degraded visual quality for images, and thus the performance of high level image-based tasks such as object detection and semantic segmentation is deteriorated. It is necessary and important to design an effective dehazing method for the computer vision system. It is well known that image haze is a function of depth and binocular images can predict the depth. Existing binocular dehazing methods conduct disparity estimation and dehazing jointly to enhance each other. However, a small error in disparity gives rise to a large variation in depth and in the estimation of haze-free images. To alleviate the problem, we propose a plain binocular image dehazing network in this paper, called BidNet, to dehaze both the left and right images simultaneously. BidNet does not explicitly perform disparity estimation that is time-consuming and well-known to be challenging. Instead, we design a stereo transformation module to mine the relationship and correlation between binocular images, making the best of varying information of cross views. Additionally, we design a Stereo Foggy Cityscapes dataset extended from the Foggy Cityscapes dataset for training the proposed BidNet. Extensive experimental results demonstrate that BidNet significantly outperforms the SOTA dehazing methods on the synthetic stereo foggy datasets as well as in real stereo foggy scenes. Experimental results show that jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. Furthermore, when applying BidNet to preprocess foggy inputs, large improvements are obtained in the performance of object detection, instance segmentation, semantic segmentation, and stereo-based 3D object detection. Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Object Detection in Foggy Images with Transmission Map Guidance
Jin Xie 0005, Jing Nie 0001 |
ICANN (7) | 3 |
| 2023 | C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object DetectionabstractMulti-modal 3D object detection that classifies and locates objects in 3D space by combining point-clouds captured by lidars and RGB images captured by cameras, serves as the basis for autonomous driving. Most of the existing methods aggregate features from point-clouds and images by plain element-wise additions or multiplications. Although these methods improve detection accuracy, such simple operations have difficulties in balancing both modalities. Further, the multi-level features from images also suffer from imbalance problems in receptive fields. To address the above problems, we propose two novel networks: cross-modality balance network (CMN) and cross-scale balance network (CSN). CMN utilizes cross-modality attention mechanisms to balance the importance and receptive field of two modalities. CSN employs cross-scale attention mechanisms to reduce the imbalance in multi-level features. Experiments are performed on the challenging benchmark: KITTI. The experimental results show consistent improvements in different 3D object detection frameworks, which verifies the effectiveness and generality of our proposed networks. Bonan Ding, Jin Xie 0005, Jing Nie 0001 |
ICASSP | 3 |
| 2023 | DDNet: Density and depth-aware network for object detection in foggy scenesabstractFog causes serious degradation in image quality that in turn can degrade the performance of object detection. The main reason can be concluded that (i) the degraded images make object localization difficult, (ii) the difficulty in extracting robust features for accurate detection results in various fog densities. To address the above two problems, in this paper, we propose a simple yet efficient network named density and depth-aware network (DDNet), which consists of a density-aware attention network (DAANet) and a depth-aware non-local contextual network (DNCNet). The DNCNet captures long-range dependencies guided by depth information to improve object localization. DAANet employs an attention mechanism guided by predicted fog densities to ensure the robustness of features under different fog densities. Experiments are performed on the FoggyDriving dataset. Our approach achieves the state-of-the-art performance. Boyi Xiao, Jin Xie 0005, Jing Nie 0001 |
IJCNN | 3 |
| 2023 | Attentive Alignment Network for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection is of great importance in various around-the-clock applications, i.e., self-driving and video surveillance. Fusing the features from RGB images and thermal infrared (TIR) images to explore the complementary information between different modalities is one of the most effective manners to improve multispectral pedestrian detection performance. However, the misalignment between different modalities in spatial dimension and modality reliability would introduce harmful information during feature fusion, limiting the performance of multispectral pedestrian detection. To address the above issues, we propose an attentive alignment network, consisting of an attentive position alignment (APA) module and an attentive modality alignment (AMA) module. Our APA module emphasizes pedestrian regions while aligning the pedestrian regions between different modalities. Our AMA module utilizes a channel-wise attention mechanism with illumination guidance to eliminate the imbalance between different modalities. The experiments are conducted on two widely used multispectral detection datasets, KASIT and CVC-14. Our approach surpasses the current state-of-the-art performance on both datasets. Nuo Chen 0003, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang |
ACM Multimedia | 3 |
| 2023 | DIIK-Net: A full-resolution cross-domain deep interaction convolutional neural network for MR image reconstruction
Yanwei Pang, Jing Nie 0001 |
Neurocomputing | 5 |
| 2023 | Context and detail interaction network for stereo rain streak and raindrop removal
Jing Nie 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang |
Neural Networks | 1 |
| 2023 | InfraNet: Accurate forehead temperature measurement framework for people in the wild with monocular thermal infrared camera
Xichuan Zhou, Dongshan Lei, Chunqiao Long, Jing Nie 0001, Haijun Liu 0001 |
Neural Networks | 4 |
| 2023 | Matrix Factorization With Framelet and Saliency Priors for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection aims to separate sparse anomalies from low-rank background components. A variety of detectors have been proposed to identify anomalies, but most of them tend to emphasize characterizing backgrounds with multiple types of prior knowledge and limited information on anomaly components. To tackle these issues, this article simultaneously focuses on two components and proposes a matrix factorization method with framelet and saliency priors to handle the anomaly detection problem. We first employ a framelet to characterize nonnegative background representation coefficients, as they can jointly maintain sparsity and piecewise smoothness after framelet decomposition. We then exploit saliency prior knowledge to measure each pixel’s potential to be an anomaly. Finally, we incorporate the pure pixel index (PPI) with Reed-Xiaoli’s (RX) method to possess representative dictionary atoms. We solve the optimization problem using a block successive upper-bound minimization (BSUM) framework with guaranteed convergence. Experiments conducted on benchmark hyperspectral datasets demonstrate that the proposed method outperforms some state-of-the-art anomaly detection methods. Xiangfei Shen, Haijun Liu 0001, Jing Nie 0001, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Latent Feature Pyramid Network for Object DetectionabstractObject detection methods based on Convolution Neural Networks (CNN) usually utilize feature pyramid networks to detect objects with various scales. The state-of-the-art feature pyramid networks improve detection accuracy by enhancing multi-level feature representations. Fusing multi-level features is the most effective manner to enhance the feature representations. However, the existing feature pyramid networks usually fuse multi-level features by element-wise operations. It leads to the lack of long-range dependencies in the feature fusion. To address the problem, we propose a simple yet efficient feature pyramid network named latent feature pyramid network (LFPN). LFPN can enhance the feature representations by modeling inner-scale and cross-scale long-range dependencies through conducting inner-scale and cross-scale feature fusion in the latent space. Comprehensive experiments are performed on two challenge object detection datasets: MS COCO and Pascal VOC. The experimental results show consistent improvements on various feature pyramid networks, backbones, and object detectors, which demonstrates the effectiveness and generality of our LFPN. Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han |
IEEE Trans. Multim. | 3 |
| 2023 | Complementary Feature Pyramid Network for Object DetectionabstractThe way of constructing a robust feature pyramid is crucial for object detection. However, existing feature pyramid methods, which aggregate multi-level features by using element-wise sum or concatenation, are inefficient to construct a robust feature pyramid. The reason is that these methods cannot be effective in discriminating the relevant semantics of objects. In this article, we propose a Complementary Feature Pyramid Network (CFPN) to aggregate multi-level features selectively and efficiently by exploring complementary information between multi-level features. Specifically, a Spatial Complementary Module (SCM) and a Channel Complementary Module (CCM) are designed and embedded in CFPN to enhance useful information and suppress irrelevant information during feature fusions along spatial and channel dimensions, respectively. CFPN is a generic feature extractor, as evidenced by its seamless integration into single-stage, two-stage, and end-to-end object detectors. Experiments conducted on the COCO and Pascal VOC datasets demonstrate that integrating our CFPN into RetinaNet, Faster RCNN, Cascade RCNN, and Sparse RCNN obtains consistent performance improvements with negligible overheads. Code and models are available at: https://github.com/VIPLab-CQU/CFPN . Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Learning a Dynamic Cross-Modal Network for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection that enables continuous (day and night) localization of pedestrians has numerous applications. Existing approaches typically aggregate multispectral features by a simple element-wise operation. However, such a local feature aggregation scheme ignores the rich non-local contextual information. Further, we argue that a local tight correspondence across modalities is desired for multi-modal feature aggregation. To address these issues, we introduce a multispectral pedestrian detection framework that comprises a novel dynamic cross-modal network (DCMNet), which strives to adaptively utilize the local and non-local complementary information between multi-modal features. The proposed DCMNet consists of a local and a non-local feature aggregation module. The local module employs dynamically learned convolutions to capture local relevant information across modalities. On the other hand, the non-local module captures non-local cross-modal information by first projecting features from both modalities into the latent space and then obtaining dynamic latent feature nodes for feature aggregation. Comprehensive experiments are performed on two challenging benchmarks: KAIST and LLVIP. Experiments reveal the benefits of the proposed DCMNet, leading to consistently improved detection performance on diverse detection paradigms and backbones. When using the same backbone, our proposed detector achieves absolute gains of 1.74% and 1.90% over the baseline Cascade RCNN on the KAIST and LLVIP datasets. Jin Xie 0005, Rao Muhammad Anwer, Hisham Cholakkal, Jing Nie 0001, Jiale Cao, Jorma Laaksonen, Fahad Shahbaz Khan |
ACM Multimedia | 4 |
| 2022 | Improving 2D object detection with binocular images for outdoor surveillance
Fuchen Chu, Yanwei Pang, Jiale Cao, Jing Nie 0001, Xuelong Li 0001 |
Neurocomputing | 4 |
| 2022 | Stereo Refinement Dehazing NetworkabstractThe performance of stereo vision tasks degrades when haze exists in the input stereo image pair. Independently applying single image dehazing algorithm on left and right images is not optimal. To overcome the problem, we propose an effective framework, called SRDNet, for simultaneously dehazing stereo images. The main idea of SRDNet is to make full use of the stereo information from cross views improving dehazing performance. It does not explicitly employ the disparity estimation and the correlation matrix. SRDNet comprises two parts: a weight-sharing coarse dehazing network (WSCDN) and a guided separated refinement network (GSRN). The WSCDN is utilized to predict a coarse dehazed image pair. Then the GSRN is introduced to predict the residues for different views by extracting the fused information of cross views and separating the features of different views with a guided channel and spatial refinement module. The residues are added to the coarse dehazed pair so as to make refinement and remove the remained haze. Experimental results demonstrate that our proposed SRDNet surpasses previous image dehazing methods by a significant margin both quantitatively and qualitatively. Moreover, our SRDNet could be a preprocessing step of the stereo-based 3D object detection and boosts the 3D detection accuracy in hazy scenes. Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Efficient Selective Context Network for Accurate Object DetectionabstractSingle-stage detectors have gained great attention due to their high detection accuracy and real-time speed. To detect multi-scale objects, single-stage detectors make scale-aware predictions based on multiple pyramid layers. However, the insufficient context exploration in shallow pyramid layers leads to the detection accuracy of small objects being far from satisfactory. To tackle this problem, we propose a scheme to selectively extract multi-scale context with attention-adaptive weights. Specifically, we propose an efficient selective context network for accurate object detection. It incorporates an enhanced context module and a triple attention module. The enhanced context module consists of multi-branches to extract original-scale, small-scale, and large-scale contextual information. To make full use of this context and filter out noisy information, the triple attention module, which contains global-level, channel-level, and spatial-level attentions, is introduced to carry out selective context fusion. The two modules are easy to implement and can efficiently boost the accuracy of object detection. The performance of our method is validated on two benchmarks: PASCAL VOC and MS COCO. For a 512×512 input, our detector with VGG16 achieves competitive results (80.9 on the Pascal VOC 2012 test set in the case of single-scale inference without MS COCO pre-training). On the MS COCO test-dev set, our detector with ResNet101 outperforms RetinaNet500 by 2.5% AP in terms of overall performance and its speed is 48 milliseconds on a Titan XP GPU. As a result, ESCNet achieves a better trade-off between accuracy and speed. Jing Nie 0001, Yanwei Pang, Shengjie Zhao 0001, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | BidNet: Binocular Image Dehazing Without Explicit Disparity EstimationabstractHeavy haze results in severe image degradation and thus hampers the performance of visual perception, object detection, etc. On the assumption that dehazed binocular images are superior to the hazy ones for stereo vision tasks such as 3D object detection and according to the fact that image haze is a function of depth, this paper proposes a Binocular image dehazing Network (BidNet) aiming at dehazing both the left and right images of binocular images within the deep learning framework. Existing binocular dehazing methods rely on simultaneously dehazing and estimating disparity, whereas BidNet does not need to explicitly perform time-consuming and well-known challenging disparity estimation. Note that a small error in disparity gives rise to a large variation in depth and in estimation of haze-free image. The relationship and correlation between binocular images are explored and encoded by the proposed Stereo Transformation Module (STM). Jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. We extend the Foggy Cityscapes dataset to a Stereo Foggy Cityscapes dataset with binocular foggy image pairs. Experimental results demonstrate that BidNet significantly outperforms state-of-the-art dehazing methods in both subjective and objective assessments. Yanwei Pang, Jing Nie 0001, Jin Xie 0005, Jungong Han, Xuelong Li 0001 |
CVPR | 2 |
| 2019 | Enriched Feature Guided Refinement Network for Object DetectionabstractWe propose a single-stage detection framework that jointly tackles the problem of multi-scale object detection and class imbalance. Rather than designing deeper networks, we introduce a simple yet effective feature enrichment scheme to produce multi-scale contextual features. We further introduce a cascaded refinement scheme which first instills multi-scale contextual features into the prediction layers of the single-stage detector in order to enrich their discriminative power for multi-scale detection. Second, the cascaded refinement scheme counters the class imbalance problem by refining the anchors and enriched features to improve classification and regression. Experiments are performed on two benchmarks: PASCAL VOC and MS COCO. For a 320×320 input on the MS COCO test-dev, our detector achieves state-of-the-art single-stage detection accuracy with a COCO AP of 33.2 in the case of single-scale inference, while operating at 21 milliseconds on a Titan XP GPU. For a 512×512 input on the MS COCO test-dev, our approach obtains an absolute gain of 1.6% in terms of COCO AP, compared to the best reported single-stage results[5]. Source code and models are available at: https://github.com/Ranchentx/EFGRNet. Jing Nie 0001, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001 |
ICCV | 1 |