VLDB 2026 Research / reviewers in the wild / expert
Jinchang Ren
dblp:25/921
· DBLP profile ↗
136ranked-venue papers
26as first author
72since 2021 · last 2026
0000-0001-6116-3194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 53 · 4 first-author · 36 since 2021Artificial intelligence and machine learning · 44 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 15 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attribution-driven background modeling network with context-aware anomaly suppression for hyperspectral anomaly detection
Yu Huo 0001, Min Zhang 0015, Jinchang Ren, Hai Wang 0015 |
Inf. Process. Manag. | 4 |
| 2026 | FACT: Feature Adaptive Continual-learning Tracker for multiple object tracking
Rongzihan Song, Zhenyu Weng, Huiping Zhuang, Jinchang Ren, Yongming Chen, Zhiping Lin 0001 |
Knowl. Based Syst. | 4 |
| 2026 | GCMNet: A global context Mamba network for long-term time series forecasting
Xiangsen Liu, Jinchang Ren, Hongming Zhang 0002, Erlei Zhang |
Pattern Recognit. | 2 |
| 2026 | Continual face forgery detection based on relation-aware spatial-frequency interaction aggregation and contrastive learning
Yanzhi Xu, Jinchang Ren, Aiqing Fang, Muhammad Irfan 0009, Jiangbin Zheng 0001 |
Pattern Recognit. | 2 |
| 2026 | DFBSNet: Dual frequency-domain branch fusion and selection network for hyperspectral anomaly detection
Dong Zhao 0005, Mingtao You, Pei Xiang, Yuta Asano, Xin Yu 0002, Huixin Zhou, Jinchang Ren |
Pattern Recognit. | 10 |
| 2026 | TBCNet: Twin-branch collaborative network for hyperspectral anomaly detection
Dong Zhao 0005, Mingtao You, Pei Xiang, Jianling Hu, Yuta Asano, Xin Yu 0002, Chih-Chung Hsu, Huixin Zhou, Jinchang Ren |
Pattern Recognit. | 9 |
| 2026 | Physics-Guided Neural Radiance Fields for Forward-Looking Sonar ImagingabstractWhile neural radiance fields (NeRFs) have achieved remarkable success in optical imaging, their extension to forwardlooking sonar (FLS) remains underexplored due to the fundamental mechanisms of different acoustic propagation physics. In this work, we present Sonar-NeRF, a physics-guided neural rendering framework tailored for high-fidelity FLS novel view synthesis. This approach replaces volume rendering with an explicit forward rendering model, derived from the active sonar equation to directly predict sonar echo intensities. A differentiable acoustic reflection model is incorporated to effectively capture specular reflections on metallic surfaces. In addition, heteroscedastic uncertainty learning based on an additive-multiplicative noise model is introduced to enable adaptive noise modelling. Experiments on synthetic and real FLS data show that our method is quantitatively and qualitatively superior to the conventional sonar simulators and existing NeRF-based methods. Cao Huang, Jinchang Ren, Hongyu Yang 0002, Yulong Ji |
IEEE Signal Process. Lett. | 2 |
| 2026 | Reliable-Teacher: Uncertainty-Guided Collaborative Learning for Nighttime Object DetectionabstractNighttime object detection presents significant challenges due to the scarcity of large-scale, high-quality annotations across diverse nighttime scenarios. To circumvent the need for manual nighttime image annotation, researchers have explored Unsupervised Domain Adaptive Object Detection (UDA-OD), which transfers knowledge from labeled daytime datasets to unlabeled nighttime data through pseudo-labeling. While existing approaches have shown promising results, their effectiveness remains limited by the low quality of pseudo labels, restricting model adaptation to nighttime conditions. To address these limitations, we propose Reliable-Teacher, a novel mutual-learning framework that comprehensively leverages target domain knowledge through Uncertainty-Guided Collaborative Learning. Specifically, our approach consists of three key components: 1) A Collaborative Pseudo-Label Construction module that intelligently integrates reliable Teacher-generated pseudo-labels into Student proposals, significantly enhancing pseudo-label quality; 2) An Uncertainty-Guided Consistency Reasoning module that enforces inter-category consistency between Teacher and Student predictions at both anchor and bounding box levels; 3) A Reliability-Weighted Classification Loss that minimizes the influence of unreliable predictions to further enhance uncertainty-guided learning. Extensive experiments demonstrate that Reliable-Teacher significantly outperforms state-of-the-art methods, achieving performance gain of up to 3.1%, 2.2% and 1.7% mAP on BDD100K [1], SHIFT [2], and VisDrone [3] benchmarks, respectively. Upon acceptance, our code will be released to facilitate further research in this domain. Wenjing Jia, Jiaqi Xiao, Jinchang Ren, Di Yuan 0002, Qiguang Miao, Xiangjian He |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PSTAN: A JND-Aware Pairwise Spatio-Temporal Alignment Network for Compressed Videos Quality EnhancementabstractCompressed video quality enhancement (CVQE) is crucial for mitigating compression artifacts and improving perceptual visual quality, especially under diverse quantization parameters (QPs) and motion patterns. However, many existing approaches insufficiently exploit long-range temporal dependencies, and their reliance on QP-specific training often leads to limited robustness when compression conditions change. In this work, we propose a just noticeable difference (JND)-aware and perception-driven learning framework for CVQE, termed the Pairwise Spatio-Temporal Alignment Network (PSTAN). PSTAN incorporates perceptual priors primarily through a JND-guided training paradigm rather than relying solely on architectural modifications, where learning is driven by perceptuallypoorvideo segments identified in the VideoSet dataset. This strategy alleviates the reliance on QP-specific supervision and promotes more stable enhancement behavior across varying compression conditions. To effectively capture temporal dependencies, PSTAN employs a pairwise spatio-temporal interaction mechanism that models each reference-target frame pair independently, enabling adaptive utilization of both nearby and distant frames. In addition, a transformer-based alignment module combining temporal mutual attention with cascaded deformable convolution is introduced to handle complex and large motions. Extensive experiments on VideoSet, MFQE 2.0 and our constructed HEVC-comperssed dataset show that PSTAN achieves consistent improvements over state-of-the-art CVQE methods in both objective and perceptual quality metrics. The code of this work is available at https://github.com/leryong/PSTAN.git. Yuan Yuan 0007, Eryong Li, Jiawei Zhang 0002, Jinchang Ren, Xu Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Hyperspectral Imaging and Machine Learning for Non-Destructive Phenolic Compounds Measurement in Peat Toward Smart Whisky ManufacturingabstractThe whisky industry heavily relies on peat as a key ingredient to impart distinctive smoky flavors to the final product. However, traditional methods for analyzing peat composition and quality are time-consuming and destructive, requiring extensive sample preparation. To address these challenges, we propose a novel nondestructive system for rapid and accurate peat analysis combining push-broom hyperspectral imaging (HSI), singular spectrum analysis (SSA), and machine learning. We introduce a faster SSA variant (SSA++) to overcome the high computational complexity of traditional SSA, enabling real-time processing of HSI data when captured in the push-broom manner. Comprehensive experiments have demonstrated the effectiveness of the proposed system, achieving a total phenol estimation of up to 99.31%R2. SSA++ maintains similar accuracy to SSA while significantly reducing computational time, enabling real-time performance. Our system offers a powerful tool for automated peat analysis, facilitating smart manufacturing and enhanced quality monitoring in the whisky industry. Yijun Yan, Jinchang Ren, Barry Harrison, Oliver Lewis, Emanuele Trucco, Guofang Wang, Yutang Ma |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | AutoFPDesigner: Automated Flight Procedure Design Based on Multi-Agent Large Language ModelabstractFlight procedures are essential to the safety and efficiency of air traffic management. However, due to the highly specialized nature of the flight procedure design process, existing methods rely heavily on manual operations and adjustments with limited automation, resulting in inefficiencies and potential safety risks. This study introduces AutoFPDesigner, a new agent-driven approach to flight procedure design, leveraging large language models (LLMs). By utilizing multi-agent collaboration, AutoFPDesigner automates Performance-Based Navigation (PBN) procedures, enabling end-to-end automation. In this framework, the designer’s role shifts from an executor to a supervisor, issuing tasks through natural language, while the system integrates specialized knowledge and uses a toolset to complete the design. Experimental results show that procedures designed with this approach meet safety requirements nearly 100%, with 75% of tasks completed in a limited number of steps. Moreover, AutoFPDesigner performs effectively across various design tasks, outperforming existing methods. Additionally, this study conducted human interaction experiments and introduced an “instruction-based” feedback method to address agent misinterpretation of human feedback. Experimental results demonstrate that the system bridges the skill gap between experts and beginners, and that the “instruction-based” feedback method enhances the accuracy of agent feedback interpretation. Code and data are available onhttps://github.com/Zhulongtao6/AutoFPDesigner-LLM Longtao Zhu, Hongyu Yang 0002, Yulong Ji, Jinchang Ren |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Cas-OVD: Cascaded Open-Vocabulary Detection of Small Objects Using Multi-Refined Region Proposal Network in Autonomous DrivingabstractAlthough text information has aided existing models to achieve promising results in open vocabulary object detection (OVD), the lack of semantic information has led to the difficulty in small objects detection (SOD). Moreover, such semantic gap also causes failure when matching texts and image features, resulting in false negative instances being detected. To address these issues, we propose a Cascade Open Vocabulary Detector (Cas-OVD), which builds upon existing multi-stage detection pipelines but specializes in text-vision alignment for small objects. In particular, we adapt a multi-refined region proposal network, guided by a non-sampled anchor strategy, to reduce the missing and false detections of small objects. Meanwhile, a deformable convolution network based feature conversion module is proposed to enhance the semantic information of small objects even the potential ones with low confidence. Unlike existing methods that rely on coarse-grained image-based features for image-text matching, Cas-OVD refines these features through a cascade alignment process, allowing each stage to build on the results of the previous one. This can progressively enhance the feature correlation between the image regions and the textual descriptions through successive error correction. On the joint BDD100K-SODA-D dataset, Cas-OVD achieved 17.95% AP$_{\mathrm{all}}$and 14.6% AP$_{\mathrm{s}}$, outperforming RegionCLIP by 3.5% AP$_{\mathrm{all}}$and 3.0% AP$_{\mathrm{s}}$, respectively. On the OV_COCO dataset, Cas-OVD has the 32.71% AP$_{\mathrm{all}}$and 17.26% AP$_{\mathrm{s}}$, surpassing the RegionCLIP by 6.6% AP$_{\mathrm{all}}$and 6.1% AP$_{\mathrm{s}}$, respectively. Zhenyu Fang, Jinchang Ren, Jiangbin Zheng 0001, Yijun Yan, Lixiang Zhang |
IEEE Trans. Multim. | 3 |
| 2025 | MDDNet: Multilevel Difference-Enhanced Denoise Network for Unsupervised Change Detection in SAR ImagesabstractChange detection in synthetic aperture radar (SAR) images is a hot yet highly challenging task in remote sensing. Existing unsupervised SAR change detection methods often struggle with inherent speckle noise and insufficiently utilize pseudo-labels, particularly neglecting uncertain areas. In this paper, we propose a multilevel difference-enhanced denoise dual-branch network (MDDNet), comprising representation learning and change detection branches. First, fuzzy c-means clustering is employed to generate pseudo-labels, categorizing the image areas as changed, nochanged, and uncertain. Second, we design a denoise representation loss function in the representation learning branch to maximize the use of pseudo-labels, while mitigating speckle noise. Furthermore, a multilevel difference computation module is proposed to focus on changes in ground objects and capture more comprehensive change information. Experimental results on three public SAR datasets show that the proposed method outperforms six state-of-the-art methods, achieving the best performance with an average overall accuracy of 98.86% and an average Kappa coefficient of 89.36%. He Zong, Erlei Zhang, Xinyu Li 0013, Hongming Zhang 0002, Jinchang Ren |
ICASSP | 5 |
| 2025 | Entropy guidance hierarchical rich-scale feature network for remote sensing image semantic segmentation of high resolution
Haoxue Zhang, Linjuan Li, Xinlin Xie, Jinchang Ren, Gang Xie 0001 |
Appl. Intell. | 5 |
| 2025 | GaitBranch: A multi-branch refinement model combined with frame-channel attention mechanism for gait recognition
Huakang Li, Yidan Qiu, Huimin Zhao 0001, Jin Zhan, Rongjun Chen 0001, Jinchang Ren, Ying Gao 0004, Wing W. Y. Ng |
Comput. Vis. Image Underst. | 6 |
| 2025 | Blind sonar image quality assessment via machine learning: Leveraging micro- and macro-scale texture and contour features in the wavelet domainabstractIn subsea environments, sound navigation and ranging (SONAR) images are widely used for exploring and monitoring infrastructures due to their robustness and insensitivity to low-light conditions. However, their quality can degrade during acquisition and transmission, where standard SONAR image processing techniques can hardly produce high-quality outcomes. An effective image quality assessment (IQA) method can assess their usefulness and aid to develop refinement techniques by identifying the degradation issues, ensuring the reliability of SONAR data. Existing methods often fail to account for degradations from noise, distortion, and resolution changes simultaneously. To address this challenge, we propose a new blind quality assessment method that measures the overall quality of SONAR images by quantifying both the perceptual and utility qualities using the micro- and macro-scale texture and contour features derived from the wavelet domain. By combining the local binary pattern (LBP) micro-scale texture features with the proposed histograms of Schmid Gabor-like edge maps as macro-scale features, a support vector regression model is learned to map from these features to subjective quality scores. Extensive experiments have demonstrated the superiority of our method over existing SONAR IQA techniques on distorted and reconstructed super-resolution side-scan, acoustic lens, and forward-looking SONAR images. Specifically, our method achieves Pearson’s and Spearman’s correlation metrics of 0.8616 and 0.8541, respectively, for distorted SONAR images, demonstrating improvements of 4.69% and 4.8%. For reconstructed super-resolution SONAR images, our method attains correlation metrics of 0.9415 and 0.9408, reflecting improvements of 0.8% and 1.6% over the second-best method, respectively. To facilitate ease of access, a comprehensive list of key abbreviations and their full names is provided in Table A.9 in the Appendix section. The source code of the proposed method will be shared at https://github.com/hfarhaditolie/BSIQA . Hamidreza Farhadi Tolie, Jinchang Ren, Rongjun Chen 0001, Huimin Zhao 0001, Eyad Elyan |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Large-scale cross-modal hashing via Kolmogorov-Arnold representation theorem and optimal transport
Rongjun Chen 0001, Chengsi Yao, Xianxian Zeng, Yongzhi Ma, Jun Yuan 0004, Jia Wen Li 0001, Huimin Zhao 0001, Xu Lu 0002, Jinchang Ren |
Knowl. Based Syst. | 9 |
| 2025 | Aligning local features from multi-view (ALFM): A hybrid self-Supervised framework for object detection via contextual distillation and global representation learning
Zhenyu Fang, Zhuowei Wang 0006, Jinchang Ren, Jiangbin Zheng 0001, Rongjun Chen 0001, Huimin Zhao 0001 |
Knowl. Based Syst. | 3 |
| 2025 | AFCMS-Net: Adaptive feature coupling and multi-level supervision network for effective image forgery localization
Yanzhi Xu, Jinchang Ren, Aiqing Fang, Muhammad Irfan 0009, Jiangbin Zheng 0001 |
Knowl. Based Syst. | 2 |
| 2025 | An Attention Architecture With Twice Attention Convolution and Simplified Transformer for Hyperspectral Image ClassificationabstractConvolutional neural network (CNN) and Transformer-based hybrid models have been successfully applied to hyperspectral image (HSI) classification, enhancing the local feature extraction capability of single Transformer-based models. However, these Transformers in the hybrid models suffer from structural redundancy in components such as positional encoding (PE) and multi-layer perceptron (MLP). To address the issue, we propose a novel attention architecture termed twice attention convolution module and simplified Transformer (TAST) for HSI classification. The proposed TAST primarily consists of a twice attention convolution module (TACM) and a simplified Transformer. TACM is designed to improve the ability to extract local features. In addition, we introduce the simplified Transformer by removing the PE and MLP components from the original Transformer, which captures long-range dependencies while simplifying the structure of the original Transformer. Experimental results on four public datasets demonstrate that the proposed TAST model outperforms both state-of-the-art CNN and Transformer models in terms of classification performance, with improvements in terms of overall accuracy (OA) around 3.87%-34.95% (Indian Pines), 0.35%-23.43% (Salinas), 0.37%-6.05% (WHU-Hi-LongKou), and 0.65%-10.79% (WHU-Hi-HongHu). Xuejiao Liao, Fangyuan Lei, Alex Hayman Ng, Jinchang Ren |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Prototype-Guided Spatial-Spectral Interaction Network for Hyperspectral Anomaly DetectionabstractIn recent years, deep learning has emerged as one of the most widely utilized techniques in hyperspectral anomaly detection (HAD) with an impressive detection accuracy. However, the investigation into the diverse background representation and the spatial-spectral interaction remains underexplored. To tackle with this, we propose a novel framework namely the prototype-guided and spatial-spectral interaction network (PSSIN) for HAD in this paper. Specifically, an adaptive anomaly mask module is utilized to mitigate the interference of the background reconstruction caused by the blending of potential anomalies. Subsequently, we design a background-guided prototype autoencoder (BP-AE) to represent the backgrounds with various land cover types, incorporating two critical components: the background prototype module (BPM) and the spatial spectral interaction block (SSIB). To characterize different typical background features by a global perspective, BPM utilizes a prototype learning strategy with a self-attention mechanism, and a multivariate ensemble loss is employed for BPM to optimize the transformation of background features and the updating of a prototype codebook. To enhance the spatial-spectral utilization of window-based approach, SSIB first introduce a spatial-spectral interaction paradigm for HAD. The window-based self-attention branch is to mine spatial features characteristics, while the depth-wise convolution branch is to extract spectral features. These two branches in a parallel configuration interact with each other's features and then perform feature fusion. SSIB architecture not only broadens the receptive fields by concurrently modeling the intra-window and cross-window relationships but also facilitates bi-directional interactions between the spatial and spectral branches. Furthermore, the comprehensive experiments conducted on six authentic datasets have fully validated its superior performance. Yu Huo 0001, Min Zhang 0015, Hai Wang 0015, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Dual Teacher: Improving the Reliability of Pseudo Labels for Semi-Supervised Oriented Object DetectionabstractOriented object detection in remote sensing is a critical task for accurately location and measurement of the interested targets. Despite of its success in object detection, deep learning-based detectors rely heavily on extensive data annotation. However, variations in object appearance significantly increase the difficulty and the cost of creating large-scale annotated datasets. Semi-supervised learning (SSL) aims to utilize unlabeled data to enhance object detectors. Among these, pseudo-label-based methods have shown promising results recently. Nonetheless, as training progresses, the accumulation of errors in pseudo labels leads to prediction bias without corrections. To tackle this particular challenge, we present a SSL pipeline, named “dual teacher,” for improving the reliability of pseudo labels in the semi-supervised oriented object detection. First, to mitigate the bias caused by limited annotated data, a global burn-in (GBI) strategy is introduced at the beginning of training, which guides the student detector to learn the feature extraction on a global scale. In addition, an online bounding box (bbox) correction module is proposed to decrease the occurrence of mislabeled instances and enhance the reliability of detection. These improvements are facilitated by an additional detector, instead of a single teacher model in the teacher-student architecture. Dual teacher reduces the dependency on the quality of pseudo labels related to the model complexity and combines the strengths of both the two-stage and one-stage detectors. With only 20% labeled data, dual teacher outperforms fully supervised rotated fully convolutional one-stage object detection (R-FCOS), you only look once X-small (YOLOX-s), and rotated region-based convolutional neural network (R-RCNN) by up to 2% on both a large-scale dataset for object detection in aerial images (DOTA) and SODA-A datasets. This reveals its potential in reducing labor-intensive tasks and enhancing robustness against environmental interference and noisy labels. The code is available at:https://github.com/ZYFFF-CV/DualTeacher-semisup.git. Zhenyu Fang, Jinchang Ren, Jiangbin Zheng 0001, Rongjun Chen 0001, Huimin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | GLVMamba: A Global-Local Visual State-Space Model for Remote Sensing Image SegmentationabstractSemantic segmentation of remote sensing images has significant advances with the adoption of deep neural networks, taking the advantages of Convolutional Neural Networks (CNNs) in local feature extraction with Transformers in global information modeling. However, due to the limitations of CNNs in long-range modeling capabilities and the computational complexity constraints of Transformers, remote sensing semantic segmentation still faces issues such as serious holes, rough edge segmentation, false and even missed detections caused by the light, shadow and other factors. To address these issues, we propose a visual state space model called GLVMamba, which employs CNNs as the encoder and the proposed Global-Local Visual State Space (GLVSS) block as the core decoder. Specifically, the GLVSS block introduces locality forward feedback and shift window mechanism to addresses the deficiency of insufficient modeling of neighboring pixel dependencies of Mamba, which enhances the integration of global and local context during feature reconstruction, boosts object perception capabilities of the model, and effectively refines edge contours. Additionally, the scale-aware pyramid pooling (SCPP) module is proposed to fully merge the features from various scales and adaptively fuse and extract the distinguishing features to mitigate the holes and false detections. The GLVMamba effectively captures global-local semantic information and multi-scale feature through the GLVSS block and the SCPP module, achieving efficient and accurate remote sensing semantic segmentation. Extensive experiments on two widely used datasets have effectively demonstrated the superiority of our proposed method over the other state-of-the-art methods. The code will be available at https://github.com/Tokisakiwlp/GLVMamba. Huajian Pan, Xiaoyong Liu 0001, Jinchang Ren, Jingjing Cao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MSLKCNN: A Simple and Powerful Multiscale Large Kernel CNN for Hyperspectral Image ClassificationabstractDeep learning-based hyperspectral image (HSI) classification models typically utilize multiple feature extraction layers to learn the features of land covers. Nevertheless, they encounter challenges, e.g., 1) Transformers require substantial computational resources, and 2) these layers are carefully assembled and designed. Recently, large kernel convolutional neural networks (LKCNNs) show excellent performance in natural visual tasks. To tackle these limitations and explore the capability of LKCNNs for HSI classification, we present a novel simple and powerful multi-scale large kernel convolutional neural network architecture (MSLKCNN) with the largest kernel size as large as 15 × 15, in contrast to commonly used 3 × 3, for HSI classification. MSLKCNN avoids these specialized designs, comprising a noise suppression module (NSM) and a multi-scale large kernel convolution (MSLKC). Specifically, NSM is first used to suppress the noise and reduce the number of the bands before extracting the features. Then, MSLKC, as the only feature extraction layer of MSLKCNN, joints three parallel convolutions to capture the features of various types (i.e. spectral, spectral-spatial) and ranges (i.e., small local, larger local, and global) from the dimension of scale: (C1) convolution with a kernel size of 1 × 1 is used to extract spectral features; (C2) multi-scale large kernel depthwise separable convolution (MLKDC) is proposed to learn the spectral-spatial features of different ranges including short-range, middle-range, and long-range; and (C3) multi-scale dilated depthwise separable convolution (MDDC) is designed to aggregate the spectral-spatial features between land covers at various distances. Extensive experimental results on three public HSI datasets demonstrate the competitiveness of the proposed MSLKCNN compared with several state-of-the-art methods. Alex Hayman Ng, Fangyuan Lei, Jinchang Ren, Zheyuan Du |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Unsupervised Domain Adaptation for VHR Urban Scene Segmentation via Prompted Foundation Model-Based Hybrid Training Joint-Optimized NetworkabstractUnsupervised Domain Adaptation for Remote Sensing Semantic Segmentation (UDA-RSSeg) is to adapt a model trained on the source domain data to the target domain samples, thereby minimizing the need for annotated data across diverse remote sensing scenes. In urban planning and monitoring, the task of UDA-RSSeg on Very-High-Resolution (VHR) images has garnered significant research interest. While recent deep learning techniques have demonstrated huge success in tackling the UDA-RSSeg task for VHR urban scenes, a persistent challenge in addressing the domain shift issue remains. Specifically, there are two primary problems: (1) severe inconsistencies in feature representation across diverse domains, characterized by notably differing data distributions, and (2) the domain gap problem due to the representation bias of the source domain patterns when translating features to predictive logits. To solve these problems, we propose a prompted foundation model based hybrid training joint-optimized network (PFM-JONet) for UDA-RSSeg on VHR urban scene. Our approach integrates the notable “Segment Anything Model” (SAM) as prompted foundation model to leverage its robust generalized representation capabilities, thereby alleviating feature inconsistencies. Based on the feature extracted by SAM-Encoder, we introduce a mapping decoder designed to convert SAM-Encoder features into predictive logits. Additionally, a prompted segmentor is employed to generate class-agnostic maps, which guide the mapping decoder’s feature representations. To efficiently optimize the entire network in an end-to-end manner, we design a hybrid training scheme that integrates feature-level and logits-level adversarial training strategies alongside a self-training mechanism. This scheme enhances the model from diverse, compatible perspectives. To evaluate the performance of our proposed PFM-JONet, we conduct extensive experiments on urban scene benchmark datasets, including ISPRS (Potsdam/Vaihingen) and CITY-OSM (Paris/Chicago). On ISPRS dataset, PFM-JONet surpasses previous SOTA methods by 1.60% in mean IoU value across four adaptation tasks. For CITY-OSM’s adaptation task, it outperforms SOTA by 4.84% in mean IoU value. These results demonstrate the effectiveness of our method. Furthermore, visualization and analysis reinforce the method’s interpretability. The code of this paper is available at https://github.com/CV-ShuchangLyu/PFM-JONet. Shuchang Lyu, Qi Zhao 0037, Yaxuan Sun, Yiwei He, Guangbiao Wang, Jinchang Ren, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | ChangeDA: Depth-Augmented Multitask Network for Remote Sensing Change Detection via Differential AnalysisabstractIn the field of remote sensing change detection (RSCD), accurately identifying significant changes between bi-temporal images is essential for environmental monitoring, urban planning, and disaster assessment. In recent years, advancements in deep learning for computer vision (CV) have transformed RSCD, significantly enhancing its effectiveness. However, existing methods often overlook the importance of depth information, focusing primarily on 2-D information. This limits their ability to capture subtle changes and structural details in 3-D space. To address these limitations, we introduce ChangeDA—a depth-augmented multitask network designed to enhance the effectiveness of RSCD. ChangeDA introduces a depth encoder module to extract implicit depth information from optical images, enabling the utilization of 3-D structural information without reliance on external data sources. Through the depth infusion module (DIM), depth information is integrated into the dual-temporal feature maps, significantly enhancing the network’s ability to perceive changes in 3-D spatial structures. In addition, ChangeDA includes a differential feature extractor (DFE) tailored to pinpoint differential features between sequential images, and an adaptive all-feature fusion (AAFF) strategy that significantly improves recognition accuracy and generalization capability through cross-level feature integration. Performance evaluations on four prominent single-modal datasets—LEVIR-CD, S2Looking, WHU-CD, and SYSU-CD—yielded state-of-the-art (SOTA)${F}1$-scores of 92.27%, 66.42%, 94.12%, and 82.74%, respectively. Furthermore, ChangeDA also achieved outstanding results on the multimodal 3DCD dataset, with an${F}1$score of 63.52% in 2-D CD and an RMSE of 1.20 in the 3-D CD task. These results demonstrate ChangeDA’s robust adaptability across diverse targets and real-world scenarios. Jiangtao Meng, Xinying Xu, Pengyue Li, Gang Xie 0001, Jinchang Ren, Yuxuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Learn From Past to Future: Exploiting Self-Training and Curriculum Learning in Remote Sensing Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation focuses on updating the segmentation model with only new-class samples. Catastrophic forgetting and background shift are the two prevalent challenges. We identify two additional issues in remote sensing data that worsen these problems: significant class distribution variability and error accumulation-induced model degradation. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). First, we introduce a self-training auxiliary branch to complement the frozen last-step model, integrating cross-step knowledge to mitigate rapid forgetting. Then, a gradient-oriented Dynamic Refined Loss is proposed to assess under-learned classes and mitigate class imbalance. Furthermore, class-balanced curriculum learning is embedded to alleviate performance degradation throughout incremental training. Extensive experiments on benchmark datasets, including DeepGlobe, iSAID, ISPRS Potsdam, and Vaihingen, demonstrate that the proposed STCL-DRNet achieves state-of-the-art (SOTA) performance. In the 1-1s setting of the DeepGlobe dataset, STCL-DRNet exceeds previous SOTA methods by 11.6% in mIoU. For the iSAID 10-1s setting, it outperforms the previous SOTA by 12.76% in mIoU. As for ISPRS Potsdam and Vaihingen, our STCL-DRNet surpasses the SOTA by 5%-8% in all settings. Visualization and analysis further validate its interpretability. Our code is available at https://github.com/cv516Buaa/STCL-DRNet. Ruimin Ren, Hongbo Zhao 0001, Shuchang Lyu, Guangbiao Wang, Qi Zhao 0037, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | FusDreamer: Label-Efficient Remote Sensing World Model for Multimodal Data ClassificationabstractWorld models significantly enhance hierarchical understanding, improving data integration and learning efficiency. To explore the potential of the world model in the remote sensing (RS) field, this article proposes a label-efficient RS world model for multimodal data fusion (FusDreamer). The FusDreamer uses the world model as a unified representation container to abstract common and high-level knowledge, promoting interactions across different types of data, that is, hyperspectral (HSI), light detection and ranging (LiDAR), and text data. Initially, a new latent-spatial multimodal generation (LaMG) paradigm is utilized for its exceptional information integration and detail retention capabilities. Subsequently, an open-world knowledge-guided consistency projection (OK-CP) module incorporates prompt representations for visually described objects and aligns language-visual features through contrastive learning. In this way, the domain gap can be bridged by fine-tuning the pre-trained world models with limited samples. Finally, an end-to-end multitask combinatorial optimization (MuCO) strategy can capture slight feature bias and constrain the diffusion process in a collaboratively learnable direction. Experiments conducted on four typical datasets indicate the effectiveness and advantages of the proposed FusDreamer. The corresponding code will be released athttps://github.com/Cimy-wang/FusDreamer. Hao Chen 0117, Jinchang Ren, Huimin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Binary Quantization Vision Transformer for Effective Segmentation of Red Tide in Multispectral Remote Sensing ImageryabstractAs a global marine disaster, red tides pose serious threats to marine ecology and the blue economy, making their monitoring crucial for preventing harmful algal blooms (HABs) and protecting the marine environment. In this study, satellite remote sensing was utilized to provide timely, large-scale, and continuous observation capabilities, overcoming the high cost and spatial and temporal limitations of in situ monitoring. However, existing remote sensing-based methods often exhibit coarse segmentation granularity and suffer from high computational complexity. To overcome these challenges, we propose a novel bimodal multispectral dynamic offset binary quantization visual transformer (DoBi-SWiP-ViT) that utilizes the ViT for global feature aggregation and parameter quantization for efficient segmentation. With the bimodal Swin-ViT with unified perceptual parsing (UPP) architecture, our model integrates data from multiple spectral bands to achieve fine-grained segmentation of large-scale remote sensing images. Additionally, we introduce a dynamic magnitude offset binary quantization ViT block to reduce the parameter redundancy and improve the computational efficiency. In addition, we validated the performance of our model through extensive comparative experiments on high-resolution imagery datasets of sea surface red tides collected from different satellite platforms. The results show that our proposed DoBi-SWiP-ViT has significantly improved the mean accuracy (mAcc) of the segmentation results. For the two test areas acquired from different satellite platforms, the improvements are 8.78% and 10.18%, respectively. This has demonstrated the superior performance of our model in detecting the red tides from high-resolution visible images, highlighting its effectiveness in capturing complex patterns and subtle features in multispectral imagery. Yefan Xie, Jinchang Ren, Xinchao Zhang, Chengcheng Ma, Jiangbin Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Frequency-Domain Guided Swin Transformer and Global-Local Feature Integration for Remote Sensing Images Semantic SegmentationabstractConvolutional neural networks (CNNs), transformers, and the hybrid methods have been significant application in remote sensing. However, existing methods are limited in effectively modeling frequency-domain information, which affects their ability to capture detailed information. Therefore, we propose a frequency-domain guided feature coupled mechanism and a global-local feature integration method (FGNet) for semantic segmentation. Specifically, a frequency-domain guided Swin (FGSwin) transformer is designed by introducing dilation group convolution, fast Fourier transform (FFT), and learnable weights to enhance the expression capability of frequency-domain and space-domain, local and global features, simultaneously. In addition, a global-local feature integration (GLFI) module is proposed for aggregating features to further enhance the discrimination of each category. Comprehensive experimental results demonstrate that compared with existing methods, the proposed method achieves superior performance in terms of mean intersection over union (mIoU), reaching 71.46% and 74.04% on ISPRS Potsdam and Vaihingen, two widely used datasets. Haoxue Zhang, Gang Xie 0001, Linjuan Li, Xinlin Xie, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | ICSF: Integrating Inter-Modal and Cross-Modal Learning Framework for Self-Supervised Heterogeneous Change DetectionabstractHeterogeneous change detection (HCD) is a process to determine the change information by analyzing heterogeneous images of the same geographic location taken at different times, which plays an important role in remote sensing applications such as disaster response and environmental monitoring. However, the different imaging mechanisms result in different visual appearances in heterogeneous images, making it difficult to accurately detect changes through direct comparison. To address this problem, we propose a inter-modal and cross-modal self-supervised dual branch learning framework (ICSF) for HCD that incorporates inter-modal and cross-modal learning. First, in the inter-modal branch, we perform contrastive learning on heterogeneous images within their respective modalities to learn the robust and discriminative features, rather than relying on the raw spectral or spatial information from these images. Second, in the cross-modal branch, we perform cross-modal reconstruction to ensure the obtained features exhibit consistent comparability, thereby facilitating the extraction of rich information on the real changes within the images. Next, the difference images (DIs) computed from both branches are further refined using a superpixel segmentation strategy to preserve the consistency of differences within the same ground object. Experimental results on five public datasets with different modality combinations and change events demonstrate the effectiveness of the proposed approach in comparison to ten state-of-the-art (SOTA) methods, achieving the best performance with an average overall accuracy (OA) of 95.88% and an average Kappa coefficient (KC) of 74.20%. Erlei Zhang, He Zong, Xinyu Li 0013, Mingchen Feng, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Low-Rank and Sparse Representation Meet Deep Unfolding: A New Interpretable Network for Hyperspectral Change DetectionabstractHyperspectral image change detection (HSI-CD) is a technique that intelligently checks the changed details in bitemporal hyperspectral images (Bi-HSIs). Deep learning (DL), with the ability to model nonlinear changing features, has achieved promising results in HSI-CD, but the feature mining mechanism is unclear and the architecture design lacks transparency in such DL models. To alleviate this problem, this paper proposes a new low-rank and sparse representation-based deep unfolding network (LRSRNet) for HSI-CD. For feature mining mechanism, the LRSRNet adopts a low-rank and sparse subnetwork (LRSnet) and a change detection sub-network (CDnet). The former is responsible for extracting low-rank features with valuable information and suppressing sparse features containing interference information, while the latter aims to obtain change information from low-rank features. For architecture design, the LRSnet formulates the HSI as a low-rank estimation, sparse estimation, and hyperspectral reconstruction in a low-rank and sparse model, and iteratively optimizes and updates the above sub-problems through deep networks. A new CDnet is designed as a concise convolutional architecture to extract change information from representative Bi-HSIs features. Experiments on three real datasets demonstrate the performance superiority of the proposed LRSRNet method over nine model-driven, datadriven, and model-data-joint-driven HSI-CD algorithms in both qualitative and quantitative evaluations. The proposed LRSRNet is available online: https://github.com/chengle-zhou/LRSRNet. Chengle Zhou, Zhi He, Jian Dong 0004, Yunfei Li 0006, Jinchang Ren, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Eye state detection based on feature fusion and attention single-shot multi-box detector
Weikun Dai, Jianbin Xiong, Jianxiang Yang, Wulue Zhang, Qianguang Zhang, Jinchang Ren |
Vis. Comput. | 9 |
| 2024 | Sparse Autoencoder Based Hyperspectral Anomaly Detection with the Singular Spectrum Analysis Based Spectral DenoisingabstractAs an effective tool for monitoring surface irregularities in remote sensing, hyperspectral anomaly detection (HAD) has garnered increasing attention. However, how to improve the detection accuracy remains a formidable challenge, due mainly to the noise and variations in the spectral domain, especially when there is lack of the labelled data for training. To tackle these difficulties, a novel unsupervised HAD method is proposed. First, 1-D Singular Spectrum Analysis (SSA) is employed to eliminate outliers in the spectral domain. Second, the SSA-smoothed hypercube undergoes a sparse autoencoder for background reconstruction, where the reconstruction error is used to extract anomalous pixels. Finally, the RX algorithm is employed to segment anomalous pixels from the background. Comprehensive experiments on four publicly available datasets have validated the superior performance of our method in effectively enhancing the separability between anomaly pixels and their respective backgrounds, outperforming a few state-of-the-art methods, particularly in terms of the detection accuracy. Yinhe Li, Jinchang Ren, Zhi Gao 0005, Genyun Sun |
IGARSS | 2 |
| 2024 | DICAM: Deep Inception and Channel-wise Attention Modules for underwater image enhancement
Hamidreza Farhadi Tolie, Jinchang Ren, Eyad Elyan |
Neurocomputing | 2 |
| 2024 | Prompting-to-Distill Semantic Knowledge for Few-Shot LearningabstractRecognizing visual patterns in low-data regime necessitates deep neural networks to glean generalized representations from limited training samples. In this letter, we propose a novel few-shot classification method, namely ProDFSL, leveraging multimodal knowledge and attention mechanism. We are inspired by recent advances of large language models and the great potential they have shown across a wide range of downstream tasks and tailor it to benefit the remote sensing community. We utilize ChatGPT to produce class-specific textual inputs for enabling CLIP with rich semantic information. To promote the adaptation of CLIP in remote sensing domain, we introduce a cross-modal knowledge generation module, which dynamically generates a group of soft prompts conditioned on the few-shot visual samples and further uses a shallow Transformer to model the dependencies between language sequences. Fusing the semantic information with few-shot visual samples, we build representative class prototypes, which are conducive to both inductive and transductive inference. In extensive experiments on standard benchmarks, our ProDFSL consistently outperforms the state of the art in few-shot learning (FSL). Zhi Gao 0005, Jinchang Ren, Xing-ao Wang, Ping Ma 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | PWDformer: Deformable transformer for long-term series forecasting
Zheng Wang 0008, Haowei Ran, Jinchang Ren, Meijun Sun |
Pattern Recognit. | 3 |
| 2024 | Feature Aggregation and Region-Aware Learning for Detection of Splicing ForgeryabstractDetection of image splicing forgery become an increasingly difficult task due to the scale variations of the forged areas and the covered traces of manipulation from post-processing techniques. Most existing methods fail to jointly multi-scale local and global information and ignore the correlations between the tampered and real regions in inter-image, which affects the detection performance of multi-scale tampered regions. To tackle these challenges, in this paper, we propose a novel method based on feature aggregation and region-aware learning to detect the manipulated areas with varying scales. In specific, we first integrate multi-level adjacency features using a feature selection mechanism to improve feature representation. Second, a cross-domain correlation aggregation module is devised to perform correlation enhancement of local features from CNN and global representations from Transformer, allowing for a complementary fusion of dual-domain information. Third, a region-aware learning mechanism is designed to improve feature discrimination by comparing the similarities and differences of the features between different regions. Extensive evaluations on benchmark datasets indicate the effectiveness in detecting multi-scale spliced tampered regions. Yanzhi Xu, Jiangbin Zheng 0001, Jinchang Ren, Aiqing Fang |
IEEE Signal Process. Lett. | 3 |
| 2024 | High-Resolution Remote Sensing Image Change Detection Based on Fourier Feature Interaction and Multiscale PerceptionabstractAs a significant means of Earth observation, change detection in high-resolution remote sensing images has received extensive attention. Nevertheless, the variability in imaging conditions introduces style discrepancies and a range of pseudochange regions between bitemporal image pairs. Furthermore, changing objects possess diverse morphological representations, which makes accurately identifying change areas and delineating their boundaries within complex object distributions increasingly difficult. In response to the aforementioned challenges, we propose the Fourier feature interaction and multiscale perception (FIMP) model for effective change detection. To mitigate the impact of style discrepancies, FIMP employs the Fourier transform to adaptively filter bitemporal features in the frequency domain while mining the optimized bitemporal features relevant to the change detection task. To enhance the ability to recognize multiscale changing objects, FIMP aggregates and emphasizes the change areas with the introduced temporal change enhancement module (TCEM). By utilizing the U-fusion change perception module (UCPM) to perform multilevel bidirectional fusion of change features at different scales, FIMP can further enhance the ability to delineate complex semantic change boundaries. Experiments on three public datasets show that our approach outperforms seven state-of-the-art methods. Shou Feng, Chunhui Zhao 0003, Nan Su 0001, Wei Li 0032, Ran Tao 0003, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Two-Click-Based Fast Small Object Annotation in Remote Sensing ImagesabstractIn the remote sensing field, detecting small objects is a pivotal task, yet achieving high performance in deep learning-based detectors heavily relies on extensive data annotation. The challenge intensifies as small objects in remote sensing imagery are typically densely distributed and numerous, leading to a substantial increase in the cost of creating large-scale annotated datasets. This elevated cost poses significant limitations on the application and advancement of small object detection. To address this issue, a point-based annotation (PBA) method is proposed, which generates bounding boxes (BBOXs) through graph-based segmentation. In this framework, user annotations categorize nodes into three distinct classes—positive, negative, and to-cut—facilitating a more intuitive and efficient annotation process. Utilizing the max-flow algorithm, our method seamlessly generates oriented BBOXs (OBBOXs) from these classified nodes. The efficacy of PBA is underscored by our empirical findings. Notably, annotation efficiency is enhanced by at least 40%, a significant leap forward. Moreover, the intersection over union (IoU) metric of our OBBOX outperforms existing methods like “segment anything model (SAM)” by 10%. Finally, when applied in training, models annotated with PBA exhibit a 3% increase in the mean average precision (mAP) compared with those using traditional annotation methods. These results not only affirm the technical superiority of PBA but also its practical impact on advancing small object detection in remote sensing. Lu Lei, Zhenyu Fang, Jinchang Ren, Paolo Gamba, Jiangbin Zheng 0001, Huimin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Detection-Driven Exposure-Correction Network for Nighttime Drone-View Object DetectionabstractDrone-view object detection (DroneDet) models typically suffer a significant performance drop when applied to nighttime scenes. Existing solutions attempt to employ an exposure-adjustment module to reveal objects hidden in dark regions before detection. However, most exposure-adjustment models are only optimized for human perception, where the exposure-adjusted images may not necessarily enhance recognition. To tackle this issue, we propose a novel Detection-driven Exposure-correction network for nighttime DroneDet, called DEDet. The DEDet conducts adaptive, nonlinear adjustment of pixel values in a spatially fine-grained manner to generate DroneDet-friendly images. Specifically, we develop a fine-grained parameter predictor (FPP) to estimate pixelwise parameter maps of the image filters. These filters, along with the estimated parameters, are used to adjust pixel values of the low-light image based on nonuniform illuminations in drone-captured images. In order to learn the nonlinear transformation from the original nighttime images to their DroneDet-friendly counterparts, we propose a progressive filtering module that applies recursive filters to iteratively refine the exposed image. Furthermore, to evaluate the performance of the proposed DEDet, we have built a dataset NightDrone to address the scarcity of the datasets specifically tailored for this purpose. Extensive experiments conducted on four nighttime datasets show that DEDet achieves a superior accuracy compared with the state-of-the-art (SOTA) methods. Furthermore, ablation studies and visualizations demonstrate the validity and interpretability of our approach. Our NightDrone dataset can be downloaded fromhttps://github.com/yuexiemail/NightDrone-Dataset. Wenjing Jia, Qiguang Miao, Junmei Feng, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Nondestructive Quantitative Measurement for Precision Quality Control in Additive Manufacturing Using Hyperspectral Imagery and Machine LearningabstractMeasuring the purity of the metal powder is essential to maintain the quality of additive manufacturing products. Contamination is a significant concern, leading to cracks and malfunctions in the final products. Conventional assessment methods focus more on physical integrity rather than material composition and can be time-consuming. By capturing spectral data from a wide frequency range along with the spatial information, hyperspectral imaging (HSI) can detect minor differences in terms of temperature, moisture, and chemical composition to tackle this challenge. In this article, we explore the application of HSI in conjunction with machine learning for nondestructive inspection of metal powders. By employing near-infrared and visible HSI cameras, we introduce the utilization of HSI for this purpose. We delve into the technical challenges encountered and present detailed solutions through three case studies, including the establishment of a spectral dictionary, contamination detection, and band selection analysis. Our experimental results demonstrate the immense potential of HSI and its synergy with machine learning for nondestructive testing in powder metallurgy, particularly in meeting the requirements of industrial manufacturing environments. Yijun Yan, Jinchang Ren, He Sun 0009 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Rapid Detection of Multi-QR Codes Based on Multistage Stepwise Discrimination and a Compressed MobileNetabstractPoor real-time performance in multi-QR codes detection has been a bottleneck in QR code decoding-based Internet of Things (IoT) systems. To tackle this issue, we propose in this article a rapid detection approach, which consists of multistage stepwise discrimination (MSD) and a Compressed MobileNet. Inspired by the object category determination analysis, the preprocessed QR codes are extracted accurately on a small scale using the MSD. Guided by the small scale of the image and the end-to-end detection model, we obtain a lightweight Compressed MobileNet in a deep weight compression manner to realize rapid inference of multi-QR codes. The average detection precision (ADP), multiple box rate (MBR) and running time are used for quantitative evaluation of the efficacy and efficiency. Compared with a few state-of-the-art methods, our approach has higher detection performance in rapid and accurate extraction of all the QR codes. The approach is conducive to embedded implementation in edge devices along with a bit of overhead computation to further benefit a wide range of real-time IoT applications. Rongjun Chen 0001, Hongxing Huang, Yongxing Yu, Jinchang Ren, Peixian Wang, Huimin Zhao 0001, Xu Lu 0002 |
IEEE Internet Things J. | 4 |
| 2023 | PCA-Domain Fused Singular Spectral Analysis for Fast and Noise-Robust Spectral-Spatial Feature Mining in Hyperspectral ClassificationabstractThe principal component analysis (PCA) and 2-D singular spectral analysis (2DSSA) are widely used for spectral- and spatial-domain feature extraction in hyperspectral images (HSIs). However, PCA itself suffers from low efficacy if no spatial information is combined, while 2DSSA can extract the spatial information yet has a high computing complexity. As a result, we propose in this letter a PCA domain 2DSSA approach for spectral–spatial feature mining in HSI. Specifically, PCA and its variation, folded PCA (FPCA) are fused with the 2DSSA, as FPCA can extract both global and local spectral features. By applying 2DSSA only on a small number of PCA components, the overall computational cost can be significantly reduced while preserving the discrimination ability of the features. In addition, with the effective fusion of spectral and spatial features, our approach can work well on the uncorrected dataset without removing the noisy and water absorption bands, even under a small number of training samples. Experiments on two publicly available datasets have fully validated the superiority of the proposed approach, in comparison to several state-of-the-art methods and deep learning models. Yijun Yan, Jinchang Ren, Qiaoyuan Liu, Huimin Zhao 0001, Haijiang Sun, Jaime Zabalza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | A Novel Gradient-guided Post-processing Method for Adaptive Image Steganography
Guoliang Xie, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001 |
Signal Process. | 2 |
| 2023 | Tensor Singular Spectrum Analysis for 3-D Feature Extraction in Hyperspectral ImagesabstractDue to the cubic structure of a hyperspectral image (HSI), how to characterize its spectral and spatial properties in three dimensions is challenging. Conventional spectral-spatial methods usually extract spectral and spatial information separately, ignoring their intrinsic correlations. Recently, some 3D feature extraction methods are developed for the extraction of spectral and spatial features simultaneously, although they rely on local spatial-spectral regions and thus ignore the global spectral similarity and spatial consistency. Meanwhile, some of these methods contain huge model parameters which require a large number of training samples. In this paper, a novel Tensor Singular Spectral Analysis (TensorSSA) method is proposed to extract global and low-rank features of HSI. In TensorSSA, an adaptive embedding operation is first proposed to construct a trajectory tensor corresponding to the entire HSI, which takes full advantage of the spatial similarity and improves the adequate representation of the global low-rank properties of the HSI. Moreover, the obtained trajectory tensor, which contains the global and local spatial and spectral information of the HSI, is decomposed by the Tensor singular value decomposition (t-SVD) to explore its low-rank intrinsic features. Finally, the efficacy of the extracted features is evaluated using the accuracy of image classification with a support vector machine (SVM) classifier. Experimental results on three publicly available datasets have fully demonstrated the superiority of the proposed TensorSSA over a few state-of-the-art 2D/3D feature extraction and deep learning algorithms, even with a limited number of training samples. Genyun Sun, Aizhu Zhang, Baojie Shao, Jinchang Ren, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | CBANet: An End-to-End Cross-Band 2-D Attention Network for Hyperspectral Change Detection in Remote SensingabstractAs a fundamental task in remote sensing observation of the earth, change detection using hyperspectral images (HSI) features high accuracy due to the combination of the rich spectral and spatial information, especially for identifying land-cover variations in bi-temporal HSIs. Relying on the image difference, existing HSI change detection methods fail to preserve the spectral characteristics and suffer from high data dimensionality, making them extremely challenging to deal with changing areas of various sizes. To tackle these challenges, we propose a cross-band 2-D self-attention Network (CBANet) for end-to-end HSI change detection. By embedding a cross-band feature extraction module into a 2-D spatial-spectral self-attention module, CBANet is highly capable of extracting the spectral difference of matching pixels by considering the correlation between adjacent pixels. The CBANet has shown three key advantages: 1) less parameters and high efficiency; 2) high efficacy of extracting representative spectral information from bi-temporal images; and 3) high stability and accuracy for identifying both sparse sporadic changing pixels and large changing areas whilst preserving the edges. Comprehensive experiments on three publicly available datasets have fully validated the efficacy and efficiency of the proposed methodology. Yinhe Li, Jinchang Ren, Yijun Yan, Qiaoyuan Liu, Ping Ma 0002, Andrei Petrovski 0001, Haijiang Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multiscale Diff-Changed Feature Fusion Network for Hyperspectral Image Change DetectionabstractFor hyperspectral image (HSI) change detection (CD), multiscale features are usually used to construct the detection models. However, the existing studies only consider the multiscale features containing changed and unchanged components, which is difficult to represent the subtle changes between bitemporal HSIs in each scale. To address this problem, we propose a multiscale diff-changed feature fusion network (MSDFFN) for HSI CD, which improves the ability of feature representation by learning the refined change components between bitemporal HSIs under different scales. In this network, a temporal feature encoder–decoder subnetwork, which combines a reduced inception (RI) module and a cross-layer attention module to highlight the significant features, is designed to extract the temporal features of HSIs. A bidirectional diff-changed feature representation (BDFR) module is proposed to learn the fine changed features of bitemporal HSIs at various scales to enhance the discriminative performance of the subtle change. A multiscale attention fusion (MSAF) module is developed to adaptively fuse the changed features of various scales. The proposed method can not only discover the subtle change in bitemporal HSIs but also improve the discriminating power for HSI CD. Experimental results on three HSI datasets show that MSDFFN outperforms a few state-of-the-art methods. Fulin Luo, Tianyuan Zhou, Tan Guo, Xiuwen Gong, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Multiscale Superpixelwise Prophet Model for Noise-Robust Feature Extraction in Hyperspectral ImagesabstractDespite of various approaches proposed to smooth the hyperspectral images (HSIs) before feature extraction, the efficacy is still affected by the noise, even using the corrected dataset with the noisy and water absorption bands discarded. In this study, a novel spectral-spatial feature mining framework, Multiscale Superpixelwise Prophet Model (MSPM), is proposed for noise-robust feature extraction and effective classification of the HSI. The prophet model is highly noise-robust for deeply digging into the complex structured features thus enlarging interclass diversity and improving intraclass similarity. First, the superpixelwise segmentation is produced from the first three principal components of an HSI to group pixels into regions with adaptively determined sizes and shapes. A multiscale prophet model is utilized to extract the multiscale informative trend components from the average spectrum of each superpixel. Taking the multiscale trend signal as the input feature, the HSI data are classified superpixelwisely, which is further refined by a majority vote based decision fusion. Comprehensive experiments on three publicly available datasets have fully validated the efficacy and robustness of our MSPM model when benchmarked with eleven state-of-the-art algorithms, including six spectral-spatial methods and five deep learning ones. Besides, MSPM also shows superiority under limited training samples, due to the combined strategies of superpixelwise fusion and multiscale fusion. Our model has provided a useful solution for noise-robust feature extraction as it achieves superior HSI classification even from the uncorrected dataset without prefiltering the water absorption and noisy bands. Ping Ma 0002, Jinchang Ren, Genyun Sun, Huimin Zhao 0001, Xiuping Jia, Yijun Yan, Jaime Zabalza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Large Kernel Spectral and Spatial Attention Networks for Hyperspectral Image ClassificationabstractCurrently, long-range spectral and spatial dependencies have been widely demonstrated to be essential for hyperspectral image (HSI) classification. Due to the transformer superior ability to exploit long-range representations, the transformer-based methods have exhibited enormous potential. However, existing transformer-based approaches still face two crucial issues that hinder the further performance promotion of HSI classification: 1) treating HSI as 1D sequences neglects spatial properties of HSI, 2) the dependence between spectral and spatial information is not fully considered. To tackle the above problems, a large kernel spectral-spatial attention network (LKSSAN) is proposed to capture the long-range 3D properties of HSI, which is inspired by the visual attention network (VAN). Specifically, a spectral-spatial attention module is first proposed to effectively exploit discriminative 3D spectral-spatial features while keeping the 3D structure of HSI. This module introduces the large kernel attention (LKA) and convolution feed-forward (CFF) to flexibly emphasize, model, and exploit the long-range 3D feature dependencies with lower computational pressure. Finally, the features from the spectral-spatial attention module are fed into the classification module for the optimization of 3D spectral-spatial representation. To verify the effectiveness of the proposed classification method, experiments are executed on four widely used HSI data sets. The experiments demonstrate that LKSSAN is indeed an effective way for long-range 3D feature extraction of HSI. Genyun Sun, Zhaojie Pan, Aizhu Zhang, Xiuping Jia, Jinchang Ren, Kai Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Crowdsourced Quality Assessment of Enhanced Underwater Images - a Pilot StudyabstractUnderwater image enhancement (UIE) is essential for a high-quality underwater optical imaging system. While a number of UIE algorithms have been proposed in recent years, there is little study on image quality assessment (IQA) of enhanced underwater images. In this paper, we conduct the first crowdsourced subjective IQA study on enhanced underwater images. We chose ten state-of-the-art UIE algorithms and applied them to yield enhanced images from an underwater image benchmark. Their latent quality scales were reconstructed from pair comparison. We demonstrate that the existing IQA metrics are not suitable for assessing the perceived quality of enhanced underwater images. In addition, the overall performance of 10 UIE algorithms on the benchmark is ranked by the newly proposed simulated pair comparison of the methods. Hanhe Lin, Hui Men, Yijun Yan, Jinchang Ren, Dietmar Saupe |
QoMEX | 4 |
| 2022 | Multiscale voting mechanism for rice leaf disease recognition under natural field conditionsabstractRice leaf disease (RLD) is one of the major factors that cause the decline in production, and the automatic recognition of such diseases under natural field conditions is of great significance for timely targeted rice management. Although many machine learning approaches have been proposed for RLD recognition, scale variation is still a challenging problem that affects prediction accuracy, especially in uncontrolled environments, such as natural fields. Also, the existing RLD data sets are collected in laboratory environments or with a constant scale, which cannot be used to develop the RLD classification algorithms under natural field conditions. To tackle these particular challenges, we propose a multiscale voting mechanism for RLD recognition under natural field conditions. First, data from 26 rice fields were collected to build a data set containing 6046 images of RLD. Afterwards, a feature pyramid was embedded into a mainstream classification architecture (EfficientNet) with a bottom-up and top-down pathway for feature fusion at different scales. To further reduce the inconsistency among multiscaled features, a multiscale voting strategy with regard to probability distribution was proposed to integrate the decisions from various scales. Each proposed module was carefully validated through an ablation study to demonstrate its effectiveness, and the proposed method was compared with a few state-of-the-art algorithms, including the Single Shot MultiBox Detector, Feature Pyramid Networks, Path Aggregation Network, and Bidirectional Feature Pyramid Network. Experimental results have shown that the classification accuracy of our model can reach 90.24%, which is 4.48% higher than that of the original EfficientNet-b0 model and 1.08% higher than that of existing multiscale networks. Finally, we exploit and demonstrate a visualized explanation for the boosted performance from the proposed model. As an extra outcome, our data set and codes are available at http://github.com/huanghsheng/multiscale-voting-mechanism to benefit the whole research community. Yu Tang 0002, Jinfei Zhao, Huasheng Huang, Jiajun Zhuang, Zhiping Tan, Chaojun Hou, Weizhao Chen, Jinchang Ren |
Int. J. Intell. Syst. | 8 |
| 2022 | DRIN: Deep Recurrent Interaction Network for click-through rate prediction
Zhao Xudong, Xinying Xu, Jinchang Ren, Li Xingbing |
Inf. Sci. | 5 |
| 2022 | Novel hyperbolic clustering-based band hierarchy (HCBH) for effective unsupervised band selection of hyperspectral images
He Sun 0009, Lei Zhang 0054, Jinchang Ren, Hua Huang 0001 |
Pattern Recognit. | 3 |
| 2022 | A Novel Robust Low-rank Multi-view Diversity Optimization Model with Adaptive-Weighting Based Manifold Learning
Junpeng Tan, Zhijing Yang, Jinchang Ren, Yongqiang Cheng 0001, Bingo Wing-Kuen Ling |
Pattern Recognit. | 3 |
| 2022 | Effective extraction of ventricles and myocardium objects from cardiac magnetic resonance images with a multi-task learning U-Net
Jinchang Ren, He Sun 0009, Huimin Zhao 0001, Hao Gao 0002, Calum MacLellan, Sophia Zhao |
Pattern Recognit. Lett. | 1 |
| 2022 | SAM-Net: Semantic probabilistic and attention mechanisms of dynamic objects for self-supervised depth and camera pose estimation in visual odometry applications
Binchao Yang, Xinying Xu, Jinchang Ren |
Pattern Recognit. Lett. | 3 |
| 2022 | SpaSSA: Superpixelwise Adaptive SSA for Unsupervised Spatial-Spectral Feature Extraction in Hyperspectral ImageabstractSingular spectral analysis (SSA) has recently been successfully applied to feature extraction in hyperspectral image (HSI), including conventional (1-D) SSA in spectral domain and 2-D SSA in spatial domain. However, there are some drawbacks, such as sensitivity to the window size, high computational complexity under a large window, and failing to extract joint spectral-spatial features. To tackle these issues, in this article, we propose superpixelwise adaptive SSA (SpaSSA), that is superpixelwise adaptive SSA for exploiting local spatial information of HSI. The extraction of local (instead of global) features, particularly in HSI, can be more effective for characterizing the objects within an image. In SpaSSA, conventional SSA and 2-D SSA are combined and adaptively applied to each superpixel derived from an oversegmented HSI. According to the size of the derived superpixels, either SSA or 2-D singular spectrum analysis (2D-SSA) is adaptively applied for feature extraction, where the embedding window in 2D-SSA is also adaptive to the size of the superpixel. Experimental results on the three datasets have shown that the proposed SpaSSA outperforms both SSA and 2D-SSA in terms of classification accuracy and computational complexity. By combining SpaSSA with the principal component analysis (SpaSSA-PCA), the accuracy of land-cover analysis can be further improved, outperforming several state-of-the-art approaches. Genyun Sun, Jinchang Ren, Aizhu Zhang, Jaime Zabalza, Xiuping Jia, Huimin Zhao 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Adaptive Distance-Based Band Hierarchy (ADBH) for Effective Hyperspectral Band SelectionabstractBand selection has become a significant issue for the efficiency of the hyperspectral image (HSI) processing. Although many unsupervised band selection (UBS) approaches have been developed in the last decades, a flexible and robust method is still lacking. The lack of proper understanding of the HSI data structure has resulted in the inconsistency in the outcome of UBS. Besides, most of the UBS methods are either relying on complicated measurements or rather noise sensitive, which hinder the efficiency of the determined band subset. In this article, an adaptive distance-based band hierarchy (ADBH) clustering framework is proposed for UBS in HSI, which can help to avoid the noisy bands while reflecting the hierarchical data structure of HSI. With a tree hierarchy-based framework, we can acquire any number of band subset. By introducing a novel adaptive distance into the hierarchy, the similarity between bands and band groups can be computed straightforward while reducing the effect of noisy bands. Experiments on four datasets acquired from two HSI systems have fully validated the superiority of the proposed framework. He Sun 0009, Jinchang Ren, Huimin Zhao 0001, Genyun Sun, Wenzi Liao, Zhenyu Fang, Jaime Zabalza |
IEEE Trans. Cybern. | 2 |
| 2022 | Fusion of PCA and Segmented-PCA Domain Multiscale 2-D-SSA for Effective Spectral-Spatial Feature Extraction and Data Classification in Hyperspectral ImageryabstractAs hyperspectral imagery (HSI) contains rich spectral and spatial information, a novel principal component analysis (PCA) and segmented-PCA (SPCA)-based multiscale 2-D-singular spectrum analysis (2-D-SSA) fusion method is proposed for joint spectral–spatial HSI feature extraction and classification. Considering the overall spectra and adjacent band correlations of objects, the PCA and SPCA methods are utilized first for spectral dimension reduction, respectively. Then, multiscale 2-D-SSA is applied onto the SPCA dimension-reduced images to extract abundant spatial features at different scales, where PCA is applied again for dimensionality reduction. The obtained multiscale spatial features are then fused with the global spectral features derived from PCA to form multiscale spectral–spatial features (MSF-PCs). The performance of the extracted MSF-PCs is evaluated using the support vector machine (SVM) classifier. Experiments on four benchmark HSI data sets have shown that the proposed method outperforms other state-of-the-art feature extraction methods, including several deep learning approaches, when only a small number of training samples are available. Genyun Sun, Jinchang Ren, Aizhu Zhang, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Novel Band Selection and Spatial Noise Reduction Method for Hyperspectral Image ClassificationabstractAs an essential reprocessing method, dimensionality reduction (DR) can reduce the data redundancy and improve the performance of hyperspectral image (HSI) classification. A novel unsupervised DR framework with feature interpretability, which integrates both band selection (BS) and spatial noise reduction method, is proposed to extract low-dimensional spectral-spatial features of HSI. We proposed a new Neighboring band Grouping and Normalized Matching Filter (NGNMF) for BS, which can reduce the data dimension whilst preserve the corresponding spectral information. An enhanced 2-D singular spectrum analysis (E2DSSA) method is also proposed to extract the spatial context and structural information from each selected band, aiming to decrease the intra-class variability and reduce the effect of noise in the spatial domain. The support vector machine (SVM) classifier is used to evaluate the effectiveness of the extracted spectral-spatial low-dimensional features. Experimental results on three publicly available HSI datasets have fully demonstrated the efficacy of the proposed NGNMF-E2DSSA method, which has surpassed a number of state-of-the-art DR methods. Aizhu Zhang, Genyun Sun, Jinchang Ren, Xiuping Jia, Zhaojie Pan, Hongzhang Ma |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Object-Based Attention Mechanism for Color Calibration of UAV Remote Sensing Images in Precision AgricultureabstractColor calibration is a critical step for Unmanned Aerial Vehicles (UAV) remote sensing, especially in precision agriculture, which relies mainly on correlating color changes to specific quality attributes, e.g., plant health, disease, and pest stresses. In UAV remote sensing, the exemplar-based color transfer is popularly used for color calibration, where the automatic search for the semantic correspondences is the key to ensuring the color transfer accuracy. However, the existing attention mechanisms encounter difficulties in building the precise semantic correspondences between the reference image and the target one, in which the normalized cross correlation is often computed for feature reassembling. As a result, the color transfer accuracy is inevitably decreased by the disturbance from the semantically unrelated pixels, leading to semantic mismatch due to the absence of semantic correspondences. In this paper, we proposed an unsupervised object-based attention mechanism (OBAM) to suppress the disturbance of the semantically unrelated pixels, along with a further introduced weight-adjusted AdaIN (WAA) method to tackle the challenges caused by the absence of semantic correspondences. By embedding the proposed modules into a photorealistic style transfer method with progressive stylization, the color transfer accuracy can be improved while better preserving the structural details. We evaluated our approach on the UAV data of different crop types including rice, beans, and cotton. Extensive experiments demonstrate that our proposed method outperforms several state-of-the-art methods. As our approach requires no annotated labels, it can be easily embedded into the off-the-shelf color transfer approaches. Relevant codes and configurations will be available at http://github.com/huanghsheng/object-based-attention-mechanism. Huasheng Huang, Yu Tang 0002, Zhiping Tan, Jiajun Zhuang, Chaojun Hou, Weizhao Chen, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Novel Gumbel-Softmax Trick Enabled Concrete Autoencoder With Entropy Constraints for Unsupervised Hyperspectral Band SelectionabstractAs an important topic in hyperspectral image (HSI) analysis, band selection has attracted increasing attention in the last two decades for dimensionality reduction in HSI. With the great success of deep learning (DL)-based models recently, a robust unsupervised band selection (UBS) neural network is highly desired, particularly due to the lack of sufficient ground truth information to train the DL networks. Existing DL models for band selection either depend on the class label information or have unstable results via ranking the learned weights. To tackle these challenging issues, in this article, we propose a Gumbel-Softmax (GS) trick enabled concrete autoencoder-based UBS framework (CAE-UBS) for HSI, in which the learning process is featured by the introduced concrete random variables and the reconstruction loss. By searching from the generated potential band selection candidates from the concrete encoder, the optimal band subset can be selected based on an information entropy (IE) criterion. The idea of the CAE-UBS is quite straightforward, which does not rely on any complicated strategies or metrics. The robust performance on four publicly available datasets has validated the superiority of our CAE-UBS framework in the classification of the HSIs. He Sun 0009, Jinchang Ren, Huimin Zhao 0001, Peter W. T. Yuen, Julius Tschannerl |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spectral-Spatial Self-Attention Networks for Hyperspectral Image ClassificationabstractThis study presents a spectral–spatial self-attention network (SSSAN) for classification of hyperspectral images (HSIs), which can adaptively integrate local features with long-range dependencies related to the pixel to be classified. Specifically, it has two subnetworks. The spatial subnetwork introduces the proposed spatial self-attention module to exploit rich patch-based contextual information related to the center pixel. The spectral subnetwork introduces the proposed spectral self-attention module to exploit the long-range spectral correlation over local spectral features. The extracted spectral and spatial features are then adaptively fused for HSI classification. Experiments conducted on four HSI datasets demonstrate that the proposed network outperforms several state-of-the-art methods. Genyun Sun, Xiuping Jia, Lixin Wu, Aizhu Zhang, Jinchang Ren, Yanjuan Yao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Bayesian Gravitation-Based Classification for Hyperspectral ImagesabstractIntegration of spectral and spatial information is extremely important for the classification of high-resolution hyperspectral images (HSIs). Gravitation describes interaction among celestial bodies which can be applied to measure similarity between data for image classification. However, gravitation is hard to combine with spatial information and rarely been applied in HSI classification. This paper proposes a Bayesian Gravitation based Classification (BGC) to integrate the spectral and spatial information of local neighbors and training samples. In the BGC method, each testing pixel is first assumed as a massive object with unit volume and a particular density, where the density is taken as the data mass in BGC. Specifically, the data mass is formulated as an exponential function of the spectral distribution of its neighbors and the spatial prior distribution of its surrounding training samples based on the Bayesian theorem. Then, a joint data gravitation model is developed as the classification measure, in which the data mass is taken to weigh the contribution of different neighbors in a local region. Four benchmark HSI datasets, i.e. the Indian Pines, Pavia University, Salinas, and Grss_dfc_2014, are tested to verify the BGC method. The experimental results are compared with that of several well-known HSI classification methods, including the support vector machines, sparse representation, and other eight state-of-the-art HSI classification methods. The BGC shows apparent superiority in the classification of high-resolution HSIs and also flexibility for HSIs with limited samples. Aizhu Zhang, Genyun Sun, Zhaojie Pan, Jinchang Ren, Xiuping Jia, Yanjuan Yao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | An Unsupervised Domain Adaptation Method Towards Multi-Level Features and Decision Boundaries for Cross-Scene Hyperspectral Image ClassificationabstractDespite success in the same-scene hyperspectral image classification (HSIC), for the cross-scene classification, samples between source and target scenes are not drawn from the independent and identical distribution, resulting in significant performance degradation. To tackle this issue, a novel unsupervised domain adaptation (UDA) framework toward multilevel features and decision boundaries (ToMF-B) is proposed for the cross-scene HSIC, which can align task-related features and learn task-specific decision boundaries in parallel. Based on the maximum classifier discrepancy, a two-stage alignment scheme is proposed to bridge the interdomain gap and generate discriminative decision boundaries. In addition, to fully learn task-related and domain-confusing features, a convolutional neural network (CNN) and Transformer-based multilevel features extractor (generator) is developed to enrich the feature representation of two domains. Furthermore, to alleviate the harm even the negative transfer to UDA caused by task-irrelevant features, a task-oriented feature decomposition method is leveraged to enhance the task-related features while suppressing task-irrelevant features, and enabling the aligned domain-invariant features can be contributed to the classification task explicitly. Extensive experiments on three cross-scene HSI benchmarks have validated the effectiveness of the proposed framework. Chunhui Zhao 0003, Boao Qin, Shou Feng, Wenxiang Zhu, Lifu Zhang 0002, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | SC2Net: A Novel Segmentation-Based Classification Network for Detection of COVID-19 in Chest X-Ray ImagesabstractThe pandemic of COVID-19 has become a global crisis in public health, which has led to a massive number of deaths and severe economic degradation. To suppress the spread of COVID-19, accurate diagnosis at an early stage is crucial. As the popularly used real-time reverse transcriptase polymerase chain reaction (RT-PCR) swab test can be lengthy and inaccurate, chest screening with radiography imaging is still preferred. However, due to limited image data and the difficulty of the early-stage diagnosis, existing models suffer from ineffective feature extraction and poor network convergence and optimisation. To tackle these issues, a segmentation-based COVID-19 classification network, namely SC2Net, is proposed for effective detection of the COVID-19 from chest x-ray (CXR) images. The SC2Net consists of two subnets: a COVID-19 lung segmentation network (CLSeg), and a spatial attention network (SANet). In order to supress the interference from the background, the CLSeg is first applied to segment the lung region from the CXR. The segmented lung region is then fed to the SANet for classification and diagnosis of the COVID-19. As a shallow yet effective classifier, SANet takes the ResNet-18 as the feature extractor and enhances high-level feature via the proposed spatial attention module. For performance evaluation, the COVIDGR 1.0 dataset is used, which is a high-quality dataset with various severity levels of the COVID-19. Experimental results have shown that, our SC2Net has an average accuracy of 84.23% and an average F1 score of 81.31% in detection of COVID-19, outperforming several state-of-the-art approaches. Huimin Zhao 0001, Zhenyu Fang, Jinchang Ren, Calum MacLellan, Yong Xia 0001, Shuo Li 0001, Meijun Sun, Kevin Ren |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Sparse learning of band power features with genetic channel selection for effective classification of EEG signals
Natasha Padfield, Jinchang Ren, Paul Murray, Huimin Zhao 0001 |
Neurocomputing | 2 |
| 2021 | Fast Blind Deblurring of QR Code Images Based on Adaptive Scale ControlabstractAbstract With the development of 5G technology, the short delay requirements of commercialization and large amounts of data change our lifestyle day-to-day. In this background, this paper proposes a fast blind deblurring algorithm for QR code images, which mainly achieves the effect of adaptive scale control by introducing an evaluation mechanism. Its main purpose is to solve the out-of-focus caused by lens shake, inaccurate focus, and optical noise by speeding up the latent image estimation in the process of multi-scale division iterative deblurring. The algorithm optimizes productivity under the guidance of collaborative computing, based on the characteristics of the QR codes, such as the features of gradient and strength. In the evaluation step, the Tenengrad method is used to evaluate the image quality, and the evaluation value is compared with the empirical value obtained from the experimental data. Combining with the error correction capability, the recognizable QR codes will be output. In addition, we introduced a scale control parameter to study the relationship between the recognition rate and restoration time. Theoretical analysis and experimental results show that the proposed algorithm has high recovery efficiency and well recovery effect, can be effectively applied in industrial applications. Rongjun Chen 0001, Zhijun Zheng, Junfeng Pan, Yongxing Yu, Huimin Zhao 0001, Jinchang Ren |
Mob. Networks Appl. | 6 |
| 2021 | Topological optimization of the DenseNet with pretrained-weights inheritance and genetic channel selection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Song Wang 0002, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2021 | EACOFT: An energy-aware correlation filter for visual tracking
Qiaoyuan Liu, Jinchang Ren, Yuru Wang, Yuanbo Wu, Haijiang Sun, Huimin Zhao 0001 |
Pattern Recognit. | 2 |
| 2021 | Iterative Enhanced Multivariance Products Representation for Effective Compression of Hyperspectral ImagesabstractEffective compression of hyperspectral (HS) images is essential due to their large data volume. Since these images are high dimensional, processing them is also another challenging issue. In this work, an efficient lossy HS image compression method based on enhanced multivariance products representation (EMPR) is proposed. As an efficient data decomposition method, EMPR enables us to represent the given multidimensional data with lower-dimensional entities. EMPR, as a finite expansion with relevant approximations, can be acquired by truncating this expansion at certain levels. Thus, EMPR can be utilized as a highly effective lossy compression algorithm for hyper spectral images. In addition to these, an efficient variety of EMPR is also introduced in this article, in order to increase the compression efficiency. The results are benchmarked with several state-of-the-art lossy compression methods. It is observed that both higher peak signal-to-noise ratio values and improved classification accuracy are achieved from EMPR-based methods. Suha Tuna, B. Ugur Töreyin, Metin Demiralp, Jinchang Ren, Huimin Zhao 0001, Stephen Marshall |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | 2D-SSA Based Multiscale Feature Fusion for Feature Extraction and Data Classification in Hyperspectral ImageryabstractSingular spectrum analysis (SSA) and its 2-D variation (2D-SSA) have been successfully applied for effective feature extraction in hyperspectral imaging (HSI). However, they both cannot effectively use the spectral-spatial information, leading to a limited accuracy in classification. To tackle this problem, a novel 2D-SSA based multiscale feature fusion method, combining with segmented principal component analysis (SPCA), is proposed in this paper. The SPCA method is used for dimension reduction and spectral feature extraction, while multiscale 2D-SSA can extract abundant spatial features at different scales. In addition, a postprocessing via SPCA is applied on fused features to enhance the spectral discriminability. Experiments on two widely used datasets show that the proposed method outperforms two conventional SSA methods and other spectral-spatial classification methods in terms of the classification accuracy and computational cost. Genyun Sun, Jinchang Ren, Jaime Zabalza, Aizhu Zhang, Yanjuan Yao |
IGARSS | 3 |
| 2020 | Generic wavelet-based image decomposition and reconstruction framework for multi-modal data analysis in smart camera applicationsabstractEffective acquisition, analysis and reconstruction of multi‐modal data such as colour and multi‐/hyper‐spectral imagery is crucial in smart camera applications, where wavelet‐based coding and compression of images are highly demanded. Many existing discrete wavelet filtering banks have fixed coefficients hence their performance is highly dependent on the signal/image being processed. To tackle this problem, a unified framework is proposed in this study, which can produce a series of discrete wavelet filtering banks, where many existing discrete wavelet filtering banks become special cases of the framework. For each generated filtering bank, it consists of two decomposition filters and two reconstruction filters through an optimisation process. The efficacy of the filtering banks produced by the framework has been validated in two case studies, including colour image decomposition and reconstruction, and hyperspectral image classification. Comprehensive experiments have demonstrated the superior performance of the proposed framework, which will benefit the efficacy of smart camera and camera network applications. Yijun Yan, Yiguang Liu, Huimin Zhao 0001, Yanmei Chai, Jinchang Ren |
IET Comput. Vis. | 6 |
| 2020 | Triple loss for hard face detection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Zheng Wang 0008, Kaizhu Huang, Bing Xiao 0005 |
Neurocomputing | 2 |
| 2020 | Boundary-aware High-resolution Network with region enhancement for salient object detection
Zheng Wang 0008, Qinghua Hu, Jinchang Ren, Meijun Sun |
Neurocomputing | 4 |
| 2020 | Combining t-Distributed Stochastic Neighbor Embedding With Convolutional Neural Networks for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs), featured by high spectral resolution over a wide range of electromagnetic spectra, have been widely used to characterize materials with subtle differences in the spectral domain. However, a large number of bands and an insufficient number of sample pixels for each class are challenging for traditional machine learning-based classifiers. As alternative tools for feature extraction, neural networks have received extensive attention. This letter proposes to combine t-distributed stochastic neighbor embedding (t-SNE) with a convolutional neural network (CNN) for HSI classification. Our framework is designed to automatically capture the potential assembly features, which are extracted from both the dimension-reduced CNN (DR-CNN) and the multiscale-CNN. Experimental results show that the proposed classification framework outperforms several state-of-the-art techniques for three real data sets. Lianru Gao, Daixin Gu, Lina Zhuang, Jinchang Ren, Bing Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | MIMN-DPP: Maximum-information and minimum-noise determinantal point processes for unsupervised hyperspectral band selection
Weizhao Chen, Zhijing Yang, Jinchang Ren, Jiang-Zhong Cao, Nian Cai, Huimin Zhao 0001, Peter W. T. Yuen |
Pattern Recognit. | 3 |
| 2020 | A Novel Intelligent Computational Approach to Model Epidemiological Trends and Assess the Impact of Non-Pharmacological Interventions for COVID-19abstractThe novel coronavirus disease 2019 (COVID-19) pandemic has led to a worldwide crisis in public health. It is crucial we understand the epidemiological trends and impact of non-pharmacological interventions (NPIs), such as lockdowns for effective management of the disease and control of its spread. We develop and validate a novel intelligent computational model to predict epidemiological trends of COVID-19, with the model parameters enabling an evaluation of the impact of NPIs. By representing the number of daily confirmed cases (NDCC) as a time-series, we assume that, with or without NPIs, the pattern of the pandemic satisfies a series of Gaussian distributions according to the central limit theorem. The underlying pandemic trend is first extracted using a singular spectral analysis (SSA) technique, which decomposes the NDCC time series into the sum of a small number of independent and interpretable components such as a slow varying trend, oscillatory components and structureless noise. We then use a mixture of Gaussian fitting (GF) to derive a novel predictive model for the SSA extracted NDCC incidence trend, with the overall model termed SSA-GF. Our proposed model is shown to accurately predict the NDCC trend, peak daily cases, the length of the pandemic period, the total confirmed cases and the associated dates of the turning points on the cumulated NDCC curve. Further, the three key model parameters, specifically, the amplitude (alpha), mean (mu), and standard deviation (sigma) are linked to the underlying pandemic patterns, and enable a directly interpretable evaluation of the impact of NPIs, such as strict lockdowns and travel restrictions. The predictive model is validated using available data from China and South Korea, and new predictions are made, partially requiring future validation, for the cases of Italy, Spain, the UK and the USA. Comparative results demonstrate that the introduction of consistent control measures across countries can lead to development of similar parametric models, reflected in particular by relative variations in their underlying sigma, alpha and mu values. The paper concludes with a number of open questions and outlines future research directions. Jinchang Ren, Yijun Yan, Huimin Zhao 0001, Ping Ma 0002, Jaime Zabalza, Zain U. Hussain, Shaoming Luo, Sophia Zhao, Aziz Sheikh, Amir Hussain 0001, Huakang Li |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | SR-GAN: Semantic Rectifying Generative Adversarial Network for Zero-shot LearningabstractThe existing Zero-Shot learning (ZSL) methods may suffer from the vague class attributes that are highly overlapped for different classes. Unlike these methods that ignore the discrimination among classes, in this paper, we propose to classify unseen image by rectifying the semantic space guided by the visual space. First, we pre-train a Semantic Rectifying Network (SRN) to rectify semantic space with a semantic loss and a rectifying loss. Then, a Semantic Rectifying Generative Adversarial Network (SR-GAN) is built to generate plausible visual feature of unseen class from both semantic feature and rectified semantic feature. To guarantee the effectiveness of rectified semantic features and synthetic visual features, a pre-reconstruction and a post reconstruction networks are proposed, which keep the consistency between visual feature and semantic feature. Experimental results demonstrate that our approach significantly outperforms the state-of-the-arts on four benchmark datasets. Zihan Ye, Fan Lyu, Qiming Fu 0001, Jinchang Ren, Fuyuan Hu |
ICME | 5 |
| 2019 | Spatial-spectral classification of hyperspectral images: a deep learning framework with Markov Random fields based modellingabstractFor the spatial‐spectral classification of hyperspectral images (HSIs), a deep learning framework is proposed in this study, which consists of convolutional neural networks (CNNs) and Markov random fields (MRFs). Firstly, a CNN model to learn the deep spectral feature from the HSI is built and the class posterior probability distribution is estimated. The CNN with a dropout layer can relieve the overfitting in classification. The CNN is utilised as a pixel‐classifier, so it only works in the spectral domain. Then, the spatial information will be encoded by MRF‐based multilevel logistic prior for regularising the classification. To derive the correlation of both spectral and spatial features for improving algorithm performance, the marginal probability distribution in HSI is learned using MRF‐based loopy belief propagation. In comparison with several state‐of‐the‐art approaches for data classification on three publicly available HSI datasets, experimental results have demonstrated the superior performance of the proposed methodology. Chunmei Qing, Jiawei Ruan, Xiangmin Xu 0001, Jinchang Ren, Jaime Zabalza |
IET Image Process. | 4 |
| 2019 | Compressive sensing based secret signals recovery for effective image Steganalysis in secure communications
Huimin Zhao 0001, Jinchang Ren, Jin Zhan, Yinyin Xiao, Sophia Zhao, Fangyuan Lei, Maher Assaad |
Multim. Tools Appl. | 2 |
| 2019 | Local Block Multilayer Sparse Extreme Learning Machine for Effective Feature Extraction and Classification of Hyperspectral ImagesabstractAlthough extreme learning machines (ELM) have been successfully applied for the classification of hyperspectral images (HSIs), they still suffer from three main drawbacks. These include: 1) ineffective feature extraction (FE) in HSIs due to a single hidden layer neuron network used; 2) ill-posed problems caused by the random input weights and biases; and 3) lack of spatial information for HSIs classification. To tackle the first problem, we construct a multilayer ELM for effective FE from HSIs. The sparse representation is adopted with the multilayer ELM to tackle the ill-posed problem of ELM, which can be solved by the alternative direction method of multipliers. This has resulted in the proposed multilayer sparse ELM (MSELM) model. Considering that the neighboring pixels are more likely from the same class, a local block extension is introduced for MSELM to extract the local spatial information, leading to the local block MSELM (LBMSELM). The loopy belief propagation is also applied to the proposed MSELM and LBMSELM approaches to further utilize the rich spectral and spatial information for improving the classification. Experimental results show that the proposed methods have outperformed the ELM and other state-of-the-art approaches. Faxian Cao, Zhijing Yang, Jinchang Ren, Weizhao Chen, Guojun Han, Yuzhen Shen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Spectral-Spatial Topographic Shadow Detection from Sentinel-2A MSI Imagery Via Convolutional Neural NetworksabstractAccurate detection of topographic shadows is of great importance, since topographic shadowing is an inevitable hamper for the interpretation of remotely sensed images covered mountainous areas. In this paper, a novel method is proposed for effective and efficient topographic shadow detection for the images obtained from Sentinel-2A multispectral imager (MSI) by combining both the spectral and spatial information. In this method, four feature indices were firstly extracted from the original Sentinel-2A spectral bands to capture the essential spectral characteristics. Specifically, we constructed a topographic shadow index (TSI) to enhance topographic shadows, focusing on the spectral signatures of shadows in Sentinel-2A MSI imagery. To further enhance the difference between shadows and other objects, the TSI is combined with the first component of the principal component analysis (FPC), the soil-adjusted vegetation index (SAVI) and the normalized water index (NDWI), representing topographic shadows, rocks, vegetation and water, respectively. Finally, a convolutional neural network (CNN) was used by operating directly on indices input due to its remarkable classification performance, which exploits the spatial contextual information and spectral features for effective topographic extraction. Our experiments with one Sentinel-2A image show that the proposed approach has led to satisfactory performances, with few errors in shadow maps and insignificant confusion with spectral-similar land covers. Hui Huang 0016, Genyun Sun, Jinchang Ren, Jun Rong, Aizhu Zhang, Yanling Hao |
IGARSS | 3 |
| 2018 | A deep-learning based feature hybrid framework for spatiotemporal saliency detection inside videos
Zheng Wang 0008, Jinchang Ren, Meijun Sun, Jianmin Jiang |
Neurocomputing | 2 |
| 2018 | A stability constrained adaptive alpha for gravitational search algorithm
Genyun Sun, Ping Ma 0002, Jinchang Ren, Aizhu Zhang, Xiuping Jia |
Knowl. Based Syst. | 3 |
| 2018 | Spectral-spatial classification of hyperspectral data using spectral-domain local binary patterns
Cailing Wang, Jinchang Ren, Yinyong Zhang |
Multim. Tools Appl. | 2 |
| 2018 | Robust information hiding in low-resolution videos with quantization index modulation in DCT-CS domain
Huimin Zhao 0001, Jinchang Ren, Wenguo Wei, Yinyin Xiao |
Multim. Tools Appl. | 3 |
| 2018 | Joint bilateral filtering and spectral similarity-based sparse representation: A generic framework for effective feature extraction and data classification in hyperspectral imaging
Zhijing Yang, Jinchang Ren, Peter W. T. Yuen, Huimin Zhao 0001, Genyun Sun, Stephen Marshall, Jón Atli Benediktsson |
Pattern Recognit. | 3 |
| 2018 | Unsupervised image saliency detection with Gestalt-laws guided optimization and visual attention based refinement
Yijun Yan, Jinchang Ren, Genyun Sun, Huimin Zhao 0001, Junwei Han 0001, Xuelong Li 0001, Stephen Marshall, Jin Zhan |
Pattern Recognit. | 2 |
| 2018 | A Dynamic Neighborhood Learning-Based Gravitational Search AlgorithmabstractBalancing exploration and exploitation according to evolutionary states is crucial to meta-heuristic search (M-HS) algorithms. Owing to its simplicity in theory and effectiveness in global optimization, gravitational search algorithm (GSA) has attracted increasing attention in recent years. However, the tradeoff between exploration and exploitation in GSA is achieved mainly by adjusting the size of an archive, named , which stores those superior agents after fitness sorting in each iteration. Since the global property of remains unchanged in the whole evolutionary process, GSA emphasizes exploitation over exploration and suffers from rapid loss of diversity and premature convergence. To address these problems, in this paper, we propose a dynamic neighborhood learning (DNL) strategy to replace the model and thereby present a DNL-based GSA (DNLGSA). The method incorporates the local and global neighborhood topologies for enhancing the exploration and obtaining adaptive balance between exploration and exploitation. The local neighborhoods are dynamically formed based on evolutionary states. To delineate the evolutionary states, two convergence criteria named limit value and population diversity, are introduced. Moreover, a mutation operator is designed for escaping from the local optima on the basis of evolutionary states. The proposed algorithm was evaluated on 27 benchmark problems with different characteristic and various difficulties. The results reveal that DNLGSA exhibits competitive performances when compared with a variety of state-of-the-art M-HS algorithms. Moreover, the incorporation of local neighborhood topology reduces the numbers of calculations of gravitational force and thus alleviates the high computational cost of GSA. Aizhu Zhang, Genyun Sun, Jinchang Ren, Xiaodong Li 0001, Xiuping Jia |
IEEE Trans. Cybern. | 3 |
| 2018 | Sparse Representation-Based Augmented Multinomial Logistic Extreme Learning Machine With Weighted Composite Features for Spectral-Spatial Classification of Hyperspectral ImagesabstractAlthough extreme learning machine (ELM) has successfully been applied to a number of pattern recognition problems, only with the original ELM it can hardly yield high accuracy for the classification of hyperspectral images (HSIs) due to two main drawbacks. The first is due to the randomly generated initial weights and bias, which cannot guarantee optimal output of ELM. The second is the lack of spatial information in the classifier as the conventional ELM only utilizes spectral information for classification of HSI. To tackle these two problems, a new framework for ELM-based spectral-spatial classification of HSI is proposed, where probabilistic modeling with sparse representation and weighted composite features (WCFs) is employed to derive the optimized output weights and extract spatial features. First, ELM is represented as a concave logarithmic-likelihood function under statistical modeling using the maximum a posteriori estimator. Second, sparse representation is applied to the Laplacian prior to efficiently determine a logarithmic posterior with a unique maximum in order to solve the ill-posed problem of ELM. The variable splitting and the augmented Lagrangian are subsequently used to further reduce the computation complexity of the proposed algorithm. Third, the spatial information is extracted using the WCFs to construct the spectral-spatial classification framework. In addition, the lower bound of the proposed method is derived by a rigorous mathematical proof. Experimental results on three publicly available HSI data sets demonstrate that the proposed methodology outperforms ELM and also a number of state-of-the-art approaches. Faxian Cao, Zhijing Yang, Jinchang Ren, Bingo Wing-Kuen Ling, Huimin Zhao 0001, Meijun Sun, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Effective Denoising and Classification of Hyperspectral Images Using Curvelet Transform and Singular Spectrum AnalysisabstractHyperspectral imaging (HSI) classification has become a popular research topic in recent years, and effective feature extraction is an important step before the classification task. Traditionally, spectral feature extraction techniques are applied to the HSI data cube directly. This paper presents a novel algorithm for HSI feature extraction by exploiting the curvelet-transformed domain via a relatively new spectral feature processing technique—singular spectrum analysis (SSA). Although the wavelet transform has been widely applied for HSI data analysis, the curvelet transform is employed in this paper since it is able to separate image geometric details and background noise effectively. Using the support vector machine classifier, experimental results have shown that features extracted by SSA on curvelet coefficients have better performance in terms of classification accuracy over features extracted on wavelet coefficients. Since the proposed approach mainly relies on SSA for feature extraction on the spectral dimension, it actually belongs to the spectral feature extraction category. Therefore, the proposed method has also been compared with some state-of-the-art spectral feature extraction techniques to show its efficacy. In addition, it has been proven that the proposed method is able to remove the undesirable artifacts introduced during the data acquisition process. By adding an extra spatial postprocessing step to the classified map achieved using the proposed approach, we have shown that the classification performance is comparable with several recent spectral–spatial classification methods. Jinchang Ren, Zheng Wang 0008, Jaime Zabalza, Meijun Sun, Huimin Zhao 0001, Shutao Li 0001, Jón Atli Benediktsson, Stephen Marshall |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Monte Carlo Convex Hull Model for classification of traditional Chinese paintings
Meijun Sun, Zheng Wang 0008, Jinchang Ren, Jesse S. Jin |
Neurocomputing | 4 |
| 2016 | Novel segmented stacked autoencoder for effective dimensionality reduction and feature extraction in hyperspectral imaging
Jaime Zabalza, Jinchang Ren, Jiangbin Zheng 0001, Huimin Zhao 0001, Chunmei Qing, Zhijing Yang, Peijun Du, Stephen Marshall |
Neurocomputing | 2 |
| 2016 | Combining Morphological Attribute Profiles via an Ensemble Method for Hyperspectral Image ClassificationabstractMorphological attribute profiles (APs) are discriminant features in the spectral–spatial classification of hyperspectral data. However, the optimal range of parameters in each filter is always a challenging yet important task, since an unsuitable range of parameters likely leads to inferior results. In order to alleviate this problem, we propose an ensemble method, which integrates multiple classification results based on a series of APs. The APs are obtained by using different filters with thresholds that are randomly selected from an arbitrarily defined range of parameters. Experimental results conducted on two hyperspectral images demonstrate the robustness and effectiveness of the proposed method. Rui Bao, Junshi Xia, Mauro Dalla Mura, Peijun Du, Jocelyn Chanussot, Jinchang Ren |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2016 | Hierarchical and multi-featured fusion for effective gait recognition under variable scenarios
Yanmei Chai, Jie Ren 0014, Huimin Zhao 0001, Yang Li 0145, Jinchang Ren, Paul Murray |
Pattern Anal. Appl. | 5 |
| 2015 | Brushstroke based sparse hybrid convolutional neural networks for author classification of Chinese ink-wash paintingsabstractA novel stroke based sparse hybrid convolutional neural networks (CNNs) method is proposed for author classification of Chinese ink-wash paintings (IWPs). As Chinese IWPs usually have many authors in several art styles, this differs from real images or western paintings and has led to a big challenge. In our work, we classify Chinese IWPs of different artists by analyzing a set of automatically extracted brushstrokes. A sparse hybrid CNNs in a deep-learning framework is then proposed to extract brushstroke features to replace the commonly used handcrafted ones such as edge, color, intensity and texture. Using 120 IWPs from six famous artists, promising results have been shown in successfully classifying authors in comparison to two other state-of-the-art approaches. Meijun Sun, Jinchang Ren, Zheng Wang 0008, Jesse S. Jin |
ICIP | 3 |
| 2015 | Effective SAR sea ice image segmentation and touch floe separation using a combined multi-stage approachabstractAccurate sea-ice segmentation from satellite synthetic aperture radar (SAR) images plays an important role for understanding the interactions between sea-ice, ocean and atmosphere in the Arctic. Processing sea-ice SAR images are challenging due to poor spatial resolution and severe speckle noise. In this paper, we present a multi-stage method for the sea-ice SAR image segmentation, which includes edge-preserved filtering for preprocessing, k-means clustering for segmentation and conditional morphology filtering for post-processing. As such, the effect of noise has been suppressed and the under-segmented regions are successfully corrected. Jinchang Ren, Byongjun Hwang, Paul Murray, Soumitra Sakhalkar, Samuel McCormack |
IGARSS | 1 |
| 2015 | Segmentation of multispectral images and prediction of CHI-A concentration for effective ocean colour remote sensingabstractWith the development of new sensors and data processing techniques, ocean colour remote sensing has undergone rapid development in more accurately measurement of coastal shelf classification and concentration of chlorophyll. In this paper, multispectral images are employed to achieve these targets, using techniques including region-growing based segmentation for pixel classification and support vector regression for ChI-a prediction. Interesting results are reported to show the great potential in using state-of-the-art data analysis techniques for effective ocean colour remote sensing. Jinchang Ren, Xuexing Zeng, David McKee 0002 |
IGARSS | 1 |
| 2015 | Multiple Depth Maps Integration for 3D Reconstruction Using Geodesic Graph CutsabstractDepth images, in particular depth maps estimated from stereo vision, may have a substantial amount of outliers and result in inaccurate 3D modelling and reconstruction. To address this challenging issue, in this paper, a graph-cut based multiple depth maps integration approach is proposed to obtain smooth and watertight surfaces. First, confidence maps for the depth images are estimated to suppress noise, based on which reliable patches covering the object surface are determined. These patches are then exploited to estimate the path weight for 3D geodesic distance computation, where an adaptive regional term is introduced to deal with the "shorter-cuts" problem caused by the effect of the minimal surface bias. Finally, the adaptive regional term and the boundary term constructed using patches are combined in the graph-cut framework for more accurate and smoother 3D modelling. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multiple depth images captured by Kinect. The experimental results have shown that our method is able to preserve the object protrusions and details while maintaining surface smoothness. Jiangbin Zheng 0001, Xinxin Zuo, Jinchang Ren, Sen Wang 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2015 | Background Prior-Based Salient Object Detection via Deep Reconstruction ResidualabstractDetection of salient objects from images is gaining increasing research interest in recent years as it can substantially facilitate a wide range of content-based multimedia applications. Based on the assumption that foreground salient regions are distinctive within a certain context, most conventional approaches rely on a number of hand-designed features and their distinctiveness is measured using local or global contrast. Although these approaches have been shown to be effective in dealing with simple images, their limited capability may cause difficulties when dealing with more complicated images. This paper proposes a novel framework for saliency detection by first modeling the background and then separating salient objects from the background. We develop stacked denoising autoencoders with deep learning architectures to model the background where latent patterns are explored and more powerful representations of data are learned in an unsupervised and bottom-up manner. Afterward, we formulate the separation of salient objects from the background as a problem of measuring reconstruction residuals of deep autoencoders. Comprehensive evaluations of three benchmark datasets and comparisons with nine state-of-the-art algorithms demonstrate the superiority of this paper. Junwei Han 0001, Dingwen Zhang, Xintao Hu, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Effective and Efficient Midlevel Visual Elements-Oriented Land-Use Classification Using VHR Remote Sensing ImagesabstractLand-use classification using remote sensing images covers a wide range of applications. With more detailed spatial and textural information provided in very high resolution (VHR) remote sensing images, a greater range of objects and spatial patterns can be observed than ever before. This offers us a new opportunity for advancing the performance of land-use classification. In this paper, we first introduce an effective midlevel visual elementsoriented land-use classification method based on “partlets,” which are a library of pretrained part detectors used for midlevel visual elements discovery. Taking advantage of midlevel visual elements rather than low-level image features, a partlets-based method represents images by computing their responses to a large number of part detectors. As the number of part detectors grows, a main obstacle to the broader application of this method is its computational cost. To address this problem, we next propose a novel framework to train coarse-to-fine shared intermediate representations, which are termed “sparselets,” from a large number of pretrained part detectors. This is achieved by building a single-hidden-layer autoencoder and a single-hidden-layer neural network with an L0-norm sparsity constraint, respectively. Comprehensive evaluations on a publicly available 21-class VHR landuse data set and comparisons with state-of-the-art approaches demonstrate the effectiveness and superiority of this paper. Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Zhenbao Liu, Shuhui Bu, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2015 | Classification of Hyperspectral Images by Exploiting Spectral-Spatial Information of Superpixel via Multiple KernelsabstractFor the classification of hyperspectral images (HSIs), this paper presents a novel framework to effectively utilize the spectral-spatial information of superpixels via multiple kernels, which is termed as superpixel-based classification via multiple kernels (SC-MK). In the HSI, each superpixel can be regarded as a shape-adaptive region, which consists of a number of spatial neighboring pixels with very similar spectral characteristics. First, the proposed SC-MK method adopts an oversegmentation algorithm to cluster the HSI into many superpixels. Then, three kernels are separately employed for the utilization of the spectral information, as well as spatial information, within and among superpixels. Finally, the three kernels are combined together and incorporated into a support vector machine classifier. Experimental results on three widely used real HSIs indicate that the proposed SC-MK approach outperforms several well-known classification methods. Leyuan Fang, Shutao Li 0001, Wuhui Duan, Jinchang Ren, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Object Detection in Optical Remote Sensing Images Based on Weakly Supervised Learning and High-Level Feature LearningabstractThe abundant spatial and contextual information provided by the advanced remote sensing technology has facilitated subsequent automatic interpretation of the optical remote sensing images (RSIs). In this paper, a novel and effective geospatial object detection framework is proposed by combining the weakly supervised learning (WSL) and high-level feature learning. First, deep Boltzmann machine is adopted to infer the spatial and structural information encoded in the low-level and middle-level features to effectively describe objects in optical RSIs. Then, a novel WSL approach is presented to object detection where the training sets require only binary labels indicating whether an image contains the target object or not. Based on the learnt high-level features, it jointly integrates saliency, intraclass compactness, and interclass separability in a Bayesian framework to initialize a set of training examples from weakly labeled images and start iterative learning of the object detector. A novel evaluation criterion is also developed to detect model drift and cease the iterative learning. Comprehensive experiments on three optical RSI data sets have demonstrated the efficacy of the proposed approach in benchmarking with several state-of-the-art supervised-learning-based object detection approaches. Junwei Han 0001, Dingwen Zhang, Gong Cheng 0003, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Novel Two-Dimensional Singular Spectrum Analysis for Effective Feature Extraction and Data Classification in Hyperspectral ImagingabstractFeature extraction is of high importance for effective data classification in hyperspectral imaging (HSI). Considering the high correlation among band images, spectral-domain feature extraction is widely employed. For effective spatial information extraction, a 2-D extension to singular spectrum analysis (2D-SSA), which is a recent technique for generic data mining and temporal signal analysis, is proposed. With 2D-SSA applied to HSI, each band image is decomposed into varying trends, oscillations, and noise. Using the trend and the selected oscillations as features, the reconstructed signal, with noise highly suppressed, becomes more robust and effective for data classification. Three publicly available data sets for HSI remote sensing data classification are used in our experiments. Comprehensive results using a support vector machine classifier have quantitatively evaluated the efficacy of the proposed approach. Benchmarked with several state-of-the-art methods including 2-D empirical mode decomposition (2D-EMD), it is found that our proposed 2D-SSA approach generates the best results in most cases. Unlike 2D-EMD that requires sequential transforms to obtain detailed decomposition, 2D-SSA extracts all components simultaneously. As a result, the execution time in feature extraction can be also dramatically reduced. The superiority in terms of enhanced discrimination ability from 2D-SSA is further validated when a relatively weak classifier, i.e., the k-nearest neighbor, is used for data classification. In addition, the combination of 2D-SSA with 1-D principal component analysis (2D-SSA-PCA) has generated the best results among several other approaches, demonstrating the great potential in combining 2D-SSA with other approaches for effective spatial-spectral feature extraction and dimension reduction in HSI. Jaime Zabalza, Jinchang Ren, Jiangbin Zheng 0001, Junwei Han 0001, Huimin Zhao 0001, Shutao Li 0001, Stephen Marshall |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Gradient-based subspace phase correlation for fast and effective image alignment
Jinchang Ren, Theodore Vlachos, Jiangbin Zheng 0001, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Singular Spectrum Analysis for Effective Feature Extraction in Hyperspectral ImagingabstractAs a very recent technique for time-series analysis, singular spectrum analysis (SSA) has been applied in many diverse areas, where an original 1-D signal can be decomposed into a sum of components, including varying trends, oscillations, and noise. Considering pixel-based spectral profiles as 1-D signals, in this letter, SSA has been applied in hyperspectral imaging for effective feature extraction. By removing noisy components in extracting the features, the discriminating ability of the features has been much improved. Experiments show that this SSA approach supersedes the empirical mode decomposition technique from which our work was originally inspired, where improved results in effective data classification using support vector machine are also reported. Jaime Zabalza, Jinchang Ren, Zheng Wang 0008, Stephen Marshall, Jun Wang 0041 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Copulas for statistical signal processing (Part II): Simulation, optimal selection and practical applications
Xuexing Zeng, Jinchang Ren, Meijun Sun, Stephen Marshall, Tariq S. Durrani |
Signal Process. | 2 |
| 2014 | Copulas for statistical signal processing (Part I): Extensions and generalization
Xuexing Zeng, Jinchang Ren, Zheng Wang 0008, Stephen Marshall, Tariq S. Durrani |
Signal Process. | 2 |
| 2013 | Effective Classification of Microcalcification Clusters Using Improved Support Vector Machine with Optimised Decision MakingabstractClassification of micro calcification clusters is very essential for early detection of breast cancer from mammograms. In this paper, an improved support vector machine (SVM) scheme is proposed, where optimized decision making is introduced for effective and more accurate data classification. Experimental results on the well-known DDSM database have shown that the proposed method can significantly increase the performance in terms of F1 and Az measurements for the successful classification of clustered micro calcifications. Jinchang Ren, Zheng Wang 0008, Meijun Sun, John J. Soraghan |
ICIG | 1 |
| 2012 | ANN vs. SVM: Which one performs better in classification of MCCs in mammogram imaging
Jinchang Ren |
Knowl. Based Syst. | 1 |
| 2012 | Effective venue image retrieval using robust feature extraction and model constrained matching for mobile robot localization
Yue Feng 0002, Jinchang Ren, Jianmin Jiang, Martin Halvey, Joemon M. Jose |
Mach. Vis. Appl. | 2 |
| 2011 | Effective recognition of MCCs in mammograms using an improved neural classifier
Jinchang Ren, Jianmin Jiang |
Eng. Appl. Artif. Intell. | 1 |
| 2011 | Performance of hidden Markov model and dynamic Bayesian network classifiers on handwritten Arabic word recognition
Jawad Hasan Yasin AlKhateeb, Olivier Pauplin, Jinchang Ren, Jianmin Jiang |
Knowl. Based Syst. | 3 |
| 2011 | Modelling of content-aware indicators for effective determination of shot boundaries in compressed MPEG videos
Jinchang Ren, Jianmin Jiang |
Multim. Tools Appl. | 2 |
| 2011 | Offline handwritten Arabic cursive text recognition using Hidden Markov Models and re-ranking
Jawad Hasan Yasin AlKhateeb, Jinchang Ren, Jianmin Jiang, Husni Al-Muhtaseb |
Pattern Recognit. Lett. | 2 |
| 2010 | Activity-driven content adaptation for effective video summarization
Jinchang Ren, Jianmin Jiang, Yue Feng 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Multi-camera video surveillance for real-time analysis and reconstruction of soccer games
Jinchang Ren, Ming Xu 0011, James Orwell, Graeme A. Jones |
Mach. Vis. Appl. | 1 |
| 2010 | High-Accuracy Sub-Pixel Motion Estimation From Noisy Images in Fourier DomainabstractIn this paper, we propose a new method for estimating sub-pixel motion via exploiting the principle of phase correlation in the Fourier domain. The method is based on linear weighting of the height of the main peak on the one hand and the difference between its two neighboring side-peaks on the other. Using both synthetic and real data we show that the proposed method outperforms many established approaches and achieves improved accuracy even in the presence of noisy samples. Jinchang Ren, Jianmin Jiang, Theodore Vlachos |
IEEE Trans. Image Process. | 1 |
| 2009 | Tracking the soccer ball using multiple fixed cameras
Jinchang Ren, James Orwell, Graeme A. Jones, Ming Xu 0011 |
Comput. Vis. Image Underst. | 1 |
| 2009 | Shot Boundary Detection in MPEG Videos Using Local and Global IndicatorsabstractShot boundary detection (SBD) plays important roles in many video applications. In this letter, we describe a novel method on SBD operating directly in the compressed domain. First, several local indicators are extracted from MPEG macroblocks, and AdaBoost is employed for feature selection and fusion. The selected features are then used in classifying candidate cuts into five sub-spaces via pre-filtering and rule-based decision making. Following that, global indicators of frame similarity between boundary frames of cut candidates are examined using phase correlation of dc images. Gradual transitions like fade, dissolve, and combined shot cuts are also identified. Experimental results on the test data from TRECVID'07 have demonstrated the effectiveness and robustness of our proposed methodology. Jinchang Ren, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Hierarchical Modeling and Adaptive Clustering for Real-Time Summarization of Rush VideosabstractIn this paper, we provide detailed descriptions of a proposed new algorithm for video summarization, which are also included in our submission to TRECVID'08 on BBC rush summarization. Firstly, rush videos are hierarchically modeled using the formal language technique. Secondly, shot detections are applied to introduce a new concept of V-unit for structuring videos in line with the hierarchical model, and thus junk frames within the model are effectively removed. Thirdly, adaptive clustering is employed to group shots into clusters to determine retakes for redundancy removal. Finally, each most representative shot selected from every cluster is ranked according to its length and sum of activity level for summarization. Competitive results have been achieved to prove the effectiveness and efficiency of our techniques, which are fully implemented in the compressed domain. Our work does not require high-level semantics such as object detection and speech/audio analysis which provides a more flexible and general solution for this topic. Jinchang Ren, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2008 | Knowledge-Supported Segmentation and Semantic Contents Extraction from MPEG Videos for Highlight-Based Annotation, Indexing and Retrieval
Jinchang Ren, Jianmin Jiang, Stanley S. Ipson |
ICIC (1) | 1 |
| 2008 | Skin Detection from Different Color Spaces for Model-Based Face Detection
Jinchang Ren, Jianmin Jiang, Stanley S. Ipson |
ICIC (3) | 2 |
| 2008 | Real-Time Modeling of 3-D Soccer Ball Trajectories From Multiple Fixed CamerasabstractIn this paper, model-based approaches for real-time 3-D soccer ball tracking are proposed, using image sequences from multiple fixed cameras as input. The main challenges include filtering false alarms, tracking through missing observations, and estimating 3-D positions from single or multiple cameras. The key innovations are: 1. incorporating motion cues and temporal hysteresis thresholding in ball detection; 2. modeling each ball trajectory as curve segments in successive virtual vertical planes so that the 3-D position of the ball can be determined from a single camera view; and 3. introducing four motion phases (rolling, flying, in possession, and out of play) and employing phase-specific models to estimate ball trajectories which enables high-level semantics applied in low-level tracking. In addition, unreliable or missing ball observations are recovered using spatio-temporal constraints and temporal filtering. The system accuracy and robustness are evaluated by comparing the estimated ball positions and phases with manual ground-truth data of real soccer sequences. Jinchang Ren, Ming Xu 0011, James Orwell, Graeme A. Jones |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Statistical Classification of Skin Color Pixels from MPEG Videos
Jinchang Ren, Jianmin Jiang |
ACIVS | 1 |
| 2007 | Detection and Recovery of Film Dirt for Archive Restoration ApplicationsabstractA novel spatio-temporal method is proposed for film dirt detection and recovery. Firstly, a more reliable confidence measurement of dirt is extracted for color films. False alarms caused by motion are filtered using consistency checks among several measurements. Then, candidate dirt is detected by filtering and thresholding this confidence measurement. Finally, bi-directional local motion compensation and ML3Dex filtering are taken for the recovery of dirt pixels. Experiments on real data demonstrate the efficiency and effectiveness of our method in terms of both detection and recovery of dirt. Jinchang Ren, Theodore Vlachos |
ICIP (4) | 1 |
| 2007 | Subspace Extension to Phase Correlation Approach for Fast Image RegistrationabstractA novel extension of phase correlation to subspace correlation is proposed, in which 2-D translation is decomposed into two 1-D motions thus only 1-D Fourier transform is used to estimate the corresponding motion. In each subspace, the first two highest peaks from 1-D correlation are linearly interpolated for subpixel accuracy. Experimental results have shown both the robustness and accuracy of our method. Jinchang Ren, Theodore Vlachos, Jianmin Jiang |
ICIP (1) | 1 |
| 2007 | Efficient detection of temporally impulsive dirt impairments in archived films
Jinchang Ren, Theodore Vlachos |
Signal Process. | 1 |
| 2007 | Segmentation-Assisted Detection of Dirt Impairments in Archived Film SequencesabstractIn this correspondence, a novel segmentation-assisted method for film-dirt detection is proposed. We exploit the fact that film dirt manifests in the spatial domain as a cluster of connected pixels whose intensity differs substantially from that of its neighborhood, and we employ a segmentation-based approach to identify this type of structure. A key feature of our approach is the computation of a measure of confidence attached to detected dirt regions, which can be utilized for performance fine tuning. Another important feature of our algorithm is the avoidance of the computational complexity associated with motion estimation. Our experimental framework benefits from the availability of manually derived as well as objective ground-truth data obtained using infrared scanning. Our results demonstrate that the proposed method compares favorably with standard spatial, temporal, and multistage median-filtering approaches and provides efficient and robust detection for a wide variety of test materials. Jinchang Ren, Theodore Vlachos |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2005 | Segmentation-Assisted Dirt Detection for the Restoration of Archived FilmsabstractA novel segmentation-assisted method for film dirt detection is proposed. Since dirt manifests as a cluster of pixels whose intensity differs from that of its neighbourhood, we employ segmentation and assume that each small region as a dirt candidate. The assumption is validated by considering raw (non-motion compensated) differences between the current frame and each of the previous and next frames which provides a measure of a confidence. Our experiments show that our method compares favourably with standard spatial, temporal and multistage median filtering approaches and provides efficient and robust detection even for fast moving sequences. 1 Jinchang Ren, Theodore Vlachos |
BMVC | 1 |
| 2005 | A User-Oriented Multimodal-Interface Framework for General Content-Based Multimedia RetrievalabstractA user-oriented multimodal interface (MMI) framework is proposed. Considering the complexities of media connotations and uncertainties of the user’s demands, content-based retrieval has intrinsic requirements for MMI for effective media-content interactions. Through integration of knowledge based conduction, learning of semantic concepts, natural language processing and analysis of users’ profiles, our framework can establish a solid basis for design and implementation of general CBR systems satisfying extensibility, condensability and inter-operability. Jinchang Ren, Theodore Vlachos, Vasileios Argyriou |
ICME | 1 |
| 2004 | Real-time 3D Football Ball Tracking from Multiple Cameras
Jinchang Ren, James Orwell, Graeme A. Jones, Ming Xu 0011 |
BMVC | 1 |
| 2004 | A general framework for 3d soccer ball estimation and tracking
Jinchang Ren, James Orwell, Graeme A. Jones, Ming Xu 0011 |
ICIP | 1 |
| 2000 | Multimodal Interface Techniques in Content-Based Multimedia Retrieval
Jinchang Ren, Rongchun Zhao, David Dagan Feng, Wan-Chi Siu |
ICMI | 1 |