EDBT 2026 Demo / reviewers in the wild / expert
Sibao Chen 0001
dblp:85/7027 · also Si-Bao Chen 0001
· DBLP profile ↗
76ranked-venue papers
22as first author
47since 2021 · last 2026
0000-0003-1481-0162ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 15 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LWGANet: Addressing Spatial and Channel Redundancy in Remote Sensing Visual Tasks with Light-Weight Grouped AttentionabstractLight-weight neural networks for remote sensing (RS) visual analysis must overcome two inherent redundancies: spatial redundancy from vast, homogeneous backgrounds, and channel redundancy, where extreme scale variations render a single feature space inefficient. Existing models, often designed for natural images, fail to address this dual challenge in RS scenarios. To bridge this gap, we propose LWGANet, a light-weight backbone engineered for RS-specific properties. LWGANet introduces two core innovations: a Top-K Global Feature Interaction (TGFI) module that mitigates spatial redundancy by focusing computation on salient regions, and a Light-Weight Grouped Attention (LWGA) module that resolves channel redundancy by partitioning channels into specialized, scale-specific pathways. By synergistically resolving these core inefficiencies, LWGANet achieves a superior trade-off between feature representation quality and computational cost. Extensive experiments on twelve diverse datasets across four major RS tasks—scene classification, oriented object detection, semantic segmentation, and change detection—demonstrate that LWGANet consistently outperforms state-of-the-art light-weight backbones in both accuracy and efficiency. Our work establishes a new, robust baseline for efficient visual analysis in RS images. Wei Lu 0032, Sibao Chen 0001 |
AAAI | 3 |
| 2026 | Background noise suppression for advanced fine-grained visual classification
Zhi-Gang Wang, Sibao Chen 0001, Bin Luo 0001 |
Neurocomputing | 2 |
| 2026 | YOFOR : You only focus on object regions for tiny object detection in aerial images
Heng Hu, Hao-Zhe Wang, Sibao Chen 0001, Jin Tang 0001 |
Neural Networks | 3 |
| 2026 | Dynamic adaptive multi-view contrastive learning for unsupervised person re-identification
Zhi-Hua Li, Xue-Yan Wang, Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Neural Networks | 3 |
| 2026 | Multi-scale feature sharing and collaborative sampling for unsupervised vehicle re-identification
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Pattern Recognit. | 2 |
| 2026 | Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector
Qing-Ling Shu, Sibao Chen 0001, Xiao Wang 0014, Zhi-Hui You, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 2 |
| 2026 | CLNS: Camera-aware label noise suppression for unsupervised visible-infrared person re-identification
Sicheng Zhao, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Futian Wang, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 4 |
| 2026 | RCNet: Reliable Co-Training Network for Weakly Supervised Change DetectionabstractFully supervised change detection (CD) methods in remote sensing (RS) perform well but depend on costly and time-consuming pixel-level annotations, which are impractical to obtain at scale. Therefore, it is essential to develop annotation-efficient alternatives that can narrow the performance gap with fully supervised methods. To this end, we propose a novel weakly supervised CD framework, named RCNet, which employs dual networks to implement reliable co-training using image-level annotations. Our framework is grounded in multi-view learning of co-training and the localization ability of class activation mapping (CAM). In our approach, two sub-nets with the same architecture perform image-level change classification and pixel-level segmentation from different views. Although CAM roughly localizes changes, ambiguity and noise in its pseudo labels may cause confirmation bias, limiting performance. Our approach mitigates this bias by introducing a feature discrepancy loss to enable cross-supervision between two sub-nets. Meanwhile, CAM tends to highlight a single object, but RS images commonly contain many dense and small changed objects with complexity, resulting in decreased reliability of pseudo labels. Therefore, we present an IoU-based reliable pseudo label screening (RPLS) strategy, which minimizes the likelihood of changed areas being misidentified as unchanged, enhancing the reliability of changed information obtained. Besides, to further improve boundary fineness and internal integrity of changed areas, we incorporate an additional strong perturbation branch for each sub-net and develop a consistency regularization loss. Extensive experiments on three challenging RS image CD datasets demonstrate that our RCNet achieves competitive performance with image-level labels. The source code is available athttps://github.com/Youzhihui/RCNet. Zhi-Hui You, Sibao Chen 0001, Chris Ding, Lili Huang 0006, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Lightweight oriented object detection with Dynamic Smooth Feature Fusion Network
Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001 |
Neurocomputing | 3 |
| 2025 | Understanding beyond outputs: A novel knowledge distillation method using Schur decomposition
Chang-Ming Pan, Sibao Chen 0001, Bo Jiang 0002, Bin Luo 0001 |
Neurocomputing | 2 |
| 2025 | Multi-view enhanced truck re-identification
Xue-Yan Wang, Run-Sen Xia, Sibao Chen 0001, Jin Tang 0001 |
Knowl. Based Syst. | 4 |
| 2025 | FSENet: Feature suppression and enhancement network for tiny object detection
Heng Hu, Sibao Chen 0001, Zhi-Hui You, Jin Tang 0001 |
Pattern Recognit. | 2 |
| 2025 | Instant pose extraction based on mask transformer for occluded person re-identification
Qing-Ling Shu, Sibao Chen 0001, Lili Huang 0006, Bin Luo 0001 |
Pattern Recognit. | 3 |
| 2025 | Camera-Proxy Enhanced Identity-Recalibration Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person Re-Identification (VI-ReID) involves querying images of the same person across visible and infrared modalities. To minimize annotation costs, Unsupervised Visible-Infrared person Re-Identification (UVI-ReID) using pseudo-label contrastive learning has emerged. Traditional UVI-ReID approaches often neglected camera domain information and relied on inadequate update strategies during training, only using cosine distance for testing, which led to incorrect mapping of cross-modal relationships. To address these issues, we propose Camera-proxy Enhanced Identity-recalibration Learning (CEIL). It consists of two main stages: first, it employs intra-modal contrastive learning in conjunction with the camera-proxy, updates the memory bank using our innovative Difficulty-aware Cluster-based Memory Updating (DCMU) strategy, and applies Camera Domain-driven Local correlation (CDL) Loss to enhance the learning process. Then utilizes cross-modal contrastive learning, featuring our Proxy-enhanced Cross-modal Mapping (PCM) module, to recalibrate the identity relationships between different modalities. Graph network-based Camera constraint adjustment Re-ranking (GCR) method is adopted during test, utilizing camera domain information to recalibrate the correspondence between identities. Extensive experiments have demonstrated that CEIL achieving state-of-the-art performance on the SYSU-MM01, RegDB, and LLCM datasets and the GCR, as a general unsupervised re-ranking method, can further enhance performance of model on these datasets. The code will be released athttps://github.com/maybeextra/CEIL. Run-Sen Xia, Xue-Yan Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | CFENet: Contextual Feature Enhancement Network for Tiny Object Detection in Aerial ImagesabstractWith the development of deep learning techniques and object detectors, the performance of object detection has been rapidly improved. However, since tiny objects contain only a small number of pixels and lack appearance information, this creates difficulties for detector recognition. Although existing research has improved detection performance by fusing different feature layers to enhance feature information of objects, this also leads to the problem of mixed feature information, especially for tiny objects where features are easily covered, which exacerbates the difficulty of recognition. To solve the above problems, we propose a contextual feature enhancement network (CFENet), which is an efficient framework built on anchor-based object detectors. In CFENet, to effectively utilize contextual information around an object to enhance the detection of tiny objects, we use poolFormer to build a backbone to extract object features. To alleviate the feature blending problem caused by feature fusion, we propose a feature suppression module (FSM) that effectively suppresses background information and redundant features to enhance tiny object features. In addition, we utilize the improved Gaussian Wasserstein distance loss to modify the loss function to obtain high-quality bounding boxes, and we further manipulate the shallow feature layer of the output and then add a detection head to enhance the detection of tiny objects. We have conducted extensive experiments on the public datasets AI-TOD, VisDrone, and DOTA to demonstrate the effectiveness of our approach. Heng Hu, Sibao Chen 0001, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Multiscale Adaptive Decoder and Diversity Selection Network for Road Extraction in Remote Sensing ImageabstractRoad extraction has been a common and challenging task in the field of remote sensing images. Due to factors such as the high resolution of remote sensing images and the subtle visibility of road features, existing methods often miss certain areas during detection and extraction. These methods struggle to capture contextual information effectively and tend to exhibit false positives and false negatives when handling objects of varying sizes. This article proposes a network based on a multi-scale adaptive decoder and diverse selection (MADSNet) to address the issue of inadequate contextual information capture. By leveraging feature diverse selection, the method minimizes errors in distinguishing between road features and background interference. Specifically, the multi-scale feature flexible extraction (MFFE) decoder utilizes the relevance inquiry attention (RIA) module and scope flexible fusion (SFF) module to enhance the ability to capture contextual information with relatively low computational demands. The optimal choice graph attention (OCGA) module aggregates neighboring nodes with similar features in a graph structure, improving focus on the single class of roads. Furthermore, a multi-level feature selection (MFS) module is proposed to activate the features relevant to the current stage while suppressing features from other stages and interfering with noise. Quantitative and qualitative experimental results on three public datasets demonstrate that the proposed MADSNet outperforms currently popular methods in terms of performance. The code will be available at https://github.com/Talent02/MADSNet. Zhen-Tao Hua, Sibao Chen 0001, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Multidimensional Remote Sensing Change Detection Based on Siamese Dual-Branch NetworksabstractDeep learning models, particularly convolutional neural networks (CNNs), have demonstrated outstanding feature learning capabilities, leading to remarkable performance in remote sensing change detection (RSCD) tasks. However, their most critical drawback lies in the lack of effective modeling of global information. This deficiency affects the model’s understanding of the overall context and structure of the entire image, making it difficult to distinguish between background and target areas, thereby leading to the erroneous identification of change regions. Second, features extracted by traditional backbone networks contain a significant amount of noise, resulting in blurred boundaries of changed objects. The challenge of effectively fusing detailed and semantic information to accurately differentiate pseudo changes remains significant. Furthermore, how to fully exploit multiscale information is another issue worth considering. We propose a full-scale multidimensional interaction network called SDSN, which enhances feature representation by leveraging both detail and semantic branches. Initially, bi-temporal images are processed by the encoder to extract coarse multiscale features. The semantic branch guides shallow-scale features, while the detail branch focuses on deep-scale features. Multikernel receptive module (MRM) aggregates global information. The detail branch utilizes a diversity variance module (DVM) and differential operations to generate refined change maps with noise reduction and background suppression. A multidimensional cross-perception module (MCM) guides the fusion of these change maps, establishing multidimensional dependencies to enrich feature representation. Compared with previous methods, SDSN demonstrates greater performance under complex environmental conditions, particularly noteworthy for its fewer parameters (4.03 M) and lower computational costs (7.94 G). The code is publicly available athttps://github.com/dpt000121/dpt. Li-Rong Shen, Sibao Chen 0001, Lili Huang 0006, Zhi-Hui You, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Real-World Remote Sensing Image Dehazing: Benchmark and BaselineabstractRemote Sensing Image Dehazing (RSID) poses significant challenges in real-world scenarios due to the complex atmospheric conditions and severe color distortions that degrade image quality. The scarcity of real-world remote sensing hazy image pairs has compelled existing methods to rely primarily on synthetic datasets. However, these methods struggle with real-world applications due to the inherent domain gap between synthetic and real data. To address this, we introduce Real-World Remote Sensing Hazy Image Dataset (RRSHID), the first large-scale dataset featuring real-world hazy and hazy-free image pairs across diverse atmospheric conditions. Based on this, we propose MCAF-Net, a novel framework tailored for real-world RSID. Its effectiveness arises from three innovative components: Multi-branch Feature Integration Block Aggregator (MFIBA), which enables robust feature extraction through cascaded integration blocks and parallel multi-branch processing; Color-Calibrated Self-Supervised Attention Module (CSAM), which mitigates complex color distortions via self-supervised learning and attention-guided refinement; and Multi-Scale Feature Adaptive Fusion Module (MFAFM), which integrates features effectively while preserving local details and global context. Extensive experiments validate that MCAF-Net demonstrates state-of-the-art performance in real-world RSID, while maintaining competitive performance on synthetic datasets. The introduction of RRSHID and MCAF-Net sets new benchmarks for real-world RSID research, advancing practical solutions for this complex task. The code and dataset are publicly available at here. Zeng-Hui Zhu, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | DEGANet: Road Extraction Using Dual-Branch Encoder With Gated Attention MechanismabstractAutomatic identification and extraction of roads from high-resolution remote sensing images (RSIs) are important in remote sensing and computer vision. Advancements in remote sensing technology have increased the information in images, making road extraction more challenging. Conventional convolutional methods have limitations, such as loss of spatial details and inadequate fusion of multiscale features. To address these challenges, the letter introduces a novel encoder-decoder architecture called dual-branch encoder with gated attention mechanism network (DEGANet), for extracting road networks in remote sensing image (RSI). First, we propose a multigated informative self-attention (MGSA) module that combines information from dual-branch encoders. By integrating the ResNet and the dynamic snake convolution (DSC) block, which conforms to road shapes, the module emphasizes slender structures similar to roads, thus enhancing the extraction of road features and focusing on capturing more road details. Second, we also introduce the cascade receptive field enhancement (CRFE) module, which optimizes both accuracy and computational complexity. This module combines various receptive field enhancement modules to improve capture long-range dependencies and spatial information perception. Comprehensive experiments conducted on various public remote sensing road datasets demonstrate that our network attains greater segmentation accuracy (intersection over union (IoU) and$F1$score) and connectivity [average path length similarity (APLS)], validating the effectiveness of our proposed method. Sibao Chen 0001, Lili Huang 0006, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Few-Shot Object Detection in Remote Sensing Images With Multiscale Spatial Selective AttentionabstractFew-shot object detection (FSOD) leverages limited labeled data and substantial unlabeled data for detection. However, these approaches mainly target natural images and ignore the spatial relationships and contextual information between objects in remote sensing images (RSIs). To overcome these challenges, this letter introduces a novel method for detecting few-shot objects in RSI. First, we propose a new attention, called multiscale spatial selective attention (MSSSA). This attention spatially selects feature maps from convolution kernels of different scales through spatial selection, focusing the network on the most relevant region of spatial context. Then, our proposed pixel-level feature extractor module (PLFEM) was used in the first stage of FSOD, providing pixel-level object position information to reduce false and missed detection. To evaluate the proposed method, we carry out comprehensive experiments on the DIOR dataset. The results show that the novel class mAP of our method reaches 38.2% in ten shots, an increase of 3.0% compared with the baseline, significantly improving the accuracy of FSOD in RSI. Yingnan Yu, Sibao Chen 0001, Lili Huang 0006, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | DLAReID: double-layer attention network for object re-identification
Sibao Chen 0001, Bin Luo 0001 |
Multim. Tools Appl. | 2 |
| 2024 | An Oriented Object Detector for Hazy Remote Sensing ImagesabstractCurrently, a lot of work is focused on aerial object detection and has achieved good results. Though these methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. Currently, there are limited approaches that combine aerial object detection with hazy conditions, and there are few publicly available datasets for real hazy weather based on aerial images. For this purpose, we propose a dataset HRSI, hazy remote sensing images in the real world, which is mainly divided into three categories: airport, large vehicle, and ship. All images in HRSI are from real hazy conditions. In addition, we propose an object detection model DFENet, a dehazing feature enhancement model for hazy remote sensing images, which is suitable for hazy weather. DFENet consists of a two-branch and a dehazing module. The two-branch structure helps to fully learn hazy and dehazing features. In order to avoid the impact of noise caused by the dehezing module, we also designed a haze-predict module (HPM) to predict the information containing haze in the image. We introduce the cross-fuse module (CFM) to utilize the information of haze to guide the feature fusion of two branches. By utilizing the information of haze, DFENet can dynamically adjust the feature weight in the two-branch to avoid the impact of noise generated by the dehazing module. Compared with traditional object detection methods, DFENet not only has good performance in hazy conditions but also improves performance in clear conditions. We tested DFENet on DOTA, HRSI, and Foggy-DOTA to demonstrate that DFENet performs better under hazy conditions. Sibao Chen 0001, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | DecoupleNet: A Lightweight Backbone Network With Efficient Feature Decoupling for Remote Sensing Visual TasksabstractIn the realm of computer vision (CV), balancing speed and accuracy remains a significant challenge. Recent efforts have focused on developing lightweight networks that optimize computational efficiency and feature extraction. However, in remote sensing (RS) imagery, where small and multiscale object detection is critical, these networks often fall short in performance. To address these challenges, DecoupleNet is proposed, an innovative lightweight backbone network specifically designed for RS visual tasks in resource-constrained environments. DecoupleNet incorporates two key modules: the feature integration downsampling (FID) module and the multibranch feature decoupling (MBFD) module. The FID module preserves small object features during downsampling, while the MBFD module enhances small and multiscale object feature representation through a novel decoupling approach. Comprehensive evaluations on three RS visual tasks demonstrate DecoupleNet’s superior balance of accuracy and computational efficiency compared to existing lightweight networks. On the NWPU-RESISC45 classification dataset, DecoupleNet achieves a top-1 accuracy of 95.30%, surpassing FasterNet by 2%, with fewer parameters and lower computational overhead. In object detection tasks using the DOTA 1.0 test set, DecoupleNet records an accuracy of 78.04%, outperforming ARC-R50 by 0.69%. For semantic segmentation on the LoveDA test set, DecoupleNet achieves 53.1% accuracy, surpassing UnetFormer by 0.70%. These findings open new avenues for advancing RS image analysis on resource-constrained devices, addressing a pivotal gap in the field. The code and pretrained models are publicly available athttps://github.com/lwCVer/DecoupleNet. Wei Lu 0032, Sibao Chen 0001, Qing-Ling Shu, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Prior Guidance and Principal Attention Network for Remote Sensing Image Change DetectionabstractIn the field of remote sensing (RS) image change detection (CD), the conventional encoder-decoder architecture networks often encounter three significant challenges. First, noise in the features extracted from traditional backbone networks leads to blurred boundaries of change objects. Second, upsampling techniques employed in the decoder, such as interpolation or deconvolution, are limited by their finite receptive fields, making it challenging to accurately distinguish pseudo-changes. Furthermore, how to merge encoder and decoder features with possible semantic gaps for the fine-grained details is a topic worth considering. To address these challenges, we introduce a prior guidance (PG) module that effectively aggregates prior high-level features as a semantic guidance map to guide encoder features for the enhancement of boundary detection. In addition, we design a principal attention (PA) module, which aggregates global information from principal regions through sparse operations and adaptively allocates this information to the upsampled and encoder features. This not only addresses the deficiency of global information in the upsampled features but also reduces the semantic gap between the encoder and decoder by establishing channel dependencies. PA does not divert attention to irrelevant regions, demonstrating excellent performance and computational efficiency. By integrating these two modules into our method, a novel PG and PA network (PGPANet) is elaborately designed. A wide range of experiments confirms the validity of our method, showcasing outstanding detection accuracy on three publicly available CD datasets: LEVIR-CD, SYSU-CD, and WHU-CD. The demo code of this work is publicly available athttps://github.com/DaGuangDaGuang/PGPANet. Qing-Ling Shu, Sibao Chen 0001, Zhi-Hui You, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Diffusion Models and Pseudo-Change: A Transfer Learning-Based Change Detection in Remote Sensing ImagesabstractRemote sensing (RS) image change detection (CD) has been a research hotspot in recent years, which plays an important role in urban planning and disaster assessment. However, since CD labels are difficult to obtain, how to utilize semantic information in RS images to improve the change prediction performance is a problem worth exploring. To solve this problem, we propose a transfer learning-based CD method that utilizes a diffusion generation model to translate high-level semantic information into low-level change information. First, we propose a pseudo-change image pair generation method that utilizes semantic labels to guide the diffusion model to generate change images. Then, the refined loss (RL) is designed to improve the model’s ability to recognize change features based on the difference between pseudo-change image pairs and unlabeled image pairs. Experimental results on WHU-CD, LEVIR-CD, and GoogleGZ-CD datasets show that the proposed method effectively transfers the semantic information into change information and finally improves the model’s feature recognition ability for change objects. Compared with recent CD and transfer learning methods, the proposed transfer learning model (TLM) achieves the best performance. The source code is available athttps://github.com/VCISwang/STCD. Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Cheng-Jie Gu, Zhi-Hui You, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Attention-Aware Sobel Graph Convolutional Network for Remote Sensing Image Change DetectionabstractIn the study of remote sensing images, the problem of change detection (CD) is crucial. Convolutional neural networks (CNNs) are well-liked feature extraction structures that are frequently used in CD. On the other hand, graph convolutional networks (GCNs) are effective in building contextual structure information. Compared with CNN, GCN can make full use of the graph structure information to capture the changing features between different areas in the graph by learning the connections and interactions between nodes. In contrast, traditional pixel-based CNNs may have difficulty modeling semantic relationships and temporal variations among features and are susceptible to noise interference. So in this article, we extract optimization information using a GCN structure. Due to the particularity of remote sensing images, edge information is often ignored, which is useful in the field of CD. In this article, we propose an attention-aware Sobel GCN (ASGCN) for remote sensing image CD. First, we use a Siamese CNN to extract primary multilevel features. Then, a dual-branch attention module (DAM) including coordinate attention and multiscale local attention module (MLAM) is proposed to focus on informative pixels, we use Sobel operator to construct graph, and the graph convolutional module can expand receptive field and extract edge information. Attention fusion module (AFM) is adopted at decoder to perform effective feature fusion. Extensive comparative experiments on three CD datasets, LEVIR-CD, WHU-CD, and DSIFN-CD, verify the effectiveness of the proposed ASGCN. Lei Wang 0095, Zhi-Hui You, Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Prototype Discriminative Learning for Semi-Supervised Change Detection in Remote Sensing ImagesabstractWith the continuous progress of deep learning in remote sensing (RS) visual tasks, considerable advancements have been achieved in RS image change detection (CD). However, prevailing CD methods heavily rely on extensive sets of fully pixelwise hand-annotated training data, a time-consuming and costly process, and they fail to fully harness the potential benefits of deep feature representations within the deep feature domain. To tackle the mentioned issues, we propose a novel semi-supervised CD method called PDLCD, which strategically leverages useful information from massive unlabeled data to complement labeled data with just a few samples. Specifically, changed objects and unchanged backgrounds of bitemporal RS images are various and complex, our approach advocates dividing each category into multiple subclasses in the deep feature domain. In this scheme, the high-level feature of each subclass follows a Gaussian distribution. Then, the prototype discriminative learning (PDL) is introduced to explicitly encourage deep features of samples closer to the nearest prototype within their respective category, and away from all prototypes of other categories. We design feature discriminative loss (FDL) to implement PDL for constructing more pronounced intraclass compactness and interclass variability. Finally, we compute the supervised loss based on a limited set of labeled data, incorporate the unsupervised loss leveraging a substantial volume of unlabeled data, and include FDL within the deep feature domain to collectively optimize the model. Extensive experiments carried out on three challenging RS image CD datasets illustrate that our proposed semi-supervised CD method obtains better CD performance than previous counterparts. The source code is available at:https://github.com/Youzhihui/PDLCD. Zhi-Hui You, Sibao Chen 0001, Jia-Xin Wang, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | GENet: Guidance Enhancement Network for 3D Shape RecognitionabstractBoth point cloud-based and view-based deep learning methods for 3D shape recognition have achieved relatively remarkable results in recent years. However, there are few methods to jointly represent 3D shapes from both point cloud and multi-view modal data. Therefore, we propose a guidance enhancement network (GENet) for 3D shape recognition based on multimodal data. On the one hand, the point cloud is encoded with features from both explicit and implicit aspects, and on the other hand, all views are encoded and constructed as a graph. In the multilayer guidance enhancement module, graph convolutional neural network (GCN) enhances each view feature, and then temporary high-level features (initially point cloud global feature) guide multiple low-level view features to obtain correlation coefficients, through which the views with higher importance are filtered as inputs for the next layer of the structure and the view features in the current layer are weighted and aggregated. The aggregated view features are then connected to the high-level features with residuals to form the enhanced high-level features. The 3D shape descriptor is finally obtained after several guidance and enhancements. The proposed GENet achieves state-of-the-art results on the 3D benchmark dataset ModelNet. Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang |
IJCNN | 7 |
| 2023 | GLCNet: Global-Local Complementary Network for 3D Shape RecognitionabstractBoth point cloud-based and multi-view-based methods have achieved remarkable results in 3D shape recognition, yet there are few methods that combine the two types of data. In this paper, a novel Global-Local Complementary Network (GLCNet) based on multimodal data is proposed. The network obtains more powerful shape descriptors by stacking multiple layers of Global-Local Complementary Module (GLC Module). More specifically, the Global-Local Relation Score Module is first used to obtain the relationship between view features and global feature. The relationship is then utilized to facilitate the aggregation of view features and to filter out the more important ones. Finally, the aggregated view features are fused with the global features to form a stronger global feature. GLCNet enables the characteristics of various data to be fully utilized and achieves a true sense of complementarity of strengths and weaknesses. Extensive experiments on the benchmark dataset ModelNet show that GLCNet achieves state-of-the-art results in 3D shape classification and retrieval. Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang |
IJCNN | 7 |
| 2023 | Improving multi-label learning by modeling Local label and feature correlationsabstractMulti-label learning deals with the problem that each instance is associated with multiple labels simultaneously, and many methods have been proposed by modeling label correlations in a global way to improve the performance of multi-label learning. However, the local label correlations and the influence of feature correlations are not fully exploited for multi-label learning. In real applications, different examples may share different label correlations, and similarly, different feature correlations are also shared by different data subsets. In this paper, a method is proposed for multi-label learning by modeling local label correlations and local feature correlations. Specifically, the data set is first divided into several subsets by a clustering method. Then, the local label and feature correlations, and the multi-label classifiers are modeled based on each data subset respectively. In addition, a novel regularization is proposed to model the consistency between classifiers corresponding to different data subsets. Experimental results on twelve real-word multi-label data sets demonstrate the effectiveness of the proposed method. Qianqian Cheng, Jun Huang 0003, Huiyi Zhang, Sibao Chen 0001 |
Intell. Data Anal. | 4 |
| 2023 | A super-resolution-based license plate recognition method for remote surveillance
Sen Pan, Sibao Chen 0001, Bin Luo 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Road Extraction by Multiscale Deformable Transformer From Remote Sensing ImagesabstractRapid progress has been made in the research of high-resolution remote sensing road extraction tasks in the past years, but due to the diversity of road types and the complexity of road context, extracting the perfect road network is still fraught with difficulties and challenges. Many Convolutional Neural Networks (CNNs) based on encoder-decoder structures have demonstrated their effectiveness. Transformer’s self-attention mechanism shows more powerful performance than CNNs in modeling global feature dependencies. In this paper, we propose a Multi-scale Deformable Transformer Network (MDTNet) based on encoder-decoder structure to extract road networks from remote sensing images. The core of MDTNet is our proposed Multi-scale Deformable Self-Attention (MDSA) mechanism. MDSA can capture more comprehensive features than conventional self-attention. In addition, roads are not present in certain blocks of areas like other objects, but are interwoven throughout the image in such a long, linear fashion that information about certain road segments may be overlooked. To minimize residual errors in road segmentations, our MDSA incorporates a deformable design on feature maps, which effectively enhances the salience of road features relative to their surroundings. Extensive experiments on several public remote sensing road datasets show that our MDTNet achieves higher segmentation [F1 score and Intersection over Union (IoU)] and connectivity [Average Path Length Similarity (APLS)] accuracy, which verifies the effectiveness of our approach. Pengcheng Hu 0001, Sibao Chen 0001, Lili Huang 0006, Guizhou Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Combinatorial online high-order interactive feature selection based on dynamic graph convolution network
Wen-Bin Wu, Jun-Jun Sun, Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Signal Process. | 3 |
| 2023 | Channel Attention TextCNN with Feature Word Extraction for Chinese Sentiment AnalysisabstractChinese short text sentiment analysis can help understand society’s views on various hot topics. Many existing sentiment analysis methods are based on sentiment dictionaries. Still, sentiment dictionaries are easily affected by subjective factors. They require a lot of time to build as well as maintenance to prevent obsolescence. For the aim of extracting rich information within texts more effectively, we propose a Channel Attention TextCNN with Feature Word Extraction model (CAT-FWE). The feature word extraction module helps us choose words that affect the sentiment of reviews. Then, these words are integrated with multi-level semantic information to enhance the information of sentences. In addition, the channel attention textCNN module that is a promotion of traditional TextCNN tends to pay more attention to those meaningful features. It eliminates the impacts of features that do not make any sense effectively. We apply our CAT-FWE model to both fine-grained classification and binary classification tasks for Chinese short texts. Experiment results show that it can improve the performance of emotion recognition. Jiangwei Liu, Zian Yan, Sibao Chen 0001, Xiao Sun 0003, Bin Luo 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | A Robust Feature Downsampling Module for Remote-Sensing Visual TasksabstractRemote sensing (RS) images present unique challenges for computer vision due to lower resolution, smaller objects, and fewer features. Mainstream backbone networks show promising results for traditional visual tasks. However, they use convolution to reduce feature map dimensionality, which can result in information loss for small objects in RS images and decreased performance. To address this problem, we propose a new and universal downsampling module named Robust Feature Downsampling (RFD). RFD fuses multiple feature maps extracted by different downsampling techniques, creating a more robust feature map with a complementary set of features. Leveraging this, we overcome the limitations of conventional convolutional downsampling, resulting in more accurate and robust analysis of RS images. We develop two versions of RFD module, Shallow RFD (SRFD) and Deep RFD (DRFD), tailored to adapt to different stages of feature capture and improve feature robustness. We replace the downsampling layers of existing mainstream backbones with RFD module and conduct comparative experiments on several public RS image datasets. The results show significant improvements compared to baseline approaches in RS image classification, object detection, and semantic segmentation. Specifically, our RFD module achieved an average performance gain of 1.5% on NWPU-RESISC45 classification dataset without utilizing any additional pretraining data, resulting in state-of-the-art performance on this dataset. Moreover, in detection and segmentation tasks on DOTA and iSAID datasets, our RFD module outperforms the baseline approaches by 2-7% when utilizing pretraining data from NWPU-RESISC45. These results highlight the value of RFD module in enhancing the performance of RS visual tasks. Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Chris Ding, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Crossed Siamese Vision Graph Neural Network for Remote-Sensing Image Change DetectionabstractThe development of deep learning in remote sensing (RS) visual tasks has led to remarkable progress in RS image change detection (CD). However, RS bi-temporal images cover complex and confusing scenes due to natural environmental factors, which presents challenges for CD task. How to effectively exploit long-range dependencies and sensitively discriminate real-changes with various scales from pseudo-changes are urgent problems. It is especially obvious for the changes of building structures man-made. This paper presents a CD approach named CSViG, which utilizes Siamese Vision Graph neural network (SViG) with crossed feature fusion. SViG acts as a feature extractor to capture richer short- and long-range dependencies. Crossed feature fusion consists of a horizontal feature fusion module (HFFM) and a vertical feature fusion module (VFFM). HFFM designs cross-concatenation (CC) way to reveal real-changes from pseudo-change in the same horizontal stage, after which global and local features are extracted by using attention mechanism and multi-scale depth-wise separable convolution. VFFM further fuses complementary content from vertical multiple stages to effectively represent change regions of different sizes (tiny or huge) by using attention mechanism. Extensive comparative experiments conducted on three available building change detection datasets demonstrate that the proposed method achieves better CD performance than previous counterparts. Zhi-Hui You, Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Guizhou Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Model Compression Based on Differentiable Network Channel PruningabstractAlthough neural networks have achieved great success in various fields, applications on mobile devices are limited by the computational and storage costs required for large models. The model compression (neural network pruning) technology can significantly reduce network parameters and improve computational efficiency. In this article, we propose a differentiable network channel pruning (DNCP) method for model compression. Unlike existing methods that require sampling and evaluation of a large number of substructures, our method can efficiently search for optimal substructure that meets resource constraints (e.g., FLOPs) through gradient descent. Specifically, we assign a learnable probability to each possible number of channels in each layer of the network, relax the selection of a particular number of channels to a softmax over all possible numbers of channels, and optimize the learnable probability in an end-to-end manner through gradient descent. After the network parameters are optimized, we prune the network according to the learnable probability to obtain the optimal substructure. To demonstrate the effectiveness and efficiency of DNCP, experiments are conducted with ResNet and MobileNet V2 on CIFAR, Tiny ImageNet, and ImageNet datasets. Yu-Jie Zheng, Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | SIECP: Neural Network Channel Pruning based on Sequential Interval Estimation
Sibao Chen 0001, Yu-Jie Zheng, Chris Ding, Bin Luo 0001 |
Neurocomputing | 1 |
| 2022 | DBRANet: Road Extraction by Dual-Branch Encoder and Regional Attention DecoderabstractAlthough widely exploited in recent decades, road extraction is still a very significant and challenging research in the field of remote sensing image processing due to the complex background and road distribution. Among the existing CNN-based methods, U-shape architectures composed of encoders and decoders have shown their effectiveness. In this letter, we propose an improved encoder–decoder method, named DBRANet, for extracting roads from remote sensing images. In the encoding phase, we present a dual-branch network module (DBNM) to construct more effective features, thus improving the fusion feature maps of different scales. One branch utilizes the residual block, and the other branch utilizes the refined asymmetric block, which effectively increases the feature extraction capability of the backbone. In the decoding phase, considering the sinuous shape and the unbalanced distribution of roads in remote sensing images, we design a novel attention module, named the regional attention network module (RANM), to automatically learn the importance of each channel according to the regional information. Extensive experiments on several public remote sensing road data sets show that our DBRANet achieves higher segmentation [$F1$score and Intersection over Union (IoU)] and connectivity [average path length similarity (APLS)] accuracy, which verifies the effectiveness of our approach. Sibao Chen 0001, Yu-Xin Ji, Jin Tang 0001, Bin Luo 0001, Weiqiang Wang 0001, Ke Lu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | BDTNet: Road Extraction by Bi-Direction Transformer From Remote Sensing ImagesabstractThe past several years have witnessed the rapid development of the task of road extraction in high-resolution remote sensing images. However, due to the complex background and road distribution, road extraction is still a challenging research in remote sensing images. In convolutional neural networks (CNNs), the U-shaped architecture network has shown its effectiveness. But the global representation cannot be captured effectively by CNNs. While in the transformer, the self-attention (SA) module can capture the long-distance feature dependencies. A hybrid encoder-decoder method called BDTNet is proposed in this letter, which enhance the extraction of global and local information in remote sensing images. Firstly, feature maps of different scales are obtained through the backbone network. And then, on the basis of reducing the computational cost of self-attention, the Bi-Direction Transformer Module (BDTM) is constructed to capture the contextual road information in feature maps of different scales. Finally, the Feature Refinement Module (FRM) is introduced to integrate the features extracted from the backbone network and BDTM, which enhances the semantic information of the feature maps and obtains more detailed segmentation results. The results show that the proposed method achieved a high IoU of 67.09% in the DeepGlobe dataset. Extensive experiments also verify the effectiveness of the proposed method on three public remote sensing road datasets. Jia-Xin Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Semi-Supervised Semantic Segmentation of Remote Sensing Images With Iterative Contrastive NetworkabstractWith the development of deep learning, semantic segmentation of remote sensing images has made great progress. However, segmentation algorithms based on deep learning usually require a huge number of labeled images for model training. For remote sensing images, pixel-level annotation usually consumes expensive resources. To alleviate this problem, this letter proposes a semi-supervised segmentation method of remote sensing images based on an iterative contrastive network. This method combines few labeled images and more unlabeled images to significantly improve the model performance. First, contrastive networks continuously learn more potential information by using better pseudo labels. Then, the iterative training method keeps the differences between models to better improve the segmentation performance. The semi-supervised experiments on different remote sensing datasets prove that this method has a better performance than the related methods. Code is available athttps://github.com/VCISwang/ICNet. Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | RanPaste: Paste Consistency and Pseudo Label for Semisupervised Remote Sensing Image Semantic SegmentationabstractWith the development of deep learning, remote sensing (RS) image segmentation has been applied with marked success. However, in the process of model training, the large number of labeled images required more expensive annotation. A key challenge is how to make full use of extensive unlabeled images available to improve the segmentation model. In this article, we propose a semisupervised remote sensing image semantic segmentation method defined as RanPaste, which combines labeled images with unlabeled images to improve segmentation performance. First, we obtain pseudo label by randomly pasting part of the ground truth label into the predicted segmentation map. Then, we combine the labeled and unlabeled images to generate rough predictions after strong augmentation. Finally, by using the semisupervised loss, we achieve better performance on remote sensing image segmentation. Our method combines consistency regularization and pseudo label and then utilizes thresholds to gradually improve the model performance. RanPaste enables the model to learn more underlying information in the unlabeled data. Experimental results on six datasets show that RanPaste can learn more latent information from unlabeled data to improve segmentation performance. Besides, our approach achieves better segmentation results on different network structures and datasets. Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Reliable Contrastive Learning for Semi-Supervised Change Detection in Remote Sensing ImagesabstractWith the development of deep learning in remote sensing (RS) image change detection (CD), the dependence of CD models on labeled data has become an important problem. To make better use of the comparatively resource-saving unlabeled data, the CD method based on semi-supervised learning (SSL) is worth further study. This article proposes a reliable contrastive learning (RCL) method for semi-supervised RS image CD. First, according to the task characteristics of CD, we design the contrastive loss based on the changed areas to enhance the model’s feature extraction ability for changed objects. Then, to improve the quality of pseudo labels in SSL, we use the uncertainty of unlabeled data to select reliable pseudo labels for model training. Combining these methods, semi-supervised CD models can make full use of unlabeled data. Extensive experiments on three widely used CD datasets demonstrate the effectiveness of the proposed method. The results show that our semi-supervised approach has a better performance than related methods. The code is available athttps://github.com/VCISwang/RC-Change-Detection. Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001, Richard C. Wilson 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Remote Sensing Scene Classification via Multi-Branch Local Attention NetworkabstractRemote sensing scene classification (RSSC) is a hotspot and play very important role in the field of remote sensing image interpretation in recent years. With the recent development of the convolutional neural networks, a significant breakthrough has been made in the classification of remote sensing scenes. Many objects form complex and diverse scenes through spatial combination and association, which makes it difficult to classify remote sensing image scenes. The problem of insufficient differentiation of feature representations extracted by Convolutional Neural Networks (CNNs) still exists, which is mainly due to the characteristics of similarity for inter-class images and diversity for intra-class images. In this paper, we propose a remote sensing image scene classification method via Multi-Branch Local Attention Network (MBLANet), where Convolutional Local Attention Module (CLAM) is embedded into all down-sampling blocks and residual blocks of ResNet backbone. CLAM contains two submodules, Convolutional Channel Attention Module (CCAM) and Local Spatial Attention Module (LSAM). The two submodules are placed in parallel to obtain both channel and spatial attentions, which helps to emphasize the main target in the complex background and improve the ability of feature representation. Extensive experiments on three benchmark datasets show that our method is better than state-of-the-art methods. Sibao Chen 0001, Qing-Song Wei, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001, Zuyuan Wang |
IEEE Trans. Image Process. | 1 |
| 2021 | Intelligent Machine Learning System for Predicting Customer ChurnabstractNowadays, customer churn issue is becoming more and more important, which is the key indicator of the business and production success. But how to predict the actual customer churn and take action before customer loss is becoming a difficult issue in the industry. At the same time, how to keep the place of production is the first problem we are facing. After the deep research, we use Artificial Intelligence (AI) and Machine Learning (ML) technology to develop a smart intelligent system and reduce the actual customer churn about the production. This paper will explain the machine learning technology which used in this smart intelligent system and the reader will learn how to use this system to reduce customer loss. In the customer’s churn prediction model aspect, the most popular predictive models have been used, namely, support vector machines, random forests, K-nearest neighbors, and Gradient boosting classifier are applied to check the effect on accuracy, AUC, and F1-score. Through the experiment, it proofs that the Gradient boosting classifier and Random forests give the highest accuracy of 95.32% and 94.29% respectively. The highest AUC score of 91% which achieved by both Gradient boosting classifier and random forests. The highest F1-score of 97.3% is achieved by the Gradient boosting classifier which outperforms over others. Chenggang He, Chris Ding, Sibao Chen 0001, Bin Luo 0001 |
ICTAI | 3 |
| 2021 | Regularization graph convolutional networks with data augmentation
Xiu-Zhi Tian, Chris Ding, Sibao Chen 0001, Bin Luo 0001, Xin Wang 0013 |
Neurocomputing | 3 |
| 2021 | Multi-ColorGAN: Few-shot vehicle recoloring via memory-augmented networks
Wei Zhou 0102, Sibao Chen 0001, Li-Xiang Xu, Bin Luo 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Lsragan: Generating Multifarious Color Photographes From SketchabstractImage translation is to translate source domain image into target domain image. Generative Adversarial Networks (GANs) have achieved appealing performances on one-to-one image translation. The task of sketch translating to real-world images is a one-to-many task. In this paper, a new translation network, named Latent Space Regularization Attention GAN (LSRAGAN), is proposed to generate multifarious color photographes from sketch. We add Gaussian prior and latent variable from target domain to generator to realize multi-modal mapping. Attention weighting is implemented on ResNet blocks in generator. L1-constrained multi-layer perceptual loss is incorporated in conditional GAN. In addition, we present a latent space regularization (LSR) loss to force generator pay attention to the influence of latent code vector, which makes the generated images more diverse. Experiments demonstrate that our method outperforms state-of-the-arts in both quantitative and visual performances. Ke Zhang 0034, Sibao Chen 0001 |
ICIP | 4 |
| 2020 | Global-Local Attention Network for Semantic Segmentation in Aerial ImagesabstractErrors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines. Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001 |
ICPR | 9 |
| 2020 | UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution ImagesabstractSemantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1]. Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001 |
ICPR | 7 |
| 2020 | LSAM: Local Spatial Attention Module
Miao-Miao Lv, Sibao Chen 0001, Bin Luo 0001 |
PRCV (3) | 2 |
| 2020 | Large-Scale Network Representation Learning Based on Improved Louvain Algorithm and Deep Autoencoder
Shou-Jiu Xiong, Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
PRCV (3) | 2 |
| 2019 | Multi-view Similarity Learning of Manifold Data
Rui-Rui Wang, Sibao Chen 0001, Bin Luo 0001, Justin Jian Zhang |
ICIG (1) | 2 |
| 2019 | Pyramid Attention Dense Network for Image Super-ResolutionabstractRecent deep convolution neural networks has made remarkable progress in single images super-resolution area. They achieved very high Peak Signal to Noise Ratio (PSNR) and structural similarity (SSIM), by improved learning of high-frequency details to enhance visual perception. However, current models usually ignore relations between adjacent pixels. In this work, we propose a network that incorporate gradients of adjacent pixels in addition to per-pixel loss and perceptual loss. In addition, we utilize multi-stage network learning to progressively generate high resolution images, by incorporate a new inter-stage feedback in the Laplacian pyramid network structure. Furthermore, we adopted recently proposed attention mechanism and dense block structure. The proposed Pyramid Attention Dense model for image super-resolution achieved state-of-the-art performance in experiments on four benchmark datasets. Sibao Chen 0001, Bin Luo 0001, Chris Ding, Shilei Huang |
IJCNN | 1 |
| 2019 | Multiple Back Propagation Network and Metric Fusion for Person Re-identificationabstractPerson re-identification (Re-ID) is a research focus in pattern recognition, which is to identify a person from another camera view. Many researches have studied feature representations and metric distances of person images, which are robust to changes of view angle and illumination. In this paper, we propose a Multiple Back Propagation (MBP) network and Metric Fusion (MF) for person Re-ID. The proposed MBP network is based on DenseNet or ResNet. Each Dense-conv layer or Conv-ID block is linked by a MBP layer. Each MBP layer is divided into two sub-streams. One sub-stream is connected to softmax loss and the other sub-stream is transferred to a convolution layer followed by triplet loss. A Metric Fusion (MF) method with an optimized weighting scheme is proposed for deep feature fusion. Furthermore, we propose a new metric Re-ranking Euclidean distance joining metric fusion. Experiments on three large-scale person Re-ID benchmark datasets, including Market1501, CUHK03 and DukeMTMC-reID, show that the proposed MBPMF method can achieve state-of-the-art performances. Sibao Chen 0001, Bin Luo 0001, Chris Ding |
IJCNN | 1 |
| 2019 | SRAGAN: Generating Colour Landscape Photograph from SketchabstractGenerating sketch from colour landscape photograph is very easy while it is hard to generate colour photograph from landscape sketch. In this paper, a new automatic conversion network, named Sparse Residual Attention Generative Adversarial Networks (SRAGAN), is proposed to generate landscape colour photograph from sketch. Besides of generator adversarial loss, we not only adopt L1-regularized per-pix loss, but also combine L1-regularized perceptual loss together into our model. Due to the sparsity of L1-norm, it can preserve boundary edge information very well, which makes our model can handle well the conversion task of sketch-to-photo. In addition, we proposed a ResAttention block to our network structure, which combines the residual learning blocks with attention module. Experiments show that the landscape colour photographes generated by our SRAGAN looks more natural with bright colour and clear edge information. At the same time, we integrate two models so that we can generate winter-style and summer-style photographes from the same landscape sketch. Experiments demonstrate that our method outperforms many state-of-the-arts both in quantitative and in visual performance. Sibao Chen 0001, Bin Luo 0001, Chris Ding, Justin Jian Zhang |
IJCNN | 1 |
| 2019 | Extended adaptive Lasso for multi-class and multi-label feature selection
Sibao Chen 0001, Yu-Mei Zhang, Chris Ding, Justin Jian Zhang, Bin Luo 0001 |
Knowl. Based Syst. | 1 |
| 2019 | Feature selection based on correlation deflation
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Neural Comput. Appl. | 1 |
| 2018 | A discriminative multi-class feature selection method via weighted l2, 1-norm and Extended Elastic Net
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Neurocomputing | 1 |
| 2018 | Linear regression based projections for dimensionality reduction
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Inf. Sci. | 1 |
| 2018 | A Nonnegative Locally Linear KNN model for image recognition
Sibao Chen 0001, Yu-Lan Xu, Chris Ding, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2018 | Non-greedy Max-min Large Margin based on L1-norm
Sibao Chen 0001, Chong Zuo, Chris Ding, Bin Luo 0001 |
Pattern Recognit. Lett. | 1 |
| 2017 | Two-Dimensional Discriminant Locality Preserving Projection Based on ℓ1-norm Maximization
Sibao Chen 0001, Cai-Yin Liu, Bin Luo 0001 |
Pattern Recognit. Lett. | 1 |
| 2015 | An algorithm framework of sparse minimization for positive definite quadratic forms
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
Neurocomputing | 1 |
| 2015 | Similarity Learning of Manifold DataabstractWithout constructing adjacency graph for neighborhood, we propose a method to learn similarity among sample points of manifold in Laplacian embedding (LE) based on adding constraints of linear reconstruction and least absolute shrinkage and selection operator type minimization. Two algorithms and corresponding analyses are presented to learn similarity for mix-signed and nonnegative data respectively. The similarity learning method is further extended to kernel spaces. The experiments on both synthetic and real world benchmark data sets demonstrate that the proposed LE with new similarity has better visualization and achieves higher accuracy in classification. Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Extended linear regression for undersampled face recognition
Sibao Chen 0001, Chris Ding, Bin Luo 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Uncorrelated LassoabstractLasso-type variable selection has increasingly expanded its machine learning applications. In this paper, uncorrelated Lasso is proposed for variable selection, where variable de-correlation is considered simultaneously with variable selection, so that selected variables are uncorrelated as much as possible. An effective iterative algorithm, with the proof of convergence, is presented to solve the sparse optimization problem. Experiments on benchmark data sets show that the proposed method has better classification performance than many state-of-the-art variable selection methods. Sibao Chen 0001, Chris Ding, Bin Luo 0001, Ying Xie 0002 |
AAAI | 1 |
| 2010 | On Dynamic Weighting of Data in Clustering with K-Alpha MeansabstractAlthough many methods of refining initialization have appeared, the sensitivity of K-Means to initial centers is still an obstacle in applications. In this paper, we investigate a new class of clustering algorithm, K-Alpha Means (KAM), which is insensitive to the initial centers. With K-Harmonic Means as a special case, KAM dynamically weights data points during iteratively updating centers, which deemphasizes data points that are close to centers while emphasizes data points that are not close to any centers. Through replacing minimum operator in K-Means by alpha-mean operator, KAM significantly improves the clustering performances. Sibao Chen 0001, Haixian Wang, Bin Luo 0001 |
ICPR | 1 |
| 2008 | Heteroscedastic discriminant analysis with two-dimensional constraintsabstractHeteroscedastic discriminant analysis (HDA) with two-dimensional (2D) constraints is proposed in this paper. HDA suffers from the small sample size problem and instability when lack of training data or feature dimension is high, even when the number of dimension is in a suitable range. Two-dimensional HDA is first proposed, then we show that 2D methods are actually a kind of structure-constrained 1D methods, and lastly, HDA with 2D constraints is proposed. Experiments on TIMIT and WSJ0 show that the proposed method outperforms other methods. Sibao Chen 0001, Yu Hu 0003, Bin Luo 0001, Renhua Wang |
ICASSP | 1 |
| 2008 | Probabilistic two-dimensional principal component analysis and its mixture model for face recognition
Haixian Wang, Sibao Chen 0001, Zilan Hu, Bin Luo 0001 |
Neural Comput. Appl. | 2 |
| 2008 | Locality-Preserved Maximum Information ProjectionabstractDimensionality reduction is usually involved in the domains of artificial intelligence and machine learning. Linear projection of features is of particular interest for dimensionality reduction since it is simple to calculate and analytically analyze. In this paper, we propose an essentially linear projection technique, called locality-preserved maximum information projection (LPMIP), to identify the underlying manifold structure of a data set. LPMIP considers both the within-locality and the between-locality in the processing of manifold learning. Equivalently, the goal of LPMIP is to preserve the local structure while maximize the out-of-locality (global) information of the samples simultaneously. Different from principal component analysis (PCA) that aims to preserve the global information and locality-preserving projections (LPPs) that is in favor of preserving the local structure of the data set, LPMIP seeks a tradeoff between the global and local structures, which is adjusted by a parameter alpha, so as to find a subspace that detects the intrinsic manifold structure for classification tasks. Computationally, by constructing the adjacency matrix, LPMIP is formulated as an eigenvalue problem. LPMIP yields orthogonal basis functions, and completely avoids the singularity problem as it exists in LPP. Further, we develop an efficient and stable LPMIP/QR algorithm for implementing LPMIP, especially, on high-dimensional data set. Theoretical analysis shows that conventional linear projection methods such as (weighted) PCA, maximum margin criterion (MMC), linear discriminant analysis (LDA), and LPP could be derived from the LPMIP framework by setting different graph models and constraints. Extensive experiments on face, digit, and facial expression recognition show the effectiveness of the proposed LPMIP method. Haixian Wang, Sibao Chen 0001, Zilan Hu, Wenming Zheng |
IEEE Trans. Neural Networks | 2 |
| 2007 | Local and Weighted Maximum Margin Discriminant AnalysisabstractIn this paper, we propose a new approach, called local and weighted maximum margin discriminant analysis (LWMMDA), to performing object discrimination. LWMMDA is a subspace learning method that identifies the underlying nonlinear manifold for discrimination. The goal of LWMMDA is to seek a transformation such that data points of different classes are projected as far as possible while points within a same class are as compact as possible. The projections are obtained by maximizing a new discriminant criterion, called local and weighted maximum margin criterion (LWMMC). Different from previous maximum margin criterion (MMC) which seeks only the globally Euclidean structure of data points, LWMMC takes the local property into account, which makes LWMMC more accurate in finding discriminant information. LWMMC has an additional weighted parameter β that further broadens the average margin between different classes. Computationally, LWMMDA completely avoids the singularity problem. Besides, LWMMDA couples the QR-decomposition into its framework, which makes LWMMDA very efficient and stable in implementation. Finally, LWMMDA framework is straightforwardly extended into the reproducing kernel Hilbert space induced by a nonlinear function ϕ. Experiments on digit visualization, face recognition, and facial expression recognition are presented to show the effectiveness of the proposed method. Haixian Wang, Wenming Zheng, Zilan Hu, Sibao Chen 0001 |
CVPR | 4 |
| 2007 | Bilateral Two-Dimensional Locality Preserving ProjectionsabstractIn this paper, we investigate locality preserving projections (LPP) in two-dimensional sense. Recently, LPP was proposed for dimensionality reduction, which can detect the intrinsic manifold structure of data and preserve the local information. When image data are concerned, they are often vectorized for LPP. However, the dimension of image data is usually very high, LPP can't be implemented due to singularity of matrix. We propose two methods for image dimensionality reduction: two-dimensional LPP (2DLPP) and bilateral two-dimensional LPP (B2DLPP), which are based directly on 2D image matrices rather than 1D vectors as LPP does. Experiments are conducted on the ORL face database, which shows higher recognition performance of the proposed methods. Sibao Chen 0001, Bin Luo 0001, Renhua Wang |
ICASSP (2) | 1 |
| 2007 | 2D-LPP: A two-dimensional extension of locality preserving projections
Sibao Chen 0001, Haifeng Zhao 0001, Min Kong, Bin Luo 0001 |
Neurocomputing | 1 |
| 2006 | LPP and LPP Mixtures for Graph Spectral Clustering
Bin Luo 0001, Sibao Chen 0001 |
PSIVT | 2 |
| 2004 | Greedy EM algorithm for robust t-mixture modelingabstractThis paper concerns a greedy EM algorithm for t-mixture modeling, which is more robust than Gaussian mixture modeling when a typical points exist or the set of data has heavy tail. Local Kullback divergence is used to determine how to insert new component. The greedy algorithm obviates the complicated initialization. The results are comparable to that of split-and-merge EM algorithm while the proposed algorithm is faster. Also the by product of a sequence of mixture models is useful for model selection. Experiments of synthetic data clustering and unsupervised color image segmentation are given. Sibao Chen 0001, Haixian Wang, Bin Luo 0001 |
ICIG | 1 |