Qiaolin Ye

dblp:44/7694 · DBLP profile ↗
← Back
106ranked-venue papers
23as first author
57since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 19 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Instance-Aware Visual Prompting helps multimodal models see better
Jingxu Wang, Liyong Fu, Qiaolin Ye
Expert Syst. Appl.5
2026 Learning clique-based inter-class affinity for compositional zero-shot learning
Chenyi Jiang, Qiaolin Ye, Zebin Wu 0001, Haofeng Zhang 0001
Pattern Recognit.2
2026 MiniBGTR: Mini-Binary Graph Transformer for Facial Expression Recognition
Xulin Song, Xiyin Wu, Qiaolin Ye
IEEE Trans. Comput. Soc. Syst.4
2026 Spatial-Temporal Self-Compensating Graph Convolutional Network for Skeleton-Based Action Recognition Under Data Constraints
abstract
Skeleton-based human action recognition has emerged as a prominent research focus in computer vision, with significant progress achieved in recent years. However, existing methods often suffer substantial performance degradation under real-world data constraints, such as body occlusion, missing frames, and noise. These limitations critically undermine the robustness of related techniques in practical applications. To address these challenges, we propose a Spatial Temporal Self-compensating Graph Convolutional Network (STSc-GCN), which skillfully utilizes the systematic and regular nature of human movement to mitigate performance degradation caused by data constraints through a data self-compensation mechanism. Specifically, STSc-GCN comprises two key modules: 1) collaborative motion spatial compensation (CMSC). This module designs multiple distinct topological relationships, primarily including Walk-probability Generality Topology and Self-organizing Particularity Topology, respectively, to deeply explore the universal and personalized collaborative relationships between human joints. These relationships help compensate for the lack of information caused by spatial data constraints and 2) meta-action sharpening temporal Compensation (MSTC). This module introduces a novel motion sharpening mechanism that enhances key dynamic information within the meta-action sequences through cross-attention technology, thereby improving model adaptability to missing-frame scenarios. STSc-GCN achieves state-of-the-art performance on four constrained datasets and shows superior results on three widely used standard datasets, confirming its effectiveness in both constrained and general scenarios. Code will be available at https://github.com/XingLi1012/STSc-GCN.git.
Xing Li 0005, Qian Huang 0008, Xin Li 0090, Jinhui Tang 0001, Qiaolin Ye
IEEE Trans. Image Process.6
2026 Learning Robust Discriminant Projections via Double Capped Lp-Norm Distance Metrics With "Min" Constraints
abstract
Recently, there has been a surge in the development of robust norm distance-based linear discriminant analysis (LDA) techniques, which have garnered significant attention in the field of feature extraction. However, a persistent issue that has yet to be resolved is that the successful suppression of outliers may inadvertently impede the accurate discrimination of normal points. To solve this problem, we, in this article, study a novel robust LDA measured by double capped $L_{p}$ -norm distance (CLD) metrics with min constraints (DCLDA) to learn robust discriminant projections, in which normal points and outliers are separately treated. To be specific, it takes a double capped $L_{p}$ -norm with "Min" constraints in the proposed model to measure the distances for between- and within-class dispersions. The proposed model effectively ensures accurate discrimination of normal points by $L_{p}$ -norm, while also eliminating the exaggerated effect of outliers that may arise from larger $p$ values. The resulted objective is not trivial because of its nonconvexity and nonsmoothness. As one of the major contributions of this article, we introduce a new reformulation that provides an objective problem theoretically equivalent to the original. By this reformulation, we develop an effective iterative algorithm to solve the proposed model. The algorithm is proven to be convergent through rigorous theoretical analysis. Extensive experiments were conducted on several real-world datasets across different image classification tasks to showcase the effectiveness of the proposed method.
Xiaobo Chen 0001, Zhao Zhang 0001, Liyong Fu, Qiaolin Ye
IEEE Trans. Neural Networks Learn. Syst.5
2025 RobustEMD: Domain robust matching for cross-domain few-shot medical image segmentation
Yazhou Zhu 0001, Minxian Li, Qiaolin Ye, Tong Xin 0002, Haofeng Zhang 0001
Artif. Intell. Medicine3
2025 Lightweight binary convolutional-transformers fusion network for facial expression recognition
Xiyin Wu, Libo Weng, Qiaolin Ye
Eng. Appl. Artif. Intell.4
2025 Interpretable deep one-class model for forest fire detection
Yangjie Xu, Yiran Ma, Qiaolin Ye, Liyong Fu, Xubing Yang
Expert Syst. Appl.3
2025 CREAM: Few-shot Object Counting with Cross REfinement and Adaptive density Map
Yuanwu Xu, Minxian Li, Qiaolin Ye, Lunbo Li, Haofeng Zhang 0001
Image Vis. Comput.3
2025 Synthetic instance segmentation from semantic image segmentation masks
Zhao Zhang 0001, Liyong Fu, Qiaolin Ye
Knowl. Based Syst.5
2025 SOD-YOLOv8n: Small Object Detection in Remote Sensing Images Based on YOLOv8n
abstract
Small target detection in remote sensing images is a significiant reserach focus within the remote sensing domain. Recently, various YOLO algorithms have demonstrated remarkable achievements in the detection of small targets in remote sensing. However, YOLO-based detection algorithms still face challenges in this context, including limited feature expression capacity, difficulties in mitigating aliasing effects and inadequate adaptability to complex-shaped targets. To address these issues, this paper proposes a novel object detection network SOD-YOLOv8n. First, we propose a novel multi-path feature fusion module (MFFM), which enhances feature extraction through diverse dimensional feature processing strategies (global, local, channel, and spatial). It also fuses complementary information across channels via channel shuffling, thereby augmenting feature representation capabilities, and boosting the detection accuracy of small targets in remote sensing. Secondly, we design an anti-aliasing module (AAM) that employs wavelet pooling technology for frequency decomposition to mitigate the aliasing effect generated during model downsampling, thereby better retaining the key high-frequency information of small targets in remote sensing. Finally, we introduce the Shape-IoU loss function, which emphasizes the edge features of the target shape (such as contours, curvature, etc.) by calculating the similarity between the true target and the target shape in the predicted box, so as to better match targets with complex shapes.We conduct extensive experiments on the AI-TOD and USOD remote sensing small target datasets, and the results show that SOD-YOLOv8n outperforms several established state-of-the-art detection models.
Qiaolin Ye, Le Sun 0002, Zebin Wu 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Multi-domain feature-enhanced attribute updater for generalized zero-shot learning
Yuyan Shi, Chenyi Jiang, Feifan Song 0004, Qiaolin Ye, Yang Long 0001, Haofeng Zhang 0001
Neural Comput. Appl.4
2025 Strengthen contrastive semantic consistency for fine-grained image classification
Yupeng Wang 0004, Yongli Wang 0002, Qiaolin Ye, Wenxi Lang, Can Xu 0006
Pattern Anal. Appl.3
2025 IIS-FVIQA: Finger Vein Image Quality Assessment with intra-class and inter-class similarity
Hengyi Ren, Xijian Fan, Qiaolin Ye
Pattern Recognit.5
2025 Non-rigid object detection via fast one-class model
Xubing Yang, Jingyao Lishen, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu
Pattern Recognit.5
2025 Traffic Agents Trajectory Prediction Based on Enhanced Bidirectional Recurrent Network and Adaptive Social Interaction Model
abstract
Accurate prediction of the future trajectory of traffic agents is imperative to the effective motion planning of autonomous vehicles and mobile robots. Despite enormous progress that has been made toward trajectory prediction, dynamic and crowded traffic scenarios pose major challenges to the understanding and forecasting of traffic agents’ motion behavior. In this paper, we propose a novel trajectory prediction method from the perspective of temporal modeling and social interaction. Specifically, we first put forward a recurrent modeling approach to learn temporal features in favor of capturing long-range and short-range temporal dependencies of individual agents. Then, we construct a social feature learning module to capture the sparse and directional interactions among agents while suppressing the spurious connections. Finally, to reduce the accumulated error during prediction, a coordinated bidirectional decoding module is developed where temporal and social features can be properly integrated into the forward and backward prediction processes. Extensive experiments are performed on four real-world trajectory prediction benchmarks, and the results demonstrate the superiority of our method compared with other competing approaches. Detailed ablation studies are also performed to evaluate the effectiveness of each model component. Note to Practitioners—Motion planning is one of the crucial components of autonomous systems, such as intelligent vehicles and mobile robots. For example, the safety and efficiency of motion planning can be drastically improved if the future trajectories of surrounding agents, e.g., pedestrians, bicyclists, cars, etc., can be accurately forecasted. Motivated by the above requirements, this article develops an advanced deep learning model that can learn temporal and social features from trajectory data and perform accuracy prediction. This work aims to enhance the bidirectional recurrent network for dealing with trajectory data with evident temporal characteristics. In addition, this work introduces a novel adaptive social interaction modeling approach that overcomes the inherent defect of the fixed threshold method. This work also addresses the error accumulation problem in the prediction. The proposed model is evaluated on several datasets, and the results demonstrate its effectiveness. Our approach has broad application prospects in autonomous driving and mobile robots.
Xiaobo Chen 0001, Yuwen Liang, Chuan Hu 0003, Hai Wang 0003, Qiaolin Ye
IEEE Trans Autom. Sci. Eng.5
2025 Multibranch Attentive Transformer With Joint Temporal and Social Correlations for Traffic Agents Trajectory Prediction
abstract
Accurately predicting the future trajectories of traffic agents is paramount for autonomous unmanned systems, such as self-driving cars and mobile robotics. Extracting abundant temporal and social features from trajectory data and integrating the resulting features effectively pose great challenges for predictive models. To address these issues, this article proposes a novel multibranch attentive transformer (MBAT) trajectory prediction network for traffic agents. Specifically, to explore and reveal diverse correlations of agents, we propose a decoupled temporal and spatial feature learning module with multibranch to extract temporal, spatial, as well as spatiotemporal features. Such design ensures each branch can be specifically tailored for different types of correlations, thus enhancing the flexibility and representation ability of features. Besides, we put forward an attentive transformer architecture that simultaneously models the complex correlations possibly occurring in historical and future timesteps. Moreover, the temporal, spatial, and spatiotemporal features can be effectively integrated based on different types of attention mechanisms. Empirical results demonstrate that our model achieves outstanding performance on public ETH, UCY, SDD, and INTERACTION datasets. Detailed ablation studies are conducted to verify the effectiveness of the model components.
Xiaobo Chen 0001, Yuwen Liang, Qiaolin Ye, Yingfeng Cai
IEEE Trans. Comput. Soc. Syst.4
2025 Robust Multiple Flat Projections Clustering With Truncated Distance Maximization Constraints
abstract
Recently, interest in flat-type projection clustering methods has grown as they improve learner's performance by exploring multiple projection subspaces. However, solvers used in previous representative works predominantly rely on greedy search strategies, which incur high computational costs and fail to consider interdependencies between projections. Moreover, these methods do not simultaneously guarantee the effective suppression of outliers and noisy data at cluster boundaries, ultimately compromising data discrimination. To address these limitations and discover a more effective subspace for each flat, we propose robust multiple flat projections clustering (RMFPC). This method computes within- and between-cluster distances using the L2,1-norm to enhance robustness against outliers. Furthermore, we propose a truncated distance maximization constraint (TDMC) to eliminate the influence of noisy data on cluster separability. The resulting objective is presented in a ratio form, which is not trivial. We provide a novel formulation to achieve a theoretically equivalent problem. Based on this reformulation, we develop an efficient non-greedy solution algorithm. In addition, a cluster center optimization mechanism is incorporated into the solution process to accurately estimate the distribution of each cluster center. The convergence analysis and proof of the proposed algorithm are provided. Experiments on both toy and real-world datasets demonstrate the effectiveness of the proposed method.
Zhao Zhang 0001, Xiaobo Chen 0001, Zhongqi Xu, Liyong Fu, Qiaolin Ye
IEEE Trans. Cybern.6
2025 Learning Cross-Task Features With Mamba for Remote Sensing Image Multitask Prediction
abstract
Multitask learning (MTL) for remote sensing (RS) image is a rapidly evolving field that requires simultaneous predictions across several related tasks. However, many existing MTL methods often overlook the exploring of cross-task features, while the strong interdependencies among tasks are critical for MTL. In this article, we propose RSMTMamba, an innovative MTL framework that integrates Mamba for multitask prediction in RS images. Our network simultaneously performs semantic segmentation, height estimation, and boundary detection within a unified architecture. The proposed architecture prioritizes the decoder, with a shared encoder for feature extraction. Specifically, a Mamba-based cross-task feature learning (MCFL) module is introduced to capture the interrelations among different tasks. Unlike transformer-based architecture, which requires significant computational resources, the MCFL module can model both local and global cross-task relationships for RS image with linear complexity. Additionally, Mamba-integrated refine decoders are utilized to aggregate features from the encoder, preliminary decoders, and the MCFL module, which enhances multitask prediction performance. The experimental results on three RS datasets demonstrate that our proposed CFLMamba achieves the state-of-the-art prediction performance, outperforming several deep neural networks in RS image analysis. The code is available athttps://github.com/sycs-2024/RSMultitaskMamba.
Liang Xiao 0001, Jianyu Chen 0003, Qian Du 0001, Qiaolin Ye
IEEE Trans. Geosci. Remote. Sens.5
2025 Nonconvex Transform-Based Low-Rank Tensor Completion With Coupled Spatiotemporal Relation Learning for Traffic Data Recovery
abstract
With the rapid development of sensor technology, Intelligent Transportation Systems (ITS) are capable of collecting vast amounts of traffic data. However, unforeseen interruptions during the process of data collection, transmission, and storage often lead to data loss, posing significant challenges to data accuracy and integrity. To address this issue, we have conducted an in-depth analysis of the unique physical characteristics of traffic data and optimized model design based on these characteristics to improve the accuracy of data recovery. This article introduces an innovative low-rank tensor completion model that leverages both global and local features of traffic data to accurately fill in missing values. Specifically, we propose using a weighted composite tensor (WCT) norm as a non-convex alternative to capture the multi-dimensional low-rank properties of traffic data in the transform domain. Additionally, to further enhance the precision of data recovery, we introduce a coupled structure regression method that combines temporal smoothness with sample similarity, aiding in revealing complex spatiotemporal data correlation patterns. To solve the resulting non-convex optimization problem, we have designed an efficient iterative algorithm based on the Alternating Direction Method of Multipliers (ADMM) and conducted a detailed theoretical analysis of its convergence and computational complexity. Experimental results show that compared to other competing algorithms, our model demonstrates significant advantages across three real-world traffic datasets, proving its superior performance in data recovery.
Xiaobo Chen 0001, Qiaolin Ye
IEEE Trans. Intell. Transp. Syst.5
2025 Convergence Analysis on Trace Ratio Linear Discriminant Analysis Algorithms
abstract
Linear discriminant analysis (LDA) may yield an inexact solution by transforming a trace ratio problem into a corresponding ratio trace problem. Most recently, optimal dimensionality LDA (ODLDA) and trace ratio LDA (TRLDA) have been developed to overcome this problem. As one of the greatest contributions, the two methods design efficient iterative algorithms to derive an optimal solution. However, the theoretical evidence for the convergence of these algorithms has not yet been provided, which renders the theory of ODLDA and TRLDA incomplete. In this correspondence, we present some rigorously theoretical insight into the convergence of the iterative algorithms. To be specific, we first demonstrate the existence of lower bounds for the objective functions in both ODLDA and TRLDA, and then establish proofs that the objective functions are monotonically decreasing under the iterative frameworks. Based on the findings, we disclose the convergence of the iterative algorithms finally.
Qiaolin Ye, Liyong Fu
IEEE Trans. Neural Networks Learn. Syst.1
2025 Cracking the Code: LoRa Physical-Layer Insights and Signal Recovery Under Cross-Technology Interference
abstract
Low-Power Wide-Area Networks (LPWANs) have emerged as a promising communication technology for the Internet of Things (IoT). However, frequency overlap among wireless networks using different radio technologies creates significant interference, compromising communication reliability. This challenge is particularly urgent in LoRa networks, which coexist in the 2.4 GHz ISM band with other IoT transmitters capable of transmitting at much higher power levels. In our study, we begin by providing a comprehensive understanding of the LoRa physical layer (PHY), including insights into modulation and demodulation mechanisms. Leveraging this knowledge, we successfully implemented a real-time LoRa PHY on the GNU Radio Software-Defined Radio platform. To address cross-technology interference during peak detection, we introduce a spectrum merging technique that maintains phase coherence between superimposed peaks, minimizing spectral leakage artifacts. Beyond that, our analysis actively enhances the performance of commercial LoRa devices. Furthermore, we systematically explore the interference dynamics between LoRa and IEEE 802.15.4g networks. Our rigorous investigation reveals LoRa’s ability to achieve high packet reception rates, even in the presence of strong IEEE 802.15.4g interference.
Demin Gao, Ye Liu 0004, Qiaolin Ye, Qing Yang 0003, Honggang Wang 0001
IEEE Trans. Wirel. Commun.3
2025 LoBee: Bidirectional Communication Between LoRa and ZigBee Based on Physical-Layer CTC
abstract
LoRa networks operating in a star topology, this configuration creates a single point of failure and may limit scalability and reliability in areas that are large and geographically dispersed. In order to improve the overall transmission capabilities of the network, recent studies show that adding LoRa to the ZigBee devices effectively disseminates network management. By doing so, the strengths of both technologies can be leveraged, with LoRa serving as the long-range transmitter and ZigBee functioning as the mesh network. In this study, we present LoBee, a novel bidirectional communication method between LoRa and ZigBee that relies on Physical-Layer Cross-Technology Communication. Despite the fact that LoRa and ZigBee utilize different modulation techniques, ZigBee devices can detect and recognize LoRa chirps through the process of sampling the received signal strength. For the transmissions from ZigBee to LoRa devices, we carefully select the input chips to generate specific waveforms, where LoBee detects the preamble of a ZigBee frame based on the locations of the repeated peaks. Our evaluation, which was conducted using USRP and commodity devices, demonstrates that LoBee is capable of achieving concurrent bidirectional wireless communications, with a data rate of approximately 639.38 bits per second from LoRa to ZigBee and from ZigBee to LoRa with more than 90% frame reception rate in the 2.4 GHz frequency band.
Demin Gao, Haoyu Wang 0015, Yongrui Chen 0001, Qiaolin Ye, Weizheng Wang 0001, Xiuzhen Guo, Shuai Wang 0008, Yunhuai Liu, Tian He 0001
IEEE Trans. Wirel. Commun.4
2024 Global superpixel-merging via set maximum coverage
Xubing Yang, Zhengxiao Zhang, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu
Eng. Appl. Artif. Intell.5
2024 Fusing spatial and frequency features for compositional zero-shot image classification
Suyi Li 0005, Chenyi Jiang, Qiaolin Ye, Wankou Yang, Haofeng Zhang 0001
Expert Syst. Appl.3
2024 Robust GEPSVM classifier: An efficient iterative optimization framework
Yan Liu 0038, Yanmeng Li, Qiaolin Ye, Dongjun Yu, Yong Qi 0002
Inf. Sci.4
2024 An Intelligent Deep Learning Framework for Traffic Flow Imputation and Short-term Prediction Based on Dynamic Features
Xianhui Zong, Yong Qi 0002, Qiaolin Ye
Knowl. Based Syst.4
2024 Mutual Balancing in State-Object Components for Compositional Zero-Shot Learning
Chenyi Jiang, Qiaolin Ye, Yuming Shen, Zheng Zhang 0006, Haofeng Zhang 0001
Pattern Recognit.2
2024 General Optimization Methods for YOLO Series Object Detection in Remote Sensing Images
abstract
The You Only Look Once (YOLO) series of object detection algorithms has attracted considerable attention for its notable advantages in speed and accuracy, resulting in widespread applications in various real-world scenarios. However, achieving outstanding accuracy on remote sensing images with densely arranged small targets and complex backgrounds remains a challenging task. To address this issue, this letter proposes two easily integrated modules suitable for the YOLO architecture, namely global semantic information extraction (GSIE) and adaptive feature fusion (AFF). The GSIE module is designed to overcome the limitation of local information in traditional methods and facilitate global semantic information interaction by introducing multi-angle feature rotation to extend the receptive field. The AFF module effectively captures fine-grained features of objects by dynamically adjusting fusion weights, thereby reducing the loss of deep semantic information during feature transfer and fusion. The experimental results on the VEDAI and LEVIR remote sensing datasets demonstrate that when embedding these two modules into YOLO series algorithms that only use the small-scale detector, there is a significant improvement in performance while reducing computational complexity.
Guozheng Nan, Yue Zhao 0036, Chengxing Lin 0002, Qiaolin Ye
IEEE Signal Process. Lett.4
2024 Composite Nonconvex Low-Rank Tensor Completion With Joint Structural Regression for Traffic Sensor Networks Data Recovery
abstract
Traffic sensor networks allow convenient collection of travel data that are of great significance for intelligent transportation systems (ITSs). However, the universality of missing data impedes the application of ITS and thus accurate missing data recovery is indispensable in practice. Typically, the global low-rankness and local spatiotemporal smoothness exist in underlying traffic tensor data. In light of this, this article proposes an improved low-rank tensor completion (LRTC) model by exploiting abundant structural information from incomplete tensors. Specifically, a logarithm power composite (LPC)-norm is first proposed as a nonconvex substitute of the rank function, leading to a flexible characterization of tensor multidimensional correlation. Then, a joint structural regression (JSR) model is presented to simultaneously leverage the intrinsic temporal continuity and profile similarity of traffic data. By doing so, we construct a novel nonconvex LRTC model by integrating the global low-rankness and fine-grained spatiotemporal structure that are complementary to each other. To solve the proposed model, following the optimization framework of the alternating direction method of multipliers (ADMMs), we develop an efficient iterative algorithm where each step can be solved in a closed form. Extensive experiments on four real-world traffic data are conducted to evaluate the effectiveness of the proposed approach. The results demonstrate that compared with other tensor completion methods, our model significantly improves the recovery performance.
Xiaobo Chen 0001, Feng Zhao 0006, Fuwen Deng, Qiaolin Ye
IEEE Trans. Comput. Soc. Syst.5
2024 RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
abstract
General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these models primarily learn low-level features and require annotated data for fine-tuning. Moreover, they are inapplicable for retrieval and zero-shot applications due to the lack of language understanding. To address these limitations, we propose RemoteCLIP, the first vision-language foundation model for remote sensing that aims to learn robust visual features with rich semantics and aligned text embeddings for seamless downstream application. To address the scarcity of pre-training data, we leverage data scaling which converts heterogeneous annotations into a unified image-caption data format based on Box-to-Caption (B2C) and Mask-to-Box (M2B) conversion. By further incorporating UAV imagery, we produce a 12 × larger pretraining dataset than the combination of all available datasets. RemoteCLIP can be applied to a variety of downstream tasks, including zero-shot image classification, linear probing,k-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. Evaluation on 16 datasets, including a newly introduced RemoteCount benchmark to test the object counting ability, shows that RemoteCLIP consistently outperforms baseline foundation models across different model scales. Impressively, RemoteCLIP beats the state-of-the-art method by 9.14% mean recall on the RSITMD dataset and 8.92% on the RSICD dataset. For zero-shot classification, our RemoteCLIP outperforms the CLIP baseline by up to 6.39% average accuracy on 12 downstream datasets.
Fan Liu 0003, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Qiaolin Ye, Liyong Fu, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 MDC-FusFormer: Multiscale Deep Cross-Fusion Transformer Network for Hyperspectral and Multispectral Image Fusion
abstract
The spatial resolution of hyperspectral images (HSIs) is usually limited due to internal imaging mechanisms. To obtain imagery with high spectral and high spatial resolutions, which is essential for subsequent HSI processing tasks, a cost-effective approach is to fuse HSI with multispectral images (MSIs). One highly effective fusion method is the convolutional neural network (CNN). However, CNNs have limitations in capturing global information and complex features. Recently, visual transformers (ViTs) have garnered interest for their ability to process non-local information. Despite this, existing HSI-MSI fusion methods suffer from insufficient spatial-spectral feature interaction, resulting in suboptimal fusion quality. To address these challenges, we propose a multiscale deep cross-fusion transformer (MDC-FusFormer) network for HSI and MSI fusion. This network effectively performs the interactive fusion of spatial-spectral features, thereby enhancing the quality of the fused images. MDC-FusFormer employs a three-branch network architecture consisting of two independent progressive feature mining modules (PFMMs), a multiscale deep cross-fusion attention module, and a spatial-spectral feature fusion module. Initially, shallow features at different scales of MSI and HSI are recursively extracted through successive up- and down-sampling using CNNs. These features then interact with the deep cross-modal information at corresponding scales through the attention block. Finally, a multidimensional refinement convolution block (MRCB) is applied to refine the feature information, which is then combined with cascaded up-sampling to reconstruct the high-resolution fused image step by step. Experimental results on five datasets indicate that, compared to nine other methods, MDC-FusFormer delivers superior performance.
Le Sun 0002, Jianxiao Zhou, Qiaolin Ye, Zebin Wu 0001, Qiao Chen 0004, Zhongqi Xu, Liyong Fu
IEEE Trans. Geosci. Remote. Sens.3
2023 Robust generalized canonical correlation analysis
Qiaolin Ye, Dongjun Yu, Yong Qi 0002
Appl. Intell.3
2023 Learning to Reduce Information Bottleneck for Object Detection in Aerial Images
abstract
Object detection in aerial images is a critical and essential task in the fields of geoscience and remote sensing. Despite the popularity of computer vision methods in detecting objects, these methods have been faced with significant limitations of aerial images such as appearance occlusion and variable object sizes. In this letter, we explore the limitations of conventional neck networks in object detection by analyzing information bottlenecks. We propose an enhanced neck network to address the information deficiency issue in current neck networks. Our proposed neck network, which serves as a bridge between the backbone network and the head network, comprises a global semantic network (GSNet) and a feature fusion refinement module (FRM). The GSNet is designed to perceive contextual surroundings and propagate discriminative knowledge through a bidirectional global pattern. The FRM is developed to exploit different levels of features to capture comprehensive location information. We validate the efficacy and efficiency of our approach through experiments conducted on two challenging datasets, DOTA and HRSC2016. Our method outperforms existing approaches in terms of accuracy and complexity, demonstrating the superiority of our proposed method.
Zhihao Song, Xuesong Jiang, Qiaolin Ye
IEEE Geosci. Remote. Sens. Lett.5
2023 Preferred vector machine for forest fire detection
Xubing Yang, Zhichun Hua, Li Zhang 0057, Xijian Fan, Fuquan Zhang 0004, Qiaolin Ye, Liyong Fu
Pattern Recognit.6
2023 Motion Stimulation for Compositional Action Recognition
abstract
Recognizing the unseen combinations of action and different objects, namely (zero-shot) compositional action recognition, is extremely challenging for conventional action recognition algorithms in real-world applications. Previous methods focus on enhancing the dynamic clues of objects that appear in the scene by building region features or tracklet embedding from ground-truths or detected bounding boxes. These methods rely heavily on manual annotation or the quality of detectors, which are inflexible for practical applications. In this work, we aim to mining the temporal clues from moving objects or hands without explicit supervision. Thus, we propose a novel Motion Stimulation (MS) block, which is specifically designed to mine dynamic clues of the local regions autonomously from adjacent frames. Furthermore, MS consists of the following three steps: motion feature extraction, motion feature recalibration, and action-centric excitation. The proposed MS block can be directly and conveniently integrated into existing video backbones to enhance the ability of compositional generalization for action recognition algorithms. Extensive experimental results on three action recognition datasets, the Something-Else, IKEA-Assembly and EPIC-KITCHENS datasets, indicate the effectiveness and interpretability of our MS block.
Yuhui Zheng, Zhao Zhang 0001, Yazhou Yao, Xijian Fan, Qiaolin Ye
IEEE Trans. Circuits Syst. Video Technol.6
2023 Multiattention Joint Convolution Feature Representation With Lightweight Transformer for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is currently a hot topic in the field of remote sensing. The goal is to utilize the spectral and spatial information from HSI to accurately identify land covers. Convolution neural network (CNN) is a powerful approach for HSI classification. However, CNN has limited ability to capture non-local information to represent complex features. Recently, vision transformers (ViTs) have gained attention due to their ability to process non-local information. Yet, under the HSI classification scenario with ultra-small sample rates, the spectral-spatial information given to ViTs for global modeling is insufficient, resulting in limited classification capability. Therefore, in this article, Multi-Attention Joint Convolution Feature Representation with Lightweight Transformer (MAR-LWFormer) is proposed, which effectively combines the spectral and spatial features of HSI to achieve efficient classification performance at ultra-small sample rates. Specifically, we use a three-branch network architecture to extract multi-scale convolved 3D-CNN, EMAP, and LBP features of HSI, respectively, by taking full exploitation of ultra-small training samples. Second, we design a series of multi-attention modules to enhance spectral-spatial representation for the three types of features and to improve the coupling and fusion of multiple features. Third, we propose an explicit feature attention tokenizer to transform the feature information, which maximizes the effective spectral-spatial information retained in the flat tokens. Finally, the generated tokens are input to the designed lightweight transformer for encoding and classification. Experimental results on three datasets validate that MAR-LWFormer has an excellent performance in HSI classification at ultra-small sample rates when compared to several state-of-the-art classifiers.
Yu Fang 0012, Qiaolin Ye, Le Sun 0002, Yuhui Zheng, Zebin Wu 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Joint Classification of Hyperspectral and LiDAR Data Using a Hierarchical CNN and Transformer
abstract
The joint use of multisource remote-sensing (RS) data for Earth observation missions has drawn much attention. Although the fusion of several data sources can improve the accuracy of land-cover identification, many technical obstacles, such as disparate data structures, irrelevant physical characteristics, and a lack of training data, exist. In this article, a novel dual-branch method, consisting of a hierarchical convolutional neural network (CNN) and a transformer network, is proposed for fusing multisource heterogeneous information and improving joint classification performance. First, by combining the CNN with a transformer, the proposed dual-branch network can significantly capture and learn spectral–spatial features from hyperspectral image (HSI) data and elevation features from light detection and ranging (LiDAR) data. Then, to fuse these two sets of data features, a cross-token attention (CTA) fusion encoder is designed in a specialty. The well-designed deep hierarchical architecture takes full advantage of the powerful spatial context information extraction ability of the CNN and the strong long-range dependency modeling ability of the transformer network based on the self-attention (SA) mechanism. Four standard datasets are used in experiments to verify the effectiveness of the approach. The experimental results reveal that the proposed framework can perform noticeably better than state-of-the-art methods. The source code of the proposed method will be available publicly athttps://github.com/zgr6010/Fusion_HCT.git.
Guangrui Zhao, Qiaolin Ye, Le Sun 0002, Zebin Wu 0001, Byeungwoo Jeon
IEEE Trans. Geosci. Remote. Sens.2
2023 Keywords-enhanced Deep Reinforcement Learning Model for Travel Recommendation
abstract
Tourism is an important industry and a popular entertainment activity involving billions of visitors per annum. One challenging problem tourists face is identifying satisfactory products from vast tourism information. Most of travel recommendation methods regard the recommendation procedure as a static process and only focus on immediate rewards. Meanwhile, they often infer user intensions from click behaviors and ignore the informative keywords of the clicked products. To this end, in this article, we present a Keywords-enhanced Deep Reinforcement Learning model (KDRL) framework. Specifically, we formalize travel recommendation as a Markov Decision Process and implement it upon the Actor–Critic framework. It integrates keyword information into the reinforcement learning–(RL) based recommendation framework by devising novel state representation and reward function and learns the travel recommendation and keywords generation simultaneously. To the best of our knowledge, this is the first time that keywords are explicitly discussed and used in RL-based travel recommendations. Extensive experiments are performed on the real-world datasets and the results clearly show the superior performance of KDRL compared with the baseline methods.
Lei Chen 0079, Jie Cao 0001, Weichao Liang, Jia Wu 0001, Qiaolin Ye
ACM Trans. Web5
2022 Class Concentration with Twin Variational Autoencoders for Unsupervised Cross-Modal Hashing
Yazhou Zhu 0001, Shengbin Liao, Qiaolin Ye, Haofeng Zhang 0001
ACCV (6)4
2022 Robust ensemble method for short-term traffic flow prediction
Liyong Fu, Yong Qi 0002, Dongjun Yu, Qiaolin Ye
Future Gener. Comput. Syst.5
2022 Flexible capped principal component analysis with applications in image recognition
Liyong Fu, Qiaolin Ye
Inf. Sci.3
2022 Learning discriminative and representative feature with cascade GAN for generalized zero-shot learning
Jingren Liu, Liyong Fu, Haofeng Zhang 0001, Qiaolin Ye, Wankou Yang, Li Liu 0004
Knowl. Based Syst.4
2022 Learning a robust classifier for short-term traffic state prediction
Liyong Fu, Yong Qi 0002, Qiaolin Ye, Dongjun Yu
Knowl. Based Syst.5
2022 Multi-view distance metric learning via independent and shared feature subspace with applications to face and forest fire recognition, and remote sensing classification
Liyong Fu, Yawen Cheng, Qiaolin Ye
Knowl. Based Syst.4
2022 3-D Contour Deformation for the Point Cloud Segmentation
abstract
The 3-D point cloud segmentation has played an important role in spatial structure analysis. Nowadays, segmentation methods either use a primitive-based strategy to fit points in predefined geometric shapes or group points based on their attributes (e.g., spatial distance). However, the required segmentation results, e.g., primitive level or object level, depend on the application. Therefore, this letter develops a semiautomatic method to extract contours for the users’ desired segmentation. First, we initialize a 3-D closed curve for the target. Second, we calculate the internal and external force based on the proposed vector flow to deform the curve. The deformation equation is solved based on the Euler equation and calculated iteratively. Finally, the curve is converged as object contours. After one removes contours, those disjoint points are grouped as the users’ desired instances. Experiments are conducted on various point clouds to demonstrate the effectiveness in terms of accuracy and consistency. Our quantitative evaluation outperformed selected primitive- and object-based methods, which presents a new viewpoint to the point cloud processing.
Sheng Xu 0003, Wen Han, Weidu Ye, Qiaolin Ye
IEEE Geosci. Remote. Sens. Lett.4
2022 Classification of 3-D Point Clouds by a New Augmentation Convolutional Neural Network
abstract
Nowadays, the classification of point clouds has become a fundamental problem in 3-D information study. Different from the deep learning process of natural images, 3-D point clouds are massive and unorganized, which can be difficultly captured features by the convolution process directly. This letter proposes a new augmentation convolutional neural network (ACNN) to classify point clouds by adding a key augmentation layer before the classical sampling and convolution structure. Input data will be augmented before each sampling layer, which brings abundant learning information to help the network capture more local structures. In order to make the augmentation more effective, we formulate the parameters of augmentation layers learnable in the learning process according to the loss function. The proposed augmentation is based on automatically tuning the magnitude of the smoothness, which plays a significant role in point cloud processing and provides local features, for example, edges, contours, and edges. Results show that we have achieved the overall accuracy of 92.52% and 89.11% in the object classification on ModelNet10 and ModelNet40, respectively, which shows our superiority over other methods. Besides, the ACNN achieves an average miscalculation error of 0.28 and cross-entropy loss of 0.48 in the classification of laser scanning point clouds, which shows high robustness to noise and density in the outdoor scene classification.
Sheng Xu 0003, Weidu Ye, Qiaolin Ye
IEEE Geosci. Remote. Sens. Lett.4
2022 Improving Deep Learning-Based Cloud Detection for Satellite Images With Attention Mechanism
abstract
Clouds in satellite images limit the ability of imagery to extract the ground information, which makes it difficult to the following image analysis tasks. Hence, cloud detection is a changeling but fundamental task in the preprocessing of satellite image processing. In this letter, an encoder–decoder neural network architecture, cloud detection with ACON and attention mechanism (CAA)-UNet, is proposed for cloud detection. CAA-UNet is based on U-Net architecture and incorporates the attention mechanism. The asymmetric encoder and decoder blocks are proposed to discover more discriminative features. Then, a modified attention gate is integrated into each skip connection to highlight salient features. These make our model distinguish between the cloud and noncloud more accurately. Furthermore, the recent new and effective activate function, ActivateOrNot (ACON), is introduced into our model, which allows each neuron to adaptively activate or not, and improves the performance remarkably. Finally, the experiment results on two cloud datasets, Landsat-8 and a high-resolution cloud (HRC) cover validation dataset, show that CAA-UNet outperforms the state-of-the-art methods, especially in comprehensive indicators: Jaccard index,$F_{1}$score, and overall accuracy.
Li Zhang 0057, Xubing Yang, Rui Jiang 0007, Qiaolin Ye
IEEE Geosci. Remote. Sens. Lett.5
2022 Robust distance metric optimization driven GEPSVM classifier for pattern classification
Liyong Fu, Tian'an Zhang, Jun Hu 0010, Qiaolin Ye, Yong Qi 0002, Dongjun Yu
Pattern Recognit.5
2022 Multiview Learning With Robust Double-Sided Twin SVM
abstract
Multiview learning (MVL), which enhances the learners' performance by coordinating complementarity and consistency among different views, has attracted much attention. The multiview generalized eigenvalue proximal support vector machine (MvGSVM) is a recently proposed effective binary classification method, which introduces the concept of MVL into the classical generalized eigenvalue proximal support vector machine (GEPSVM). However, this approach cannot guarantee good classification performance and robustness yet. In this article, we develop multiview robust double-sided twin SVM (MvRDTSVM) with SVM-type problems, which introduces a set of double-sided constraints into the proposed model to promote classification performance. To improve the robustness of MvRDTSVM against outliers, we take L1-norm as the distance metric. Also, a fast version of MvRDTSVM (called MvFRDTSVM) is further presented. The reformulated problems are complex, and solving them are very challenging. As one of the main contributions of this article, we design two effective iterative algorithms to optimize the proposed nonconvex problems and then conduct theoretical analysis on the algorithms. The experimental results verify the effectiveness of our proposed methods.
Qiaolin Ye, Zhao Zhang 0001, Yuhui Zheng, Liyong Fu, Wankou Yang
IEEE Trans. Cybern.1
2022 BASNet: Burned Area Segmentation Network for Real-Time Detection of Damage Maps in Remote Sensing Images
abstract
Since remote sensing images of post-fire vegetation are characterized by high resolution, multiple interferences, and high similarities between the background and the target area, it is difficult for existing methods to detect and segment the burned area in these images with sufficient speed and accuracy. In this paper, we apply Salient Object Detection (SOD) to burned area segmentation, the first time this has been done, and propose an efficient burned area segmentation network (BASNet) to improve the performance of unmanned aerial vehicle (UAV) high-resolution image segmentation. BASNet comprises positioning module and refinement module. The positioning module efficiently extracts high-level semantic features and general contextual information via global average pooling layer and convolutional block to determine the coarse location of the salient region. The refinement module adopts the convolutional block attention module to effectively discriminate the spatial location of objects. In addition, to effectively combine edge information with spatial location information in the lower layer of the network and the high-level semantic information in the deeper layer, we design the residual fusion module to perform feature fusion by level to obtain the prediction results of the network. Extensive experiments on two UAV datasets collected from Chongli in China and Andong in South Korea, demonstrate that our proposed BASNet significantly outperforms state-of-the-art SOD methods quantitatively and qualitatively. BASNet also achieves a promising prediction speed for processing high-resolution UAV images, thus providing wide-ranging applicability in post-disaster monitoring and management.
Weihao Bo, Xijian Fan, Tardi Tjahjadi, Qiaolin Ye, Liyong Fu
IEEE Trans. Geosci. Remote. Sens.5
2022 Robust Least Squares Twin Support Vector Regression With Adaptive FOA and PSO for Short-Term Traffic Flow Prediction
abstract
Accurate short-term traffic flow prediction plays an important role in the field of modern Intelligent Transportation Systems. Since various uncontrollable factors (e.g.weather, traffic jams or accidents), collected traffic data inevitably contain outliers. This makes it challenge to achieve satisfactory results for traffic flow prediction. Least Squares Twin Support Vector Regression (LSTSVR) has been shown to provide a powerful potential in nonlinear prediction problems. This is especially true when using appropriate heuristic algorithms to determine the parameters of nonlinear LSTSVR. In view of this, a novel LSTSVR model based on the robust$\text{L}_{2,\mathrm {p}}$-norm ($0< p\le 2$) distance is proposed to alleviate the negative effect of traffic data with outliers, called PLSTSVR. An iterative algorithm is designed to solve the optimization problem of PLSTSVR, which has great potential for solving other relevant optimization problems. To search the parameters of constructed PLSTSVR, this paper constructs two traffic flow prediction models based on PLSTSVR and heuristic algorithms (Fruit Fly Optimization Algorithm and Particle Swarm Optimization), called PLSTSVR-FOA and PLSTSVR-PSO. Extensive experiments demonstrate that the constructed models are more effective and robust than other competing models in various experimental settings.
Yong Qi 0002, Qiaolin Ye, Dongjun Yu
IEEE Trans. Intell. Transp. Syst.3
2022 Learning Robust Discriminant Subspace Based on Joint L₂, ₚ- and L₂, ₛ-Norm Distance Metrics
abstract
-norm as the distance metric. However, both of their robustness and discriminant power are limited. In this article, we present a new robust discriminant subspace (RDS) learning method for feature extraction, with an objective function formulated in a different form. To guarantee the subspace to be robust and discriminative, we measure the within-class distances based on [Formula: see text]-norm and use [Formula: see text]-norm to measure the between-class distances. This also makes our method include rotational invariance. Since the proposed model involves both [Formula: see text]-norm maximization and [Formula: see text]-norm minimization, it is very challenging to solve. To address this problem, we present an efficient nongreedy iterative algorithm. Besides, motivated by trace ratio criterion, a mechanism of automatically balancing the contributions of different terms in our objective is found. RDS is very flexible, as it can be extended to other existing feature extraction techniques. An in-depth theoretical analysis of the algorithm's convergence is presented in this article. Experiments are conducted on several typical databases for image classification, and the promising results indicate the effectiveness of RDS.
Liyong Fu, Zechao Li, Qiaolin Ye, Qingwang Liu, Xiaobo Chen 0001, Xijian Fan, Wankou Yang, Guowei Yang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2021 Discriminative Additive Scale Loss for Deep Imbalanced Classification and Embedding
abstract
Real-world data in emerging applications may suffer from highly-skewed class imbalanced distribution, however how to deal with this kind of problem appropriately through deep learning needs further investigation. In this paper, we mainly propose a novel cross-entropy based loss function, referred to as Additive Scale Loss (ASL), for deep representation learning and imbalanced classification. To deal with the class imbalanced problem, ASL aims at increasing the loss in case of misclassification, which can avoid the superimposed loss values caused by the large amount of easily classified data in the unbalanced database to dominate the loss value of misclassified data. Moreover, in real-world applications, one data source may be used for multiple scenarios, such as classification and embedding learning, however training two separable models to handle these problems is costly, especially in deep learning area. To tackle this issue, we present and integrate a discriminative inter-class separation term into ASL, and propose a discriminative ASL (D-ASL), which can not only improve the classification performance, but also obtain discriminative representations simultaneously. The discriminative inter-class separation term is general, and can be easily integrated to other loss functions, such as CE and FL, as the byproducts. Finally, a new deep convolutional neural network equipped with D-ASL and a fully-connected (FC) layer is proposed, which can classify the imbalanced image data and obtain the discriminative representations at the same time. Extensive experimental results verified the superior performance of our method.
Zhao Zhang 0001, Weiming Jiang, Yang Wang 0023, Qiaolin Ye, Ming-Bo Zhao, Mingliang Xu 0001, Meng Wang 0001
ICDM4
2021 Pixel-level automatic annotation for forest fire image
Xubing Yang, Run Chen, Fuquan Zhang 0004, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu
Eng. Appl. Artif. Intell.6
2021 Double L2, p-norm based PCA for feature extraction
Pu Huang 0004, Qiaolin Ye, Fanlong Zhang, Guowei Yang 0002, Zhangjing Yang
Inf. Sci.2
2021 Recurrent Thrifty Attention Network for Remote Sensing Scene Recognition
abstract
The self-attention mechanism has been empirically shown its effectiveness in a wide range of computer vision applications. However, it is usually criticized for the expensive computation cost. Although some revised methods are proposed in the recent past, they are not maturely applicable to remote sensing scene (RSS) images. To address this problem, in this article, we propose a simple yet effective context acquisition module, named thrifty attention, which can capture the long-range dependence efficiently and effectively. Moreover, a recurrent version for thrifty attention, termed recurrent thrifty attention (RTA), is further proposed to take the long-range multihop communications in space–time for RSS images. RTA is a general global contextual information acquisition module that can be used in any hierarchy of deep convolutional neural networks. To demonstrate its superiority, we deploy it to the classical ResNet and establish our proposed RTA Network (RTANet). Extensive experiments are carried out on two levels of the RSS recognition tasks, i.e., the image-level RSS classification and the instance-level RSS object detection. Compared with the standard self-attention mechanism, RTA can reduce at most 0.43 M model parameters while increasing a slight of model floating-point operations per second (FLOPs). Furthermore, results on RSS classification and object detection further verify the accuracy superiority of RTANet.
Liyong Fu, Qiaolin Ye
IEEE Trans. Geosci. Remote. Sens.3
2020 Uncertainty-aware Cross-dataset Facial Expression Recognition via Regularized Conditional Alignment
abstract
Cross-dataset facial expression recognition (FER) has remained a challenging problem due to the obvious biases caused by diverse subjects and various collection conditions. To this end, domain adaption can be adopted as an effective solution by learning invariant representations across domains (datasets). However, FER requires special consideration of its specific problems e.g., uncertainties caused by ambiguous facial images, and diverse inter- and intra-class relationship. Such uncertainties already exist in single dataset FER, and could be significantly aggravated by enlarged class-wise discrepancies under cross-dataset scenarios. To mitigate this problem, this paper proposes an unsupervised domain adaptation method via regularized conditional alignment for FER, which adversarially reduces domain- and class-wise discrepancies while explicitly dealing with uncertainties within and across domain. Specifically, the proposed method effectively suppresses uncertainties in FER transfer tasks via: 1) semantics-preserving adaptation framework which enforces both domain-invariant learning and class-level semantic consistency between source and target expression data, where discriminative cluster structures are simultaneously retained; 2) auxiliary uncertainty regularization which further constrains the ambiguity of cluster boundaries to guarantee the transferring reliability, thus discouraging the negative transfer brought by divergent facial images. Evaluation experiments on publicly available datasets demonstrate that the proposed method significantly outperforms the current state-of-the-art methods.
Linyi Zhou, Xijian Fan, Tardi Tjahjadi, Qiaolin Ye
ACM Multimedia5
2020 Multi-view generalized support vector machine via mining the inherent relationship between views with applications to face and fire smoke recognition
Yawen Cheng, Liyong Fu, Qiaolin Ye, Fan Liu 0003
Knowl. Based Syst.4
2020 Robust discriminant feature selection via joint L2, 1-norm distance minimization and maximization
Zhangjing Yang, Qiaolin Ye, Qiao Chen 0004, Xu Ma 0005, Liyong Fu, Guowei Yang 0002, Fan Liu 0003
Knowl. Based Syst.2
2020 Positional Context Aggregation Network for Remote Sensing Scene Classification
abstract
To capture the long-range dependence of an input image for remote sensing scene (RSS) classification, in this letter, we propose a general positional context aggregation (PCA) module in deep convolutional neural networks. The PCA module is with the form of self-attention mechanism, in which two proposed blocks, the spatial context aggregation (SCA) and the relative position encoding (RPE), are used to capture the spatial-dipartite contextual aggregation information and the RPE information. Therefore, compared with the classical self-attention mechanism, global attention maps extracted by PCA not only have the advantage of regional distinction but also satisfy the translation equivariance that is proven to benefit scene classification. To demonstrate the superiority of the PCA module, we implement it on the pretrained ResNet [i.e., the so-called PCA network (PCANet)] and report the results on five popular RSS classification benchmarks. Experimental results show that the PCA module can improve the RSS classification performance significantly, and PCANet50 achieves the state-of-the-art results on these data sets.
Qiaolin Ye
IEEE Geosci. Remote. Sens. Lett.3
2020 Unsupervised deep triplet hashing with pseudo triplets for scalable image retrieval
Haofeng Zhang 0001, Zheng Zhang 0006, Qiaolin Ye
Multim. Tools Appl.4
2020 Improved multi-view GEPSVM via Inter-View Difference Maximization and Intra-view Agreement Minimization
Yawen Cheng, Qiaolin Ye, Liyong Fu, Zhangjing Yang
Neural Networks3
2020 Recursive Discriminative Subspace Learning With $\ell_{1}$ -Norm Distance Constraint
abstract
In feature learning tasks, one of the most enormous challenges is to generate an efficient discriminative subspace. In this paper, we propose a novel subspace learning method, named recursive discriminative subspace learning with an ℓ1-norm distance constraint (RDSL). RDSL can robustly extract features from the contaminated images and learn a discriminative subspace. With the use of an inequation-based ℓ1-norm distance metric constraint, the minimized ℓ1-norm distance metric objective function with slack variables induces samples in the same class to cluster as close as possible, meanwhile samples from different classes can be separated from each other as far as possible. By utilizing ℓ1-norm items in both the objective function and the constraint, RDSL can well handle the noisy data and outliers. In addition, the large margin formulation makes the proposed method insensitive to initializations. We describe two approaches to solve RDSL with a recursive strategy. Experimental results on six benchmark datasets, including the original data and the contaminated data, demonstrate that RDSL outperforms the state-of-the-art methods.
Yunlian Sun, Qiaolin Ye, Jinhui Tang 0001
IEEE Trans. Cybern.3
2020 Robust Triple-Matrix-Recovery-Based Auto-Weighted Label Propagation for Classification
abstract
The graph-based semisupervised label propagation (LP) algorithm has delivered impressive classification results. However, the estimated soft labels typically contain mixed signs and noise, which cause inaccurate predictions due to the lack of suitable constraints. Moreover, the available methods typically calculate the weights and estimate the labels in the original input space, which typically contains noise and corruption. Thus, the encoded similarities and manifold smoothness may be inaccurate for label estimation. In this article, we present effective schemes for resolving these issues and propose a novel and robust semisupervised classification algorithm, namely the triple matrix recovery-based robust auto-weighted label propagation framework (ALP-TMR). Our ALP-TMR introduces a TMR mechanism to remove noise or mixed signs from the estimated soft labels and improve the robustness to noise and outliers in the steps of assigning weights and predicting the labels simultaneously. Our method can jointly recover the underlying clean data, clean labels, and clean weighting spaces by decomposing the original data, predicted soft labels, or weights into a clean part plus an error part by fitting noise. In addition, ALP-TMR integrates the auto-weighting process by minimizing the reconstruction errors over the recovered clean data and clean soft labels, which can encode the weights more accurately to improve both data representation and classification. By classifying samples in the recovered clean label and weight spaces, one can potentially improve the label prediction results. Extensive simulations verified the effectivenss of our ALP-TMR.
Zhao Zhang 0001, Ming-Bo Zhao, Qiaolin Ye, Min Zhang 0005, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2019 Efficient and robust TWSVM classification via a minimum L1-norm distance metric criterion
Qiaolin Ye, Dongjun Yu
Mach. Learn.2
2019 A random-weighted plane-Gaussian artificial neural network
Xubing Yang, Fuquan Zhang 0004, Xijian Fan, Qiaolin Ye
Neural Comput. Appl.5
2019 Robust auto-weighted projective low-rank and sparse recovery for visual representation
Lei Wang 0124, Bangjun Wang, Zhao Zhang 0001, Qiaolin Ye, Liyong Fu, Guangcan Liu, Meng Wang 0001
Neural Networks4
2019 Robust capped L1-norm twin support vector machine
Chunyan Wang 0018, Qiaolin Ye, Ning Ye 0001, Liyong Fu
Neural Networks2
2019 Flexible non-greedy discriminant subspace feature extraction
Henghao Zhao, Liyong Fu, Qiaolin Ye, Zhangjing Yang, Xubing Yang
Neural Networks4
2019 Nonpeaked Discriminant Analysis for Data Representation
abstract
Of late, there are many studies on the robust discriminant analysis, which adopt L1-norm as the distance metric, but their results are not robust enough to gain universal acceptance. To overcome this problem, the authors of this article present a nonpeaked discriminant analysis (NPDA) technique, in which cutting L1-norm is adopted as the distance metric. As this kind of norm can better eliminate heavy outliers in learning models, the proposed algorithm is expected to be stronger in performing feature extraction tasks for data representation than the existing robust discriminant analysis techniques, which are based on the L1-norm distance metric. The authors also present a comprehensive analysis to show that cutting L1-norm distance can be computed equally well, using the difference between two special convex functions. Against this background, an efficient iterative algorithm is designed for the optimization of the proposed objective. Theoretical proofs on the convergence of the algorithm are also presented. Theoretical insights and effectiveness of the proposed method are validated by experimental tests on several real data sets.
Qiaolin Ye, Zechao Li, Liyong Fu, Zhao Zhang 0001, Wankou Yang, Guowei Yang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2018 Robust Adaptive Low-Rank and Sparse Embedding for Feature Representation
abstract
Most existing low-rank sparse embedding models extract features of data in the original input space and usually separate the manifold preservation step from the coding process, which may result in the decreased performance. In this paper, a novel Robust Adaptive Low-rank and Sparse Embedding (RALSE) framework is technically proposed for salient feature extraction of the high-dimensional data by seamlessly integrating the joint low-rank and sparse recovery with the robust adaptive salient feature extraction. Specifically, our RALSE integrates the joint low-rank and sparse representation, adaptive neighborhood preserving graph weight learning and the robustness-promoting representation into a unified framework. For accurate similarity measure, RALSE computes the adaptive weights by minimizing the reconstruction error over the noise-removed data and salient features simultaneously, where L1-norm is regularized to ensure the sparse properties of learnt weights. RALSE can also ensure the learnt projection to preserve local neighborhood information of embedded features clearly and adaptively. The projection is not only modeled under joint low-rank and sparse regularization, but also computed from a clean subspace, making it powerful for the salient feature extraction. Thus, the learnt low-rank sparse features would be more accurate for subsequent classification. Extensive results demonstrate the effectiveness of our RALSE formulation for data representation and classification.
Lei Wang 0124, Zhao Zhang 0001, Guangcan Liu, Qiaolin Ye, Jie Qin 0004, Meng Wang 0001
ICPR4
2018 Rotational Invariant Discriminant Subspace Learning For Image Classification
abstract
A novel discriminant analysis technique for feature extraction, referred to as Robust Discriminant Subspace (RDS) with L2,p+s-Norm Distance Maximization-Minimization (maxmin) is posed. In its objective, the within-class and between-class distances are measured by L2,p-norm and L2,s-norm, respectively, such that it is robust and rotational invariant. An efficient iterative algorithm is designed to solve the resulted objective, which is non-greedy. We also conduct some insightful analysis on the convergence of the proposed algorithm. Theoretical insights and effectiveness of our RDS are further supported by promising experimental results on several images databases.
Qiaolin Ye, Zhao Zhang 0001
ICPR1
2018 Robust Adaptive Label Propagation by Double Matrix Decomposition
abstract
In this paper, we investigate the robust transductive label prediction problem. Technically, a Robust Adaptive Label Propagation framework by Double Matrix Decomposition, called ALP-MD, is proposed for the semi-supervised data classification. Compared with existing transductive label propagation models, our ALP-MD improves the classification power by performing label prediction in the clean data space and clean label space at the same time. More specifically, our ALP-MD clearly integrates the idea of double matrix decomposition into the process of label prediction for the noise removal. Since the predicted soft labels usually contains noise and mixed signs, our ALP-MD explicitly decomposes the predicted soft label matrix into a clean soft label matrix and a noise term and then estimates the hard label based on the clean soft label matrix for more accurate classification. In addition, ALP-MD also involves a regularization term to model the noise in data, integrates the adaptive weights learning into the process of robust label prediction and moreover performs the weights learning in the clean data space. Thus, our ALP-MD can explicitly ensure the learned weights to be informative as much as possible and to be joint optimal for both representation and classification, and potentially enhance the label prediction ability. Extensive comparisons demonstrated its effectiveness.
Zhao Zhang 0001, Sheng Li 0001, Qiaolin Ye, Ming-Bo Zhao, Meng Wang 0001
ICPR4
2018 Graph regularized local self-representation for missing value imputation with applications to on-road traffic sensor data
Xiaobo Chen 0001, Yingfeng Cai, Qiaolin Ye, Lei Chen 0011
Neurocomputing3
2018 A discriminative dynamic framework for facial expression recognition in video sequences
Xijian Fan, Xubing Yang, Qiaolin Ye, Yin Yang 0002
J. Vis. Commun. Image Represent.3
2018 Lp- and Ls-Norm Distance Based Robust Linear Discriminant Analysis
Qiaolin Ye, Liyong Fu, Zhao Zhang 0001, Henghao Zhao, Meem Abdullah Naiem
Neural Networks1
2018 Adaptive non-negative projective semi-supervised learning for inductive classification
Zhao Zhang 0001, Lei Jia 0002, Ming-Bo Zhao, Qiaolin Ye, Min Zhang 0005, Meng Wang 0001
Neural Networks4
2018 A Feature Selection Method for Projection Twin Support Vector Machine
Rui Yan 0010, Qiaolin Ye, Liyan Zhang 0001, Xiangbo Shu
Neural Process. Lett.2
2018 L1-Norm GEPSVM Classifier Based on an Effective Iterative Algorithm for Classification
Qiaolin Ye, Tian'an Zhang, Dongjun Yu, Yiqing Xu
Neural Process. Lett.2
2018 Least squares twin bounded support vector machines based on L1-norm distance metric for classification
Qiaolin Ye, Tian'an Zhang, Dongjun Yu, Xia Yuan, Yiqing Xu, Liyong Fu
Pattern Recognit.2
2018 L1-Norm Distance Linear Discriminant Analysis Based on an Effective Iterative Algorithm
abstract
Recent works have proposed two L1-norm distance measure-based linear discriminant analysis (LDA) methods, L1-LD and LDA-L1, which aim to promote the robustness of the conventional LDA against outliers. In LDA-L1, a gradient ascending iterative algorithm is applied, which, however, suffers from the choice of stepwise. In L1-LDA, an alternating optimization strategy is proposed to overcome this problem. In this paper, however, we show that due to the use of this strategy, L1-LDA is accompanied with some serious problems that hinder the derivation of the optimal discrimination for data. Then, we propose an effective iterative framework to solve a general L1-norm minimization-maximization (minmax) problem. Based on the framework, we further develop a effective L1-norm distance-based LDA (called L1-ELDA) method. Theoretical insights into the convergence and effectiveness of our algorithm are provided and further verified by extensive experimental results on image databases.
Qiaolin Ye, Jian Yang 0003, Fan Liu 0003, Chunxia Zhao, Ning Ye 0001, Tongming Yin
IEEE Trans. Circuits Syst. Video Technol.1
2018 Underlying Connections Between Algorithms for Nongreedy LDA-L1
abstract
To solve the essential objective of LDA-L1, NLDA-L1 proposes a nongreedy algorithm by constructing an auxiliary function. In this correspondence, we show that essentially, this algorithm directly solves the objective using a gradient ascending procedure, meaning that the auxiliary function may be not necessary. Then, we further show that NLDA-L1 is a special case of ILDA-L1, which applies the same iterative procedure of ILDA-L1.
Qiaolin Ye, Henghao Zhao, Liyong Fu, Shangbing Gao
IEEE Trans. Image Process.1
2018 L1-Norm Distance Minimization-Based Fast Robust Twin Support Vector k-Plane Clustering
abstract
Twin support vector clustering (TWSVC) is a recently proposed powerful k-plane clustering method. It, however, is prone to outliers due to the utilization of squared L2-norm distance. Besides, TWSVC is computationally expensive, attributing to the need of solving a series of constrained quadratic programming problems (CQPPs) in learning each clustering plane. To address these problems, this brief first develops a new k-plane clustering method called L1-norm distance minimization-based robust TWSVC by using robust L1-norm distance. To achieve this objective, we propose a novel iterative algorithm. In each iteration of the algorithm, one CQPP is solved. To speed up the computation of TWSVC and simultaneously inherit the merit of robustness, we further propose Fast RTWSVC and design an effective iterative algorithm to optimize it. Only a system of linear equations needs to be computed in each iteration. These characteristics make our methods more powerful and efficient than TWSVC. We also conduct some insightful analysis on the existence of local minimum and the convergence of the proposed algorithms. Theoretical insights and effectiveness of our methods are further supported by promising experimental results.
Qiaolin Ye, Henghao Zhao, Zechao Li, Xubing Yang, Shangbing Gao, Tongming Yin, Ning Ye 0001
IEEE Trans. Neural Networks Learn. Syst.1
2017 Graph regularized multilayer concept factorization for data representation
Xiaobo Shen 0001, Zhenqiu Shu, Qiaolin Ye, Chunxia Zhao
Neurocomputing4
2016 Recursively global and local discriminant analysis for semi-supervised and unsupervised dimension reduction with image analysis
Shangbing Gao, Yunyang Yan, Qiaolin Ye
Neurocomputing4
2016 Recursive Dimension Reduction for semisupervised learning
Qiaolin Ye, Tongming Yin, Shangbing Gao, Jiajia Jing, Cui-Ping Sun
Neurocomputing1
2016 Comments on "Joint Global and Local Structure Discriminant Analysis"
abstract
Joint global and local structure discriminant analysis (JGLDA) is a recently-developed linear discriminant analysis approach, which considers the diversity of data across classes. In this communication, however, we will show that the discussion on the characterization of within-class locality and within-class diversity of previous efforts on local linear discriminant analysis (LLDA) and JGLDA is flawed.
Qiaolin Ye, Jiajia Jing
IEEE Trans. Inf. Forensics Secur.1
2016 Can the Virtual Labels Obtained by Traditional LP Approaches Be Well Encoded in WLR?
abstract
Semisupervised dimension reduction via virtual label regression first derives the virtual labels of unlabeled data by employing a newly designed label propagation (LP) approach (called Special random walk (SRW)) and then encodes them in a weighted linear regression model. Nie et al. (2011) highlighted two important characteristics of SRW nonexistent in the previous LP approaches: outlier detection and probability value output, which guarantee the elegant encoding of the resultant virtual labels in the weighted label regression. However, in this brief, we show that the relationship between the SRW and the previous work on LP is very close. Naturally, a problem deserving investigation is whether traditional LP approaches are indeed unable to share the above two characteristics of SRW. We aim to address this problem.
Qiaolin Ye, Jian Yang 0003, Tongming Yin, Zhao Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2015 Fast orthogonal linear discriminant analysis with application to image classification
Qiaolin Ye, Ning Ye 0001, Tongming Yin
Neurocomputing1
2014 Recursive soft margin subspace learning
abstract
In this paper, we propose a recursive soft margin (RSM) subspace learning framework for dimension reduction of high-dimensional data, which has strong recognition ability. RSM is motivated by the soft margin criterion of support vector machines (SVMs), which allows some training samples to be misclassified for a certain cost to achieve higher recognition results. Instead of maximizing the sum of squares of Euclidean interclass (called intracluster in unsupervised learning) pairwise distances over all the similar points in previous work, RSM seeks to maximize every pairwise interclass distance between two similar points, and this distance is represented in absolute. Then, we introduce a symmetrical Hingle loss function into the RSM framework. Doing so is to allow some pairwise interclass distances to violate the maximization constraint, such that we can get satisfactory classification performance by losing some training performance. To find multiple projection vectors, a recursive procedure is designed. Our framework is illustrated with Graph Embedding (GE). For any dimension reduction method expressible by the GE, it can thus be generalized by the proposed framework to boost their recognition power by reformulating the original problems.
Qiaolin Ye, Chunxia Zhao
IJCNN1
2014 Fast orthogonal linear discriminant analysis with applications to image classification
abstract
Orthogonalized variant of Linear Discriminant Analysisis (LDA) is an effective statistical learning tool for dimension reduction. However, existing orthogonalized LDA algorithms suffer from various drawbacks, including the requirement for expensive computing time. This paper develops an efficient algorithm for dimension reduction, referred to as Fast Orthogonal Linear Discriminant Analysis (FOLDA), which adopts an iterative procedure to extract the orthogonal projection vectors. Different from previous efforts, this new approach applies QR decomposition and regression to solve for a new projection vector in each time of iterations, leading to the by far cheaper computational cost. FOLDA can achieve comparable recognition rates to existing orthogonal LDA algorithms. Experimental results on image databases, such as MNIST, COIL20, MEPG-7, and OUTEX, show the effectiveness and efficiency of FOLDA.
Qiaolin Ye, Ning Ye 0001, Haofeng Zhang 0001, Chunxia Zhao
IJCNN1
2014 Feature selection for least squares projection twin support vector machine
Jianhui Guo, Ping Yi, Ruili Wang 0001, Qiaolin Ye, Chunxia Zhao
Neurocomputing4
2014 Flexible orthogonal semisupervised learning for dimension reduction with image classification
Qiaolin Ye, Ning Ye 0001, Chunxia Zhao, Tongming Yin, Haofeng Zhang 0001
Neurocomputing1
2014 Enhanced multi-weight vector projection support vector machine
Qiaolin Ye, Ning Ye 0001, Tongming Yin
Pattern Recognit. Lett.1
2012 Recursive robust least squares support vector regression based on maximum correntropy criterion
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004, Qiaolin Ye
Neurocomputing4
2012 Density-based weighting multi-surface least squares classification with its applications
Qiaolin Ye, Ning Ye 0001, Shangbing Gao
Knowl. Inf. Syst.1
2012 Smooth twin support vector regression
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004, Qiaolin Ye
Neural Comput. Appl.4
2012 Weighted Twin Support Vector Machines with Local Information and its application
Qiaolin Ye, Chunxia Zhao, Shangbing Gao
Neural Networks1
2012 Recursive "concave-convex" Fisher Linear Discriminant with applications to face, handwritten digit and terrain recognition
Qiaolin Ye, Chunxia Zhao, Haofeng Zhang 0001, Xiaobo Chen 0001
Pattern Recognit.1
2011 Distance difference and linear programming nonparallel plane classifier
Qiaolin Ye, Chunxia Zhao, Haofeng Zhang 0001, Ning Ye 0001
Expert Syst. Appl.1
2011 1-Norm least squares twin support vector machines
Shangbing Gao, Qiaolin Ye, Ning Ye 0001
Neurocomputing2
2011 Localized twin SVM via convex minimization
Qiaolin Ye, Chunxia Zhao, Xiaobo Chen 0001
Neurocomputing1
2011 Recursive projection twin support vector machine via within-class variance minimization
Xiaobo Chen 0001, Jian Yang 0003, Qiaolin Ye, Jun Liang 0004
Pattern Recognit.3
2010 Iterative support vector machine with guaranteed accuracy and run time
abstract
Abstract:Using a conjugate gradient method, a novel iterative support vector machine (FISVM) is proposed, which is capable of generating a new non‐linear classifier. We attempt to solve a modified primal problem of proximal support vector machine (PSVM) and show that the solution of the modified primal problem reduces to solving just a system of linear equations as opposed to a quadratic programming problem in SVM. This algorithm not only has no requirement for special optimization solvers, such as linear or quadratic programming tools, but also guarantees fast convergence. The full algorithm merely needs four lines of MATLAB codes, which gives results that are similar to or better than that of several new learning algorithms, in terms of classification accuracy. Besides, the proposed stand‐alone approach is capable of dealing with instability of classification performance of smooth support vector machine, generalized proximal support vector machine, PSVM and reduced support vector machine. Experiments carried out on UCI datasets show the effectiveness of our approach.
Qiaolin Ye, Chunxia Zhao, Yannan Chen
Expert Syst. J. Knowl. Eng.1
2010 Multi-weight vector projection support vector machines
Qiaolin Ye, Chunxia Zhao, Yannan Chen
Pattern Recognit. Lett.1