EDBT 2026 Demo / reviewers in the wild / expert
Liyong Fu
dblp:155/9293
· DBLP profile ↗
42ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0002-5794-9458ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Instance-Aware Visual Prompting helps multimodal models see better
Jingxu Wang, Liyong Fu, Qiaolin Ye |
Expert Syst. Appl. | 3 |
| 2026 | Learning Robust Discriminant Projections via Double Capped Lp-Norm Distance Metrics With "Min" ConstraintsabstractRecently, there has been a surge in the development of robust norm distance-based linear discriminant analysis (LDA) techniques, which have garnered significant attention in the field of feature extraction. However, a persistent issue that has yet to be resolved is that the successful suppression of outliers may inadvertently impede the accurate discrimination of normal points. To solve this problem, we, in this article, study a novel robust LDA measured by double capped $L_{p}$ -norm distance (CLD) metrics with min constraints (DCLDA) to learn robust discriminant projections, in which normal points and outliers are separately treated. To be specific, it takes a double capped $L_{p}$ -norm with "Min" constraints in the proposed model to measure the distances for between- and within-class dispersions. The proposed model effectively ensures accurate discrimination of normal points by $L_{p}$ -norm, while also eliminating the exaggerated effect of outliers that may arise from larger $p$ values. The resulted objective is not trivial because of its nonconvexity and nonsmoothness. As one of the major contributions of this article, we introduce a new reformulation that provides an objective problem theoretically equivalent to the original. By this reformulation, we develop an effective iterative algorithm to solve the proposed model. The algorithm is proven to be convergent through rigorous theoretical analysis. Extensive experiments were conducted on several real-world datasets across different image classification tasks to showcase the effectiveness of the proposed method. Xiaobo Chen 0001, Zhao Zhang 0001, Liyong Fu, Qiaolin Ye |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Interpretable deep one-class model for forest fire detection
Yangjie Xu, Yiran Ma, Qiaolin Ye, Liyong Fu, Xubing Yang |
Expert Syst. Appl. | 4 |
| 2025 | Synthetic instance segmentation from semantic image segmentation masks
Zhao Zhang 0001, Liyong Fu, Qiaolin Ye |
Knowl. Based Syst. | 4 |
| 2025 | Non-rigid object detection via fast one-class model
Xubing Yang, Jingyao Lishen, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu |
Pattern Recognit. | 6 |
| 2025 | Robust Multiple Flat Projections Clustering With Truncated Distance Maximization ConstraintsabstractRecently, interest in flat-type projection clustering methods has grown as they improve learner's performance by exploring multiple projection subspaces. However, solvers used in previous representative works predominantly rely on greedy search strategies, which incur high computational costs and fail to consider interdependencies between projections. Moreover, these methods do not simultaneously guarantee the effective suppression of outliers and noisy data at cluster boundaries, ultimately compromising data discrimination. To address these limitations and discover a more effective subspace for each flat, we propose robust multiple flat projections clustering (RMFPC). This method computes within- and between-cluster distances using the L2,1-norm to enhance robustness against outliers. Furthermore, we propose a truncated distance maximization constraint (TDMC) to eliminate the influence of noisy data on cluster separability. The resulting objective is presented in a ratio form, which is not trivial. We provide a novel formulation to achieve a theoretically equivalent problem. Based on this reformulation, we develop an efficient non-greedy solution algorithm. In addition, a cluster center optimization mechanism is incorporated into the solution process to accurately estimate the distribution of each cluster center. The convergence analysis and proof of the proposed algorithm are provided. Experiments on both toy and real-world datasets demonstrate the effectiveness of the proposed method. Zhao Zhang 0001, Xiaobo Chen 0001, Zhongqi Xu, Liyong Fu, Qiaolin Ye |
IEEE Trans. Cybern. | 5 |
| 2025 | Local-Global Information Perception Network for Salient Object Detection in Optical Remote Sensing ImagesabstractIn the field of salient object detection (SOD), optical remote sensing images (ORSI) differ significantly from natural sensing images (NSI). Existing research in ORSI-based SOD is constrained by the limitations of convolutional neural networks (CNNs) in feature extraction and by the underutilization of feature information in Transformer-based approaches. To address these challenges, this paper presents a Transformer-based Local-Global Information Perception Network (LGIPNet) for ORSI, which enhances encoder-generated features at multiple levels to highlight salient targets through three specialized feature enhancement modules. The Edge Adaptive Enhancement Module (EAEM) focuses on extracting local edge features to guide precise edge generation. The Multi-scale Grouped Weighted Attention Module (MGWAM) extracts local information from low-level features, scales features, and uses multi-scale channel and learnable weighted spatial attention to locate salient targets. For high-level features, the Dual-Domain Attention Module (DDAM) integrates a Channel Enhancement Attention Block (CEA) and an Adaptive Spatial Attention Block (ASA) to refine both local and global information. Specifically, the EAEM sharpens the edges of salient objects, thereby ensuring the precision of boundary detection. The MGWAM, on the other hand, enriches the feature representation across multiple scales, enhancing the network’s capability to encapsulate both fine-grained details and broader contextual information. The DDAM further strengthens the balance between local and global information, preserving feature integrity across levels. Finally, multi-scale features are cascaded to produce the final saliency map. This holistic strategy empowers LGIPNet to accurately identify and emphasize salient objects in ORSI. Experiments on three datasets demonstrate that LGIPNet outperforms existing state-of-the-art methods, establishing its effectiveness and robustness in ORSI-based SOD. The source code is available at https://github.com/sCauliflower/LGIPNet.git. Le Sun 0002, Hongxin Liu, Yuhui Zheng, Qiao Chen 0004, Zebin Wu 0001, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Convergence Analysis on Trace Ratio Linear Discriminant Analysis AlgorithmsabstractLinear discriminant analysis (LDA) may yield an inexact solution by transforming a trace ratio problem into a corresponding ratio trace problem. Most recently, optimal dimensionality LDA (ODLDA) and trace ratio LDA (TRLDA) have been developed to overcome this problem. As one of the greatest contributions, the two methods design efficient iterative algorithms to derive an optimal solution. However, the theoretical evidence for the convergence of these algorithms has not yet been provided, which renders the theory of ODLDA and TRLDA incomplete. In this correspondence, we present some rigorously theoretical insight into the convergence of the iterative algorithms. To be specific, we first demonstrate the existence of lower bounds for the objective functions in both ODLDA and TRLDA, and then establish proofs that the objective functions are monotonically decreasing under the iterative frameworks. Based on the findings, we disclose the convergence of the iterative algorithms finally. Qiaolin Ye, Liyong Fu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Validation and Evaluation of the Three-Dimensional Analytical Radiative Transfer Model ESRT in Heterogeneous Forest CanopyabstractThe complicated spatial heterogeneity in forest scenes is an important challenge in radiative transfer modeling. To overcome the excessive simplification of forest scenes by classical analytical models and the efficiency limitations of computer simulation models, we developed a three-dimensional (3D) analytical radiative transfer model called ESRT (Stochastic Radical Transfer model for forests with hEterogeneous canopy structure) by extending the stochastic radiative transfer theory. At present, the application of ESRT in different complex forest scenes still needs to be explored, and more field data is essential for model verification. In this study, we validated the performance of ESRT in simulating different kinds of heterogeneous forest canopy reflectance based on survey data from 21 mixed forest plots and 50 pest-stress forest plots, and evaluated the effect of mixing and pest levels on forest canopy reflectance. The results showed that compared to the original SRT model, the extended ESRT model can more accurately simulate the canopy reflectance of mixed forests and pest-damaged forests, showing better consistency with the measured spectra from the sample plots. The conifer-broadleaf ratio and the vertical distribution of damaged foliage can both affect the canopy spectral signals. This study provides a theoretical basis for applications of ESRT in complex forest simulation. Zhuoli Zhang, Bingxiang Tan, Liyong Fu |
IGARSS | 4 |
| 2024 | Global superpixel-merging via set maximum coverage
Xubing Yang, Zhengxiao Zhang, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | RemoteCLIP: A Vision Language Foundation Model for Remote SensingabstractGeneral-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these models primarily learn low-level features and require annotated data for fine-tuning. Moreover, they are inapplicable for retrieval and zero-shot applications due to the lack of language understanding. To address these limitations, we propose RemoteCLIP, the first vision-language foundation model for remote sensing that aims to learn robust visual features with rich semantics and aligned text embeddings for seamless downstream application. To address the scarcity of pre-training data, we leverage data scaling which converts heterogeneous annotations into a unified image-caption data format based on Box-to-Caption (B2C) and Mask-to-Box (M2B) conversion. By further incorporating UAV imagery, we produce a 12 × larger pretraining dataset than the combination of all available datasets. RemoteCLIP can be applied to a variety of downstream tasks, including zero-shot image classification, linear probing,k-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. Evaluation on 16 datasets, including a newly introduced RemoteCount benchmark to test the object counting ability, shows that RemoteCLIP consistently outperforms baseline foundation models across different model scales. Impressively, RemoteCLIP beats the state-of-the-art method by 9.14% mean recall on the RSITMD dataset and 8.92% on the RSICD dataset. For zero-shot classification, our RemoteCLIP outperforms the CLIP baseline by up to 6.39% average accuracy on 12 downstream datasets. Fan Liu 0003, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Qiaolin Ye, Liyong Fu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Multiscale 3-D-2-D Mixed CNN and Lightweight Attention-Free Transformer for Hyperspectral and LiDAR ClassificationabstractThe effective combination of hyperspectral image (HSI) and light detection and ranging (LiDAR) data can be utilized for land cover classification. Recently, deep learning-based classification methods, especially those utilizing Transformer networks, have achieved remarkable success. However, deep learning classification methods for multi-source data still encounter various technical challenges, such as the comprehensive utilization of multi-scale information, the lightweight network design, and the efficient fusion strategies for heterogeneous data. To address these challenges, we propose a novel and efficient deep neural network, namely multi-scale 3D-2D mixed CNN feature extraction and multi-source data lightweight attention-free fusion network (M2FNet) based on CNN and Transformer. Through end-to-end training, this network effectively combines heterogeneous information from multiple sources, leading to improved performance in joint classification. Specifically, M2FNet employs a multi-scale 3D-2D mixed CNN design to extract both the spatial-spectral features of HSI and the depth-based elevation features of LiDAR data. Subsequently, the extracted features are fed into a novel encoder comprising a feature enhancement module, designed with mathematical morphology and a dilated convolutional module derived from the self-attention of the conventional Transformer encoder (DConvformer), which plays a crucial role in integrating multi-source information within the network. The well-designed architecture enables the network to acquire multi-scale depth and high-order features, significantly reducing the number of training parameters. Comparative experimental results and ablation studies demonstrate that M2FNet outperforms other advanced methods. The source code is publicly available at https://github.com/cupid6868/M2FNet.git. Le Sun 0002, Yuhui Zheng, Zebin Wu 0001, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MDC-FusFormer: Multiscale Deep Cross-Fusion Transformer Network for Hyperspectral and Multispectral Image FusionabstractThe spatial resolution of hyperspectral images (HSIs) is usually limited due to internal imaging mechanisms. To obtain imagery with high spectral and high spatial resolutions, which is essential for subsequent HSI processing tasks, a cost-effective approach is to fuse HSI with multispectral images (MSIs). One highly effective fusion method is the convolutional neural network (CNN). However, CNNs have limitations in capturing global information and complex features. Recently, visual transformers (ViTs) have garnered interest for their ability to process non-local information. Despite this, existing HSI-MSI fusion methods suffer from insufficient spatial-spectral feature interaction, resulting in suboptimal fusion quality. To address these challenges, we propose a multiscale deep cross-fusion transformer (MDC-FusFormer) network for HSI and MSI fusion. This network effectively performs the interactive fusion of spatial-spectral features, thereby enhancing the quality of the fused images. MDC-FusFormer employs a three-branch network architecture consisting of two independent progressive feature mining modules (PFMMs), a multiscale deep cross-fusion attention module, and a spatial-spectral feature fusion module. Initially, shallow features at different scales of MSI and HSI are recursively extracted through successive up- and down-sampling using CNNs. These features then interact with the deep cross-modal information at corresponding scales through the attention block. Finally, a multidimensional refinement convolution block (MRCB) is applied to refine the feature information, which is then combined with cascaded up-sampling to reconstruct the high-resolution fused image step by step. Experimental results on five datasets indicate that, compared to nine other methods, MDC-FusFormer delivers superior performance. Le Sun 0002, Jianxiao Zhou, Qiaolin Ye, Zebin Wu 0001, Qiao Chen 0004, Zhongqi Xu, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Video Moment Retrieval With Noisy LabelsabstractVideo moment retrieval (VMR) aims to localize the target moment in an untrimmed video according to the given nature language query. The existing algorithms typically rely on clean annotations to train their models. However, making annotations by human labors may introduce much noise. Thus, the video moment retrieval models will not be well trained in practice. In this article, we present a simple yet effective video moment retrieval framework via bottom-up schema, which is in end-to-end manners and robust to noisy label training. Specifically, we extract the multimodal features by syntactic graph convolutional networks and multihead attention layers, which are fused by the cross gates and the bilinear approach. Then, the feature pyramid networks are constructed to encode plentiful scene relationships and capture high semantics. Furthermore, to mitigate the effects of noisy annotations, we devise the multilevel losses characterized by two levels: a frame-level loss that improves noise tolerance and an instance-level loss that reduces adverse effects of negative instances. For the frame level, we adopt the Gaussian smoothing to regard noisy labels as soft labels through the partial fitting. For the instance level, we exploit a pair of structurally identical models to let them teach each other during iterations. This leads to our proposed robust video moment retrieval model, which experimentally and significantly outperforms the state-of-the-art approaches on standard public datasets ActivityCaption and textually annotated cooking scene (TACoS). We also evaluate the proposed approach on the different manual annotation noises to further demonstrate the effectiveness of our model. Wenwen Pan 0003, Zhou Zhao 0001, Wencan Huang, Liyong Fu, Jun Yu 0002, Fei Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Preferred vector machine for forest fire detection
Xubing Yang, Zhichun Hua, Li Zhang 0057, Xijian Fan, Fuquan Zhang 0004, Qiaolin Ye, Liyong Fu |
Pattern Recognit. | 7 |
| 2023 | CRNet: Channel-Enhanced Remodeling-Based Network for Salient Object Detection in Optical Remote Sensing ImagesabstractDespite the remarkable progress made by the salient object detection of natural sensing images (NSI-SOD), the complex background and scale diversity issues of remote sensing images (RSIs) still pose a substantial obstacle. In this study, we build an end-to-end channel-enhanced remodeling-based network (CRNet) for optical RSIs (ORSIs) to highlight salient objects through feature augmentation. First, the backbone convolutional block is used to suggest the fundamental characteristics. Then, we use the channel enhance module (CEM) to enhance the shallow features. CEM primarily relies on the channel attention mechanism and employs a no-downscaling strategy to produce local cross-channel interaction, which lowers model complexity while enhancing extraction performance. Meanwhile, we use the redefined feature module (RFM) to reconstruct the deep features and generate global attention features by dimensional transformation and feature relationship aggregation to achieve the role of locating salient targets. Finally, the cascade combines the multi-scale features to provide the final saliency map. To further enhance the representational power of the network, we use a hybrid loss function to improve performance. The proposed approach outperforms current state-of-the-art methods, as shown by several experiments on three available datasets. The source code of the proposed CRNet is available publicly at https://github.com/hilitteq/CRNet.git. Le Sun 0002, Yuwen Chen 0001, Yuhui Zheng, Zebin Wu 0001, Liyong Fu, Byeungwoo Jeon |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Robust ensemble method for short-term traffic flow prediction
Liyong Fu, Yong Qi 0002, Dongjun Yu, Qiaolin Ye |
Future Gener. Comput. Syst. | 2 |
| 2022 | Flexible capped principal component analysis with applications in image recognition
Liyong Fu, Qiaolin Ye |
Inf. Sci. | 2 |
| 2022 | Learning discriminative and representative feature with cascade GAN for generalized zero-shot learning
Jingren Liu, Liyong Fu, Haofeng Zhang 0001, Qiaolin Ye, Wankou Yang, Li Liu 0004 |
Knowl. Based Syst. | 2 |
| 2022 | Learning a robust classifier for short-term traffic state prediction
Liyong Fu, Yong Qi 0002, Qiaolin Ye, Dongjun Yu |
Knowl. Based Syst. | 2 |
| 2022 | Multi-view distance metric learning via independent and shared feature subspace with applications to face and forest fire recognition, and remote sensing classification
Liyong Fu, Yawen Cheng, Qiaolin Ye |
Knowl. Based Syst. | 2 |
| 2022 | Robust distance metric optimization driven GEPSVM classifier for pattern classification
Liyong Fu, Tian'an Zhang, Jun Hu 0010, Qiaolin Ye, Yong Qi 0002, Dongjun Yu |
Pattern Recognit. | 2 |
| 2022 | Unabridged adjacent modulation for clothing parsing
Chengting Zuo, Qianhao Wu, Liyong Fu, Xinguang Xiang |
Pattern Recognit. | 4 |
| 2022 | MMatch: Semi-Supervised Discriminative Representation Learning for Multi-View ClassificationabstractSemi-supervised multi-view learning has been an important research topic due to its capability to exploit complementary information from unlabeled multi-view data. This work proposes MMatch, a new semi-supervised discriminative representation learning method for multi-view classification. Unlike existing multi-view representation learning methods that seldom consider the negative impact caused by particular views with unclear classification structures (weak discriminative views). MMatch jointly learns view-specific representations and class probabilities of training data. The representations concatenated to integrate multiple views’ information to form a global representation. Moreover, MMatch performs the smoothness constraint on the class probabilities of the global representation to improve pseudo labels, whereas the pseudo labels regularize the structure of view-specific representations. A discriminative global representation is mined with the training process, and the negative impact of weak discriminative views is overcome. Besides, MMatch learns consistent classification while preserving diverse information from multiple views. Experiments on several multi-view datasets demonstrate the effectiveness of MMatch. Xiaoli Wang 0003, Liyong Fu, Yudong Zhang 0001, Yongli Wang 0002, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Multiview Learning With Robust Double-Sided Twin SVMabstractMultiview learning (MVL), which enhances the learners' performance by coordinating complementarity and consistency among different views, has attracted much attention. The multiview generalized eigenvalue proximal support vector machine (MvGSVM) is a recently proposed effective binary classification method, which introduces the concept of MVL into the classical generalized eigenvalue proximal support vector machine (GEPSVM). However, this approach cannot guarantee good classification performance and robustness yet. In this article, we develop multiview robust double-sided twin SVM (MvRDTSVM) with SVM-type problems, which introduces a set of double-sided constraints into the proposed model to promote classification performance. To improve the robustness of MvRDTSVM against outliers, we take L1-norm as the distance metric. Also, a fast version of MvRDTSVM (called MvFRDTSVM) is further presented. The reformulated problems are complex, and solving them are very challenging. As one of the main contributions of this article, we design two effective iterative algorithms to optimize the proposed nonconvex problems and then conduct theoretical analysis on the algorithms. The experimental results verify the effectiveness of our proposed methods. Qiaolin Ye, Zhao Zhang 0001, Yuhui Zheng, Liyong Fu, Wankou Yang |
IEEE Trans. Cybern. | 5 |
| 2022 | BASNet: Burned Area Segmentation Network for Real-Time Detection of Damage Maps in Remote Sensing ImagesabstractSince remote sensing images of post-fire vegetation are characterized by high resolution, multiple interferences, and high similarities between the background and the target area, it is difficult for existing methods to detect and segment the burned area in these images with sufficient speed and accuracy. In this paper, we apply Salient Object Detection (SOD) to burned area segmentation, the first time this has been done, and propose an efficient burned area segmentation network (BASNet) to improve the performance of unmanned aerial vehicle (UAV) high-resolution image segmentation. BASNet comprises positioning module and refinement module. The positioning module efficiently extracts high-level semantic features and general contextual information via global average pooling layer and convolutional block to determine the coarse location of the salient region. The refinement module adopts the convolutional block attention module to effectively discriminate the spatial location of objects. In addition, to effectively combine edge information with spatial location information in the lower layer of the network and the high-level semantic information in the deeper layer, we design the residual fusion module to perform feature fusion by level to obtain the prediction results of the network. Extensive experiments on two UAV datasets collected from Chongli in China and Andong in South Korea, demonstrate that our proposed BASNet significantly outperforms state-of-the-art SOD methods quantitatively and qualitatively. BASNet also achieves a promising prediction speed for processing high-resolution UAV images, thus providing wide-ranging applicability in post-disaster monitoring and management. Weihao Bo, Xijian Fan, Tardi Tjahjadi, Qiaolin Ye, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Learning Robust Discriminant Subspace Based on Joint L₂, ₚ- and L₂, ₛ-Norm Distance Metricsabstract-norm as the distance metric. However, both of their robustness and discriminant power are limited. In this article, we present a new robust discriminant subspace (RDS) learning method for feature extraction, with an objective function formulated in a different form. To guarantee the subspace to be robust and discriminative, we measure the within-class distances based on [Formula: see text]-norm and use [Formula: see text]-norm to measure the between-class distances. This also makes our method include rotational invariance. Since the proposed model involves both [Formula: see text]-norm maximization and [Formula: see text]-norm minimization, it is very challenging to solve. To address this problem, we present an efficient nongreedy iterative algorithm. Besides, motivated by trace ratio criterion, a mechanism of automatically balancing the contributions of different terms in our objective is found. RDS is very flexible, as it can be extended to other existing feature extraction techniques. An in-depth theoretical analysis of the algorithm's convergence is presented in this article. Experiments are conducted on several typical databases for image classification, and the promising results indicate the effectiveness of RDS. Liyong Fu, Zechao Li, Qiaolin Ye, Qingwang Liu, Xiaobo Chen 0001, Xijian Fan, Wankou Yang, Guowei Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Pixel-level automatic annotation for forest fire image
Xubing Yang, Run Chen, Fuquan Zhang 0004, Li Zhang 0057, Xijian Fan, Qiaolin Ye, Liyong Fu |
Eng. Appl. Artif. Intell. | 7 |
| 2021 | Recurrent Thrifty Attention Network for Remote Sensing Scene RecognitionabstractThe self-attention mechanism has been empirically shown its effectiveness in a wide range of computer vision applications. However, it is usually criticized for the expensive computation cost. Although some revised methods are proposed in the recent past, they are not maturely applicable to remote sensing scene (RSS) images. To address this problem, in this article, we propose a simple yet effective context acquisition module, named thrifty attention, which can capture the long-range dependence efficiently and effectively. Moreover, a recurrent version for thrifty attention, termed recurrent thrifty attention (RTA), is further proposed to take the long-range multihop communications in space–time for RSS images. RTA is a general global contextual information acquisition module that can be used in any hierarchy of deep convolutional neural networks. To demonstrate its superiority, we deploy it to the classical ResNet and establish our proposed RTA Network (RTANet). Extensive experiments are carried out on two levels of the RSS recognition tasks, i.e., the image-level RSS classification and the instance-level RSS object detection. Compared with the standard self-attention mechanism, RTA can reduce at most 0.43 M model parameters while increasing a slight of model floating-point operations per second (FLOPs). Furthermore, results on RSS classification and object detection further verify the accuracy superiority of RTANet. Liyong Fu, Qiaolin Ye |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Dominant Trees Analysis Using UAV LiDAR and PhotogrammetryabstractDominant trees compose the upper story of forest canopies, and are one of the key factors that affect the light redistribution for forest ecosystem. UAV lidar and photogrammetry can be used to measure spatial variation of upper crowns of dominant trees with details. How about the differences of tree crowns using lidar and photogrammetry with very different observation geometry? This paper aims to extract parameters of dominant trees and analyze the differences using lidar and photogrammetry. The lidar and photogrammetry-based CHMs were smoothed by Gaussian algorithm for three times. The result indicated that the positions of potential tops of dominant trees from lidar-based CHM were near to that from photogrammetry. The mean and standard deviation of position differences were 0.5m and 0.3m respectively. The heights of potential tops from lidar-based CHM were highly correlated with that from photogrammetry-based CHM. The height differences had the mean of -0.4m and standard deviation of 0.4m. The smoothing of crowns will weaken the effects of height variation within crowns on detection of dominant trees. Qingwang Liu, Xin Tian 0005, Liyong Fu |
IGARSS | 4 |
| 2020 | Multi-view generalized support vector machine via mining the inherent relationship between views with applications to face and fire smoke recognition
Yawen Cheng, Liyong Fu, Qiaolin Ye, Fan Liu 0003 |
Knowl. Based Syst. | 2 |
| 2020 | Robust discriminant feature selection via joint L2, 1-norm distance minimization and maximization
Zhangjing Yang, Qiaolin Ye, Qiao Chen 0004, Xu Ma 0005, Liyong Fu, Guowei Yang 0002, Fan Liu 0003 |
Knowl. Based Syst. | 5 |
| 2020 | Improved multi-view GEPSVM via Inter-View Difference Maximization and Intra-view Agreement Minimization
Yawen Cheng, Qiaolin Ye, Liyong Fu, Zhangjing Yang |
Neural Networks | 5 |
| 2020 | Improving Estimation of Forest Canopy Cover by Introducing Loss Ratio of Laser Pulses Using Airborne LiDARabstractForest canopy cover (CC) directly and indirectly influences various processes of forest ecosystems. Airborne light detection and ranging (LiDAR) can be used to characterize forest spatial structures and further obtain estimates of forest CC. However, nonreturn laser pulses from targets of interest impact the estimation accuracy of forest CC. The objective of this article was to develop a novel method of estimating the nonreturn laser pulses to improve the estimation accuracy of forest CC using LiDAR data. The improved models for estimating forest CC were developed by introducing the loss ratio of laser pulses and a CC coefficient into the original models. The forest CC reference data were collected and used to validate the forest CC estimates. The results show that the loss ratio for forested areas was much higher than that for open ground areas. The range between the sensor and a target was a crucial factor that caused the loss of returns. The relationship between the range and the loss ratio was nonlinear in both open ground and forested areas. Compared with the original models, the improved models combining the loss ratio and the CC coefficient statistically significantly increased the estimation accuracy of the forest CC. Moreover, the forest CC estimates from the canopy height model (CHM) were more accurate than those from the height normalized point cloud (NPC) data. In addition, the simplified models were more generalized than the other models. This article is novel and has great potential to improve mapping of forest CC. Qingwang Liu, Liyong Fu, Guangxing Wang 0003, Zengyuan Li, Erxue Chen, Yong Pang 0002, Kailong Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Robust auto-weighted projective low-rank and sparse recovery for visual representation
Lei Wang 0124, Bangjun Wang, Zhao Zhang 0001, Qiaolin Ye, Liyong Fu, Guangcan Liu, Meng Wang 0001 |
Neural Networks | 5 |
| 2019 | Robust capped L1-norm twin support vector machine
Chunyan Wang 0018, Qiaolin Ye, Ning Ye 0001, Liyong Fu |
Neural Networks | 5 |
| 2019 | Flexible non-greedy discriminant subspace feature extraction
Henghao Zhao, Liyong Fu, Qiaolin Ye, Zhangjing Yang, Xubing Yang |
Neural Networks | 2 |
| 2019 | Nonpeaked Discriminant Analysis for Data RepresentationabstractOf late, there are many studies on the robust discriminant analysis, which adopt L1-norm as the distance metric, but their results are not robust enough to gain universal acceptance. To overcome this problem, the authors of this article present a nonpeaked discriminant analysis (NPDA) technique, in which cutting L1-norm is adopted as the distance metric. As this kind of norm can better eliminate heavy outliers in learning models, the proposed algorithm is expected to be stronger in performing feature extraction tasks for data representation than the existing robust discriminant analysis techniques, which are based on the L1-norm distance metric. The authors also present a comprehensive analysis to show that cutting L1-norm distance can be computed equally well, using the difference between two special convex functions. Against this background, an efficient iterative algorithm is designed for the optimization of the proposed objective. Theoretical proofs on the convergence of the algorithm are also presented. Theoretical insights and effectiveness of the proposed method are validated by experimental tests on several real data sets. Qiaolin Ye, Zechao Li, Liyong Fu, Zhao Zhang 0001, Wankou Yang, Guowei Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | How trees allocate carbon for optimal growth: insight from a game-theoretic modelabstractHow trees allocate photosynthetic products to primary height growth and secondary radial growth reflects their capacity to best use environmental resources. Despite substantial efforts to explore tree height-diameter relationship empirically and through theoretical modeling, our understanding of the biological mechanisms that govern this phenomenon is still limited. By thinking of stem woody biomass production as an ecological system of apical and lateral growth components, we implement game theory to model and discern how these two components cooperate symbiotically with each other or compete for resources to determine the size of a tree stem. This resulting allometry game theory is further embedded within a genetic mapping and association paradigm, allowing the genetic loci mediating the carbon allocation of stemwood growth to be characterized and mapped throughout the genome. Allometry game theory was validated by analyzing a mapping data of stem height and diameter growth over perennial seasons in a poplar tree. Several key quantitative trait loci were found to interpret the process and pattern of stemwood growth through regulating the ecological interactions of stem apical and lateral growth. The application of allometry game theory enables the prediction of the situations in which the cooperation, competition or altruism is an optimal decision of a tree to fully use the environmental resources it owns. Liyong Fu, Lidan Sun, Libo Jiang, Meixia Ye, Shouzheng Tang, Minren Huang, Rongling Wu |
Briefings Bioinform. | 1 |
| 2018 | Lp- and Ls-Norm Distance Based Robust Linear Discriminant Analysis
Qiaolin Ye, Liyong Fu, Zhao Zhang 0001, Henghao Zhao, Meem Abdullah Naiem |
Neural Networks | 2 |
| 2018 | Least squares twin bounded support vector machines based on L1-norm distance metric for classification
Qiaolin Ye, Tian'an Zhang, Dongjun Yu, Xia Yuan, Yiqing Xu, Liyong Fu |
Pattern Recognit. | 7 |
| 2018 | Underlying Connections Between Algorithms for Nongreedy LDA-L1abstractTo solve the essential objective of LDA-L1, NLDA-L1 proposes a nongreedy algorithm by constructing an auxiliary function. In this correspondence, we show that essentially, this algorithm directly solves the objective using a gradient ascending procedure, meaning that the auxiliary function may be not necessary. Then, we further show that NLDA-L1 is a special case of ILDA-L1, which applies the same iterative procedure of ILDA-L1. Qiaolin Ye, Henghao Zhao, Liyong Fu, Shangbing Gao |
IEEE Trans. Image Process. | 3 |