Yiyou Guo

dblp:190/4792 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-3497-6939ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Context-Enhanced Zero-Shot Video Temporal Grounding with Adaptive Boundary Refinement
abstract
In this paper, we introduce a novel training-free framework for Video Temporal Grounding (VTG) that combines pre-trained Visual Language Models (VLMs) and Large Language Models (LLMs). Existing methods often struggle with capturing the semantics of natural language queries and identifying the dynamic transitions at event boundaries. To address these challenges, our approach uses VLMs to generate detailed contextual descriptions of video content, providing richer prompts for LLMs to understand and reason about event temporal relations. Furthermore, we introduce an adaptive event boundary refinement strategy, ensuring better coverage of the full event phases. Our framework demonstrates superior performance in zero-shot settings on several benchmark datasets, including Charades-STA and ActivityNet Captions, and exhibits remarkable robustness in out-of-distribution (OOD) scenarios.
Fangkai Li, Feiyu Pan, Yiyou Guo, Xiankai Lu
ICME5
2025 Remote Sensing Scene Classification via Pseudo-Category-Relationand Orthogonal Feature Learning
abstract
Remote sensing (RS) scene classification is a crucial component in the analysis of Earth observation data, aiding in a deeper understanding and monitoring of our dynamic planet. Its applications extend across various fields, including land management, urban analysis, and environmental monitoring. The complex semantic information in RS scene images and the relationships between different scene categories present significant challenges to improving scene classification tasks. Unlike previous methods that only focus on network structure or feature encoding, our approach emphasizes the association of scene categories, integrating feature learning and knowledge transfer together to enhance the analysis of scenes at a higher semantic level. To this end, we propose an RS scene classification scheme based on pseudo-scene category-relation reasoning and orthogonal feature (OF) learning modules, capturing the inherent semantic connections among diverse scene classes. Additionally, we introduce cascaded attention (CA) and selected separation modules to strategically optimize the network, targeting challenging classes with high feature similarities. Knowledge is then distilled across different branches, guiding to enhance the model’s robustness and prediction accuracy. Experiments are conducted on three challenging RS scene datasets of AID30, UCMerced21, and NWPU-RESISC45 to validate the effectiveness of the learned pseudo-category relationships. The results demonstrate that the proposed framework outperforms existing hierarchical approaches in leveraging the hierarchical structure of RS scene images.
Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Gongping Yang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 From Coarse to Fine: Learning Semantic Relations for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) is consisted of many narrow spectral bands which are capable of recording abundant features including both the spectral and spatial signatures information which have been widely used in various fields, such as urban planning, disaster monitoring. Due to the large number and high similarity of the collected spectral brands, many methods are developed to handle the problem of extracting effective features. Recently, many CNN-based methods have been proposed by exploiting the spectral-spatial signatures of the HSIs data and achieved promising results. Although many methods adopt patch-based input pattern to emphasize the importance of the spatial neighbor information of each pixel, the relations are still limited to a small area around the pixel and the latent relations among the pixels belonging to different semantic categories at the boundary are still not well exploited. To explore the relationships between pixels from a more global perspective, a neighbor-based relation mining framework is proposed to explore the long-range relations among different local regions. Experiments are conducted on two hyperspectral image classification datasets and the results demonstrate the effectiveness of the proposed long-range relations mining scheme by comparison with some state-of-the-art methods.
Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Huan Xie 0001
IGARSS4
2023 From Coarse to Fine: Knowledge Distillation for Remote Sensing Scene Classification
abstract
Scene classification is one of the most commonly studied areas of parsing the earth observation data. How to effectively interpreting the remote sensing images and extracting informative features are the great challenges for remote sensing image classification. Many important applications, such as land management and urban analysis, are based on the performance of remote sensing classification model. Recently, a lot of CNN based methods have been proposed and achieve promising results. Inspired by the success of knowledge distillation which transfers the learned information from a teacher model to a student model, a knowledge distillation based framework is proposed in this paper to handle the task of remote sensing scene classification from coarse to fine. Specifically, the learned knowledge from the teacher network is transformed into the coarse soft label and fine output mask to better guiding the student network to learn more informative features. Experiments are conducted on two widely used remote sensing scene datasets to evaluate the effectiveness of the proposed method and achieve comparable results compared with some state-of-the-art methods.
Jinsheng Ji, Xiaoming Xi, Xiankai Lu, Yiyou Guo, Huan Xie 0001
IGARSS4
2022 Dpnet: end-to-end Aerial Image Segmentation Via Deformable Point Network
abstract
Aerial image Segmentation segmentation faces intrinsic foreground-background imbalance and background clutter distraction. To guide the segmentation model to learn more discriminative foreground ability and more invariant back-ground representation features, we design a Deformable Point Network (DPNet). It is an end-to-end segmentation network and consists of a multi-head deformable attention module that simultaneously considers foreground object information and background suppression. Specifically, we first employ a feature pyramid network to aggregate multiple-layer features to handle scale variants. And then, we further investigate deformable convolution to select some representative points for each layer and propose a differential module to implement it automatically instead of traditional dense fusion. Moreover, we incorporate the multi-head mechanism in the feature fusion to focus on the key contents from different representation regions. Experimental results on the representative iSAID, Vaihingen, and Postdam datasets demonstrate that our DPNet achieves competitive performance. Also, the multiple-head deformable attention facilitates the network convergence significantly.
Yiyou Guo, Zheyun Qin, Yongtai Yang, Xiankai Lu, Huan Xie 0001, Xiaohua Tong
IGARSS1
2022 An Anchor-Free Network With Density Map and Attention Mechanism for Multiscale Object Detection in Aerial Images
abstract
Accurate detection of the multiple classes in aerial images has become possible with the use of anchor-based object detectors. However, anchor-based object detectors place a large number of preset anchors on images and regress the target bounding box while anchor-free object detections predict the location of objects directly and avoid the carefully predefined anchor box parameters. Object detection in aerial images is faced with two main challenges: 1) the scale diversity of the geospatial objects; and 2) the cluttered background in complex scenes. In this letter, to address these challenges, we present a novel Anchor-Free Network with a Density map and attention mechanism (DA2FNet). Considering the extreme density variations of the detection instances among the different categories in aerial images, the proposed DA2FNet model conducts density map estimation with image-level supervision for the geospatial object counting, to acquire global knowledge about the scale information. A simple and effective image-level global counting loss function is also introduced. In addition, a compositional attention network is further introduced to enhance the saliency of the foreground objects. The proposed DA2FNet method was compared with the state-of-the-art object detection models, achieving excellent performance on the NWPU VHR-10, RSOD, and DOTA datasets.
Yiyou Guo, Xiaohua Tong, Xiong Xu 0001, Sicong Liu 0001, Yongjiu Feng, Huan Xie 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 CubeNet: X-shape connection for camouflaged object detection
Mingchen Zhuge, Xiankai Lu, Yiyou Guo, Zhihua Cai, Shuhan Chen
Pattern Recognit.3
2021 Multi-level dictionary learning for fine-grained images categorization with attention model
Jinsheng Ji, Yiyou Guo, Zhen Yang 0012, Tao Zhang 0027, Xiankai Lu
Neurocomputing2
2021 MFALNet: A Multiscale Feature Aggregation Lightweight Network for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Semantic segmentation labels each pixel in high-resolution remote sensing (HRRS) images with a category. To tackle with the large size and complexity of HRRS images, this letter presents a novel multiscale feature aggregation lightweight network (MFALNet) for semantic segmentation. Unlike standard convolution, asymmetric depth-wise separable convolution residual (ADCR) unit is used to reduce the parameter size of the network and makes the optimized structure deeper but lightweight and less complex. The proposed network is an encoder–decoder structure, where multiscale feature aggregation is implemented in both the encoder and the decoder. The spatial self-attention block helps to capture long-range contextual information, and the gated convolution modules are further used for refining features when aggerating high- and low-level feature maps in the decoder. The proposed MFALNet has evaluated on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam 2-D semantic labeling contest open benchmark data set, and the experimental results prove that the scheme can obtain a better tradeoff between segmentation accuracy and computational efficiency compared with the state-of-the-art semantic segmentation models.
Yiyou Guo, Tengfei Bao, Chenqin Fu
IEEE Geosci. Remote. Sens. Lett.2
2021 Multi-view feature learning for VHR remote sensing image classification
Yiyou Guo, Jinsheng Ji, Qiankun Ye, Huan Xie 0001
Multim. Tools Appl.1
2020 Geospatial Object Detection with Single Shot Anchor-Free Network
abstract
Geospatial object detection has made considerable progress with the use of anchor-based object detectors. In such a situation, the detection performance relies heavily on the parameter settings of anchor boxes. We present a Single Shot Anchor-Free Network (SSAFNet) to tackle with this problem. By eliminating the anchor boxes, the SSAFNet completely avoids the carefully predefined anchor boxes parameters and the computation for adapting the huge scale variation of geospatial objects. A compositional attention network is further introduced to enhance the saliency of foreground objects. We evaluate the SSAFNet on the representative NWPU VHR-10 and RSOD datasets, achieving competitive performance with state-of-the-art anchor-based detection methods.
Yiyou Guo, Jinsheng Ji, Xiankai Lu, Huan Xie 0001, Xiaohua Tong
IGARSS1
2020 A survey on deep learning methods for scene flow estimation
Ruihong Wu, Qingyun Zhao, Yiyou Guo, Long Chen 0005
Pattern Recognit.5
2019 AggregationNet: Identifying Multiple Changes Based on Convolutional Neural Network in Bitemporal Optical Remote Sensing Images
Qiankun Ye, Xiankai Lu, Lihong Wan, Yiyou Guo
PAKDD (3)5
2018 Non-convex joint bilateral guided depth upsampling
Xiankai Lu, Yiyou Guo, Na Liu 0007, Lihong Wan
Multim. Tools Appl.2
2017 Supervised Adaptive Incremental Clustering for data stream of chunks
Laiwen Zheng, Yiyou Guo
Neurocomputing3