EDBT 2026 Demo / reviewers in the wild / expert
Guanyi Li
dblp:38/8385
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-2790-9962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Image recognition and object detection · 31% Video understanding and tracking · 15% Motion planning and robot control · 15% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
camouflaged object detection |
1.0 | 1 | 2026 | Learning Compact Representations With an Information Bottleneck for Camouflaged Object Detection · IEEE Trans. Multim. 2026 |
Computer vision › Image recognition and object detection › object detection › infrared object detection
infrared small target detection |
1.0 | 1 | 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion · IEEE Trans. Image Process. 2026 |
Computer vision › Video understanding and tracking
motion analysis |
1.0 | 1 | 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion · IEEE Trans. Image Process. 2026 |
Robotics › Motion planning and robot control › robot control
motion compensation |
1.0 | 1 | 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion · IEEE Trans. Image Process. 2026 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.8 | 1 | 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene Classification · IEEE Trans. Multim. 2024 |
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition |
0.8 | 1 | 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene Classification · IEEE Trans. Multim. 2024 |
Computer vision › Image recognition and object detection › scene recognition
remote sensing scene classification |
0.8 | 1 | 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene Classification · IEEE Trans. Multim. 2024 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.3 | 1 | 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion · IEEE Trans. Image Process. 2026 |
Methods — techniques the papers use, named apart from their topics
information bottleneck · 2.0shifted neighborhood compensation · 1.0progressive distillation · 1.0frequency-domain feature analysis · 1.0cross-domain interaction · 1.0parameter-efficient fine-tuning · 0.8multimodal foundation model · 0.8frequency distribution · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language ModelsabstractOpen-vocabulary camouflaged object segmentation (OVCOS) seeks to segment and classify camouflaged objects in arbitrary categories, presenting unique challenges due to visual ambiguity and unseen categories. Recent approaches typically adopt a two-stage paradigm: they first segment objects, and then classify the segmented regions using vision language models (VLMs). However, such methods (i) suffer from a domain gap caused by the mismatch between VLMs' full-image training and cropped-region inferencing, and (ii) depend on generic segmentation models optimized for well-delineated objects which are less effective for camouflaged objects. Without explicit guidance, generic segmentation models often overlook subtle boundaries, leading to imprecise segmentation. In this paper, we introduce a novel VLM-guided cascaded framework to address these issues in OVCOS. For segmentation, we leverage the segment anything model (SAM), guided by the VLM. Our framework uses VLM-derived features as explicit prompts to SAM, effectively directing attention to camouflaged regions and significantly improving localization accuracy. For classification, we avoid the domain gap introduced by hard cropping. Instead, we treat the segmentation output as a soft spatial prior using the alpha channel. This retains the full image context while providing precise spatial guidance, leading to more accurate and context-aware classification of camouflaged objects. The same VLM is shared between segmentation and classification to ensure efficiency and semantic consistency. Extensive experiments on both OVCOS and conventional camouflaged object segmentation benchmarks demonstrate the clear superiority of our method, highlighting the effectiveness of leveraging rich VLM semantics for both segmentation and classification of camouflaged objects. Our code and models are open-sourced at https://github.com/intcomp/camouflaged-vlm. Kai Zhao 0012, Wubang Yuan, Zheng Wang 0059, Guanyi Li, Xiaoqiang Zhu, Deng-Ping Fan, Dan Zeng 0001 |
Comput. Vis. Media | 4 |
| 2026 | MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex MotionabstractMotion cues play a vital role in multi-frame infrared small target detection (MISTD). However, most targets in existing datasets exhibit regular and slow motion, which cannot reflect the complex and diverse motion patterns in real-world scenarios. This biased data distribution makes recent data-driven methods highly rely on simplified motion assumptions that tend to fail in irregular or fast motion, resulting in noisy feature representations cluttered with target-irrelevant factors. Hence, we stress that methods for MISTD should also work when targets are in complex motion. To enable this research, we propose a large-scale dataset called MIST for airborne infrared detection scenarios. The dataset is built on a synthetic data engine that models variations in pose, size, and intensity of moving targets while seamlessly blending them into real backgrounds for physical, geometric, and visual realism. Targets in MIST exhibit low signal-to-clutter ratios and complex motion, making it a promising yet challenging benchmark for developing algorithms focused on motion analysis. To tackle the challenges of MIST, we develop MISTNet, a robust baseline based on the Information Bottleneck theory. To handle irregular and fast motion, we propose a shifted neighborhood compensation block to efficiently model multi-scale correspondences for implicit motion compensation. To distill compact representations free from irrelevant cues, we design a progressive distillation decoder to hierarchically filter out redundancy while preserving target-relevant information. We benchmark 31 state-of-the-art methods and find that their performance on MIST drops significantly compared with that on the widely used NUDT-MIRSDT dataset. Our MISTNet outperforms all other methods by a large margin, with an over 6% gain in the IoU metric, demonstrating its superiority. The dataset, code, and model weights are available at https://github.com/GR-ray/MIST. Meihong Zhang, Gongyang Li, Guanyi Li, Kai Zhao 0012, Xianchao Zhang 0002, Dan Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Learning Compact Representations With an Information Bottleneck for Camouflaged Object DetectionabstractFrequency domain-based methods have demonstrated promising performance in Camouflaged Object Detection (COD) tasks because of their enhanced power for distinguishing between objects and the background in the frequency domain. However, these methods often overlook the interference caused by task-irrelevant cues such as background textures. These extraneous factors are learned alongside task-relevant features by the employed network, increasing the number of false positives. Therefore, we propose a camouflaged object detection method based on the Information Bottleneck (IB) theory. The aim is to obtain a robust representation that retains the essential features needed for prediction while minimizing the redundant information derived from both the RGB and frequency domains. Specifically, we propose a Feature Selection Information Bottleneck Module (FSIBM). By explicit supervision, this module minimizes the mutual information between the fused feature from two domains and the predictive features, thereby weakening task-irrelated information. Simultaneously, the FSIBM maximizes the mutual information between the predictive features and the ground truth (i.e., emphasizing task-related elements). Additionally, we introduce a Cross-Domain Awareness Interaction Module (CDAIM), which establishes self-reinforcement for the object attributes within each domain and facilitates cross-domain complementarity. This enables the capture of sufficient discriminative features from both domains. To verify the generalization ability of the proposed method, we applied it to three benchmark datasets, on which our method outperformed the corresponding state-of-the-art methods. Our code is released athttps://github.com/KwunYat/CODIB. Guanyi Li, Junjie Zhang 0002, Wubang Yuan, Gloria Jin, Dan Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Learning Instructive Frequency Spectral and Curvature Features for Cloud DetectionabstractCurrent cloud detection methods often treat all spectral bands equally, which limits their ability to capture instructive clues necessary for accurate detection. As a result, distinguishing clouds from snow in coexisting environments remains challenging. Moreover, most approaches struggle to adaptively model the boundaries of clouds, which is crucial for detecting thin clouds with ambiguous edges. To address these challenges, we propose a novel approach for cloud detection called FSCFNet, which captures guiding visual features from frequency and curvature computations. FSCFNet comprises two key modules: the Frequency Spectral Feature Enhancement Module (FSFEM) and the Curvature-based Edge-Awareness Module (CEAM). The FSFEM leverages the distinct characteristics of spectral bands to extract instructive visual cues, enabling the network to learn robust discriminative features for ice, snow, and clouds. In contrast, the CEAM adaptively identifies texture-rich regions using curvature, enhancing the ability to delineate thin cloud boundaries. Comprehensive quantitative and qualitative experiments on the Landsat 8 and MODIS datasets demonstrate that FSCFNet consistently outperforms state-of-the-art methods. Our code is publicly available at https://github.com/wanjuanhu/FSCFNet/tree/main. Wanjuan Hu, Guanyi Li, Guoguo Zhang, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Cold SegDiffusion: A novel diffusion model for medical image segmentation
Minglei Li 0002, Jiusi Zhang, Guanyi Li, Yuchen Jiang 0001, Hao Luo 0003 |
Knowl. Based Syst. | 4 |
| 2024 | Leveraging Frequency-Guided Mixer and Target-Aware Attention for Ground-Based Cloud DetectionabstractCompared to satellite imagery, ground-based cameras capture cloud data (ground-to-sky data) with higher temporal and spatial resolutions, providing more detailed cloud information. However, the spectral information available in ground-to-sky data is limited. Therefore, extracting features with strong discrimination from optical remote sensing images (ORSIs) is challenging. Currently, deep learning-based cloud detection methods face two main challenges. Firstly, although Convolutional Neural Networks (CNNs) effectively extract high-frequency (HF) components from images through convolutions, they struggle to capture low-frequency (LF) components, which are capable of representing global features and target structures. Secondly, in ORSIs, the spectral characteristics of thin clouds and the sky are similar, making it difficult to distinguish cloud regions from the background. To address these challenges, we propose a network consisting of two main modules: the Mixer Module (MM) and the Cloud Aware Attention Module (CAAM). The MM comprises a HF and a LF components extraction branch. The HF branch extracts local textures through max-pooling and parallel convolution operations. The LF branch captures long-range dependency by decomposing a large kernel convolution. It leverages the advantages of both convolution and self-attention to effectively capture global features. In addition, we introduce the CAAM, which quantifies images into histograms to separate clouds from the background and enhances the perception of clouds using attention mechanism. We conducted experiments using both daytime and nighttime cloud image data from the SWINySeg dataset with mIoU reaching 88.93% and OA reaching 93.97%. The results demonstrate that our proposed method achieves promising performance compared to state-of-the-art cloud detection methods. Chenyu Dong, Guanyi Li, Yixiao Gu, Junjie Zhang 0002, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Multi-Level Information Fusion Network With Edge Information Injection for Single-Band Cloud DetectionabstractCurrent cloud detection methods have demonstrated effectiveness by utilizing the rich spectral features of multi-spectral images. Compared to multispectral images, single-band infrared images offer higher efficiency in terms of sampling and processing speed. However, single-band cloud detection methods have not been fully developed, and existing methods based on multispectral cloud detection have some limitations when applied directly to single-band images: Firstly, they often blend shallow features containing spatial details with deep features providing high-level semantic information, yet struggle to disentangle features with strong discrimination representing cloud edges and bodies from limited information. Additionally, the correlation between features at different aspects is not fully reasoned, resulting in blurred boundary segmentation. To address these issues, we introduce a Multi-level Information Fusion Network (MIFNet) with an integrated edge information injection strategy. Our method effectively decouples clouds into their fundamental components: body and edge (Low-Frequency (LF) and High-Frequency (HF) components), enabling the comprehensive acquisition of strong discriminative features. Specifically, we propose an Edge Feature Extraction Module (EFEM) that isolates the cloud body through low-pass filtering, while the cloud’s edge is extracted by subtracting lower-level features from LF components. Furthermore, we employ a Feature Refinement Module (FRM) to locate the cloud body’s position precisely. Building upon this foundation, we devise a Graph Reasoning Module (GRM) to facilitate the full inference of feature correlations at different levels and to model the global interdependence between edges and semantics. Through comprehensive evaluations on benchmark datasets comprising infrared band images from Landsat 8 and MODIS satellites, we demonstrate that our proposed MIFNet outperforms state-of-the-art methods, yielding promising results in cloud detection accuracy. Our code is publicly available at https://github.com/KwunYat/MIFNet. Guanyi Li, Junjie Zhang 0002, Enquan Yang, Dan Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene ClassificationabstractFew-shot open-set recognition, as a new paradigm, leveraging a limited amount of supervised data to identify specific Remote Sensing (RS) scene categories and generalize to novel ones. However, the data bias induced by the small sample size not only causes severe overfitting within base classes, but also impairs the capacity for inference to identify RS scenes in hitherto unobserved categories. Furthermore, owing to environmental influences, RS images frequently manifest notable intra-class disparities and comparatively low inter-class distinctions, intensifying the challenge in obtaining suitable classifiers. To address above issues, we investigate the utilization of a Multi-modal Foundational Model (MFM) infused with essential domain knowledge to mitigate the generalization limitations encountered in few-shot scenarios. Recognizing that existing MFMs with a visual-text dual-branch structure are primarily tailored for natural scenes, we propose a custom Frequency Distribution-based Multi-modal Fine-Tuning strategy (FreqDiMFT) in a parameter-efficient manner. More specifically, within the vision branch, we address the high inter-class similarity and intra-class diversity in RS images by embedding the local-global frequency distribution information to facilitate the recognition of RS scenes. To further amplify the model's generalization ability post transfer, we introduce an adaptive feature refinement module designed for Transformers, proficient in filtering redundant features resulting from domain disparities. To mitigate the domain drift on the textual branch, we adopt an input format that combines basic templates with domain expertise from RS end to generate more discriminative class prototypes. To fully verify the effectiveness of our FreqDiMFT in a more practical setting, we collect a Large-Scale hybrid dataset (LSRS). Extensive experiments demonstrate that, even with a scant number of training samples, our strategy yields advanced performances compared to state-of-the-art models. Junjie Zhang 0002, Yutao Rao, Xiaoshui Huang, Guanyi Li, Dan Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | A Pyramid Attention Network With Edge Information Injection for Remote-Sensing Object DetectionabstractRemote sensing images (RSIs) are often characterized by the high spatial resolution, strong object scale effects, and complex scenes, which poses great challenges to the object detection. Although mainstream neural network-based methods work well in detecting common objects, they often fail to fully exploit the detailed structural information in the spatial domain, leading to the poor performance for objects with diverse scales and distributions under complicated backgrounds. To address the above issue, we propose a pyramid attention network with edge information injection for remote sensing object detection. Considering each object is composed of the inner body and outer profile parts that corresponding to the low and high frequency components of image respectively, the difference between the original image and its low frequency component is beneficial for obtaining the high frequency counterpart. We design the Edge Information Extraction Module (EIEM) to mine the detailed edge features at multiple scales, and subsequently inject them into features at corresponding scales in the backbone network. As for promoting the performance in complex scenes, we introduce a Pyramid Feature Fusion (PFF) module, which leverages both local and global attention for establishing the long-range channel dependency, thereby highlighting objects that need to be concentrated on. To verify the effectiveness of our proposed method, we conduct extensive experiments on DIOR and RSOD datasets with mean Average Precision (mAP) reaching 74.93% and 96.44% respectively, demonstrating that our model achieved SOTA performance compared to mainstream methods. Junjie Zhang 0002, Anqi Ding, Guanyi Li, Liangang Zhang, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Scenario Context-Aware-Based Bidirectional Feature Pyramid Network for Remote Sensing Target DetectionabstractCompared with ordinary optical images, the situation of remote sensing images is much more complicated. The problems caused by the shooting angles over the Earth’s surface are: 1) some target categories with more complex shooting environments greatly increase the difficulty of detection and 2) the remote sensing images with large and small targets at the same time leading to large changes in the target scale are difficult to handle. In this letter, we designed a novel scenario context-aware-based bidirectional feature pyramid network (SCBi-FPN) to address the above problems. There are two key modules of the proposed network: the scene context-aware module uses pyramid pooling to aggregate contextual information of the different regions to obtain better global contextual information. The bidirectional feature pyramid network (Bi-FPN) module with squeeze and excitation (SE) blocks connects feature layers at different scales in a cross-scale manner and performs weighted feature map fusion before passing through the SE blocks to enable the network to obtain more accurate information. The experiments demonstrate that our designed network has good results compared with the state-of-the-art methods. In particular, we achieved mean average precision (mAP) of 92.92 on the publicly available NWPU VHR-10 dataset. Guanyi Li, Baohua Jin, Qiqiang Chen, Junru Yin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Possibilistic Rules from Fuzzy Prototypes
Guanyi Li, Jonathan Lawry |
IPMU (2) | 1 |
| 2005 | Remote sensing and GIS in runoff coefficient estimation in Binjiang BasinabstractRemote sensing and GIS technologies were integrated to estimate runoff for Binjiang basin, China. Remotely sensed images from the Landsat satellites were used to develop land cover maps of the study area for the years 1990, 1995 and 2000. GIS analysis, based on land-cover and soil map data, was used to estimate SCS curve numbers on a 30-meter grid and to compute runoff depths for a 10 -year maximum rainfall event. Runoff coefficients were computed by the Rational method for each year of the study. Temporal changes in spatial distribution of land cover, runoff coefficients, and runoff volume were estimated. A runoff hydrograph, based on the above information and digital elevation model, was developed for the study area. The land cover for the study region showed an increase in urban area and a corresponding decrease in vegetation area. A increase in modeled runoff volume and peak flow is attributed to this change in land cover. The results of this study indicate a need for protecting vegetation to reduce the effect of floods in this basin. Yulin Zhan, Changyao Wang, Zheng Niu, Pifu Cong, Guanyi Li |
IGARSS | 5 |