VLDB 2026 Research / reviewers in the wild / expert
Fangli Guan
dblp:276/8070
· DBLP profile ↗
13ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-7409-2129ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chain-of-Search: Parameter-Efficient Reasoning for Zero-Shot Object NavigationabstractZero-shot object navigation tasks agents with locating target objects in unseen environments—a core capability of embodied intelligence. While recent vision-language navigation methods leverage Large Language Models (LLMs) for multimodal reasoning, they suffer from two key limitations: (1) semantic misalignment between language-grounded maps and real-world layouts, and (2) inefficiency due to LLMs’ lack of specialization for navigation-specific tasks. To address these challenges, we propose Chain-of-Search (CoS), a novel parameter-efficient framework that enables human-like decision-making via iterative semantic reasoning. First, CoS replaces traditional global maps with an optimal-benefit multi-map construction that continuously balances expected gain and cost throughout the navigation process. Second, we introduce a Parameter-Efficient Intent Aligner (PEIA), trained via a prompt-guided paradigm to align directional decisions with navigation intent. PEIA injects semantic cues into benefit-aware maps, enabling more rational and goal-consistent exploration. Finally, a Reflection-Guided Destination Verifier (RDV) confirms whether the target is reached via language-driven reasoning and corrects potential errors through self-reflection. CoS achieves state-of-the-art performance on HM3D (+2.8% SR) and MP3D (+1.2% SR) without relying on LLMs, demonstrating the effectiveness of lightweight, reasoning-centered navigation. Hanrui Chen, Liqi Yan, Qifan Wang 0001, Fangli Guan, Pan Li 0001 |
AAAI | 5 |
| 2026 | Benefit-cost frontier-aware semantic reasoning for zero-shot object navigation
Hanrui Chen, Liqi Yan, Qifan Wang 0001, Fangli Guan, Pan Li 0001 |
Appl. Intell. | 5 |
| 2025 | HiSSC: Hierarchical Scene-Aware Interaction for Indoor Semantic Scene CompletionabstractSemantic Scene Completion (SSC) stands as a critical 3D perception task that enable applications in indoor scene fine modeling, robot indoor autonomous navigation, AR/VR and other interactive applications. But it still faces unique challenges in indoor scenarios: complex vertical structures, severe occlusions, and functional dependencies between objects. Therefore, we propose HiSSC, a Hierarchical Scene-Aware Interaction Framework for Indoor Semantic Completion. Specifically, we introduce (1) Hierarchical Birds'-Eye-View Modeling; (2) Refined Complementary Interaction to better model complex environments. Validation tests on NYUv2 datatset show that our method successfully distinguishes stacked objects in complex indoor environment and overcomes the shortcomings of existing methods. Yunzhan Fu, Enyu Bao, Zao Hu, Danni Zhang, Fangli Guan |
CW | 5 |
| 2025 | STaR: Multi-Granular Spatio-Temporal Reasoning for Long-Form Dense Video CaptioningabstractDense video captioning is crucial for enhancing video understanding in daily applications and presents a significant challenge in multimodal analysis. Existing methods often overlook video-to-dynamic-space mapping at varying scales, resulting in captions that lack specificity and remain overly general, failing to capture real-world physical detail. To address this limitation, we propose a multi-granularity Spatio-Temporal Reasoning (STaR) approach, which integrates: (i) efficient global feature integration to model long-term temporal dependencies, (ii) spatial attention mechanisms with position encoding to capture absolute spatial information, and (iii) cross-modal feature fusion to align and unify global, local, and spatial representations. Moreover, we enhance the framework using a Large Language Model (LLM) to improve the richness and naturalness of the generated descriptions. Comparative experiments have been conducted to evaluate the effectiveness of the proposed method on SoccerNet dataset. Experimental results demonstrate that our model effectively enhances localization accuracy and generates captions with superior temporal and spatial detail fidelity. The code is available at https://github.com/bread-555/STaR. Chenhuan Cai, Liqi Yan, Huapeng Li, Qifan Wang 0001, Fangli Guan, Pan Li 0001 |
ECAI | 8 |
| 2025 | Optimal Distributed Training With Co-Adaptive Data Parallelism in Heterogeneous EnvironmentsabstractThe computational power required for training deep learning models has been skyrocketing in the past decade as they scale with big data, and has become a very expensive and scarce resource. Therefore, distributed training, which can leverage distributed available computational power, is vital for efficient large-scale model training. However, most previous distributed training frameworks like DDP and DeepSpeed are primarily designed for co-located clusters under homogeneous computing and communication conditions, and hence cannot account for geo-distributed clusters with both computing and communication heterogeneity. To address this challenge, we develop a new data parallel based distributed training framework called Co-Adaptive Data Parallelism (C-ADP). First, we consider a data owner and parameter server that distributes data to and coordinates the collaborative learning across all the computing devices. We employ local training and delayed parameter synchronization to reduce communication costs. Second, we formulate a data parallel scheduling optimization problem to minimize the training time by optimizing data distribution. Third, we devise an efficient algorithm to solve this scheduling problem, and formally prove that the obtained solution is optimal in the asymptotic sense. Experiments on the ImageNet100 dataset demonstrate that C-ADP achieves fast convergence in heterogeneous distributed training environments. Compared to Distributed Data Parallel (DDP) and DeepSpeed, C-ADP achieves 21.6 times and 26.3 times improvements in FLOPS, respectively, and a reduction in training time of about 72% and 47%, respectively. Lifang Chen, Zhichao Chen 0002, Liqi Yan, Yanyu Cheng, Fangli Guan, Pan Li 0001 |
IJCAI | 5 |
| 2025 | F-DDIM: A Featurized Denoising Diffusion Implicit Model for Facial Image SteganographyabstractFacial image steganography is crucial for privacy-preserving media transmission. Traditional embedding methods degrade image quality and are vulnerable to steganalysis, while GAN-based non-embedding approaches lack controllability and realism. Diffusion-based methods using textual prompts face two key issues: (1) security risks from interpretable prompts and (2) poor preservation of facial details. This paper presents Featurized Denoising Diffusion Implicit Models (F-DDIM), a novel non-embedding steganography framework. First, F-DDIM replaces explicit textual prompts with implicit image-based encoding, enhancing security. Second, it selectively refines facial regions for natural and high-quality recovery through iterative reconstruction. Third, it enables indistinguishable encryption without secret key sharing via a novel sub-code embedding algorithm. Fourth, a refinement step post-decoding improves the clarity and accuracy of recovered facial image details. Experimental results demonstrate that F-DDIM achieves superior image fidelity and robustness against transmission interference. Liqi Yan, Xuebin Li, Fangli Guan, Kanglei Peng, Pan Li 0001 |
ACM Multimedia | 4 |
| 2025 | BuildingSAM: A Dual-Branch Feature-Augmented Segment Anything Model for Remote Sensing Building ExtractionabstractWe propose a segment anything model for building extraction (BuildingSAM) as a general solution for building extraction (BE) from high-resolution remote-sensing images. Unlike previous methods, BuildingSAM is constructed based on the Segment Anything Model (SAM) for large-scale images and parameter-efficient fine-tuning (PEFT), representing a novel research paradigm for BE. Although transformer-based architecture excels at processing global and low-frequency information, it can overlook local details and introduce biases in feature learning. To address these shortcomings, BuildingSAM incorporates a dual-branch module, in which one branch employs the image encoder of the lightweight next-generation semantic segmentation network SegNeXt to focus on capturing local building details, and the other uses a vision transformer (ViT) image encoder to extract global features. Furthermore, we introduce the Conv-LoRA method, which integrates ultra-lightweight convolutional parameters into low-rank adaptation (LoRA) to inject image-related inductive biases into the ViT image encoder. This approach enhances the ability of BuildingSAM to learn building boundary features in complex regions efficiently. In experiments conducted using two BE benchmark datasets, the proposed method significantly outperformed existing state-of-the-art BE models. Comprehensive ablation studies further validated the superior performance of BuildingSAM, at minimal additional computational cost. Wenqing Feng, Fangli Guan, Jihui Tu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Enhanced YOLOv8 by integrating MSFF and MSAFM for UAV-Perspective Small Object DetectionabstractDetecting small objects is a formidable challenge in earth observation and remote sensing. To address the complexities associated with small object recognition in remote sensing imagery, this paper introduces the Multi-Scale Attention Feature Map (MSAFM) method. This approach effectively merges high-level and low-level feature maps by exploiting local features from the low-level maps and global features from the high-level maps, thereby substantially enhancing the fused feature map. Furthermore, we propose a feature map enhancement technique named Multi-Scale Feature Fusion (MSFF), which builds a feature pyramid within the same feature map using convolutional kernels of varying sizes and employs a bidirectional sampling strategy across different temporal sequences to augment the feature representation. When integrated into the YOLOv8 framework and evaluated on the VisDrone2019 dataset, these modules yielded notable improvements in accuracy, with mAP50 increasing by 11% and mAP50-95 by 7%, along with significant enhancements in detection accuracy across all categories. Guijun Chen, Fangli Guan |
CW | 2 |
| 2024 | Integrating Depth-Anything-V2 Depth Estimation and Sobel Operator Matrix for UAV Landing-site detectionabstractRescue UAVs can support important tasks such as rapid response and manoeuvrability, search and rescue, monitoring and assessment, and delivery of relief materials and communication. However, in natural disasters and accidents scenarios, there are difficulties in the autonomous selection of UAV landing sites. This study proposes a landing point detection method that considers UAV attitude and fuses visual depth estimation with matrix computation. Considering image local visual features and global features, the highlight is to propose a multi-layer feature fusion framework, including visual depth estimation module based on Depth-Anything-V2, planar matrix estimation based on Sobel operator, and optimal landing point detection considering UAV attitude. At present, the initial validation has been completed at UseGeo, and the landing point selection is as expected. Zao Hu, Xuan Rong, Danni Zhang, Zhang Xu, Fangli Guan |
CW | 6 |
| 2024 | A ConvNeXt-based Spatial-Enhanced Attention and Convolution Combination for Cross-View Geo-LocalizationabstractCross-view matching of satellite and drone images can address the reliability and safety of navigation in global navigation satellite system(GNSS) denied environments. Visual understanding is a promising research topic that can improve the accuracy of cross-view matching. However, most of the existing methods ignore the connection of contextual information in the visual image. Therefore, we propose a novel ConvNeXt-based spatial-enhanced attention and convolution combination model for solving this task. In addition, we use a new hybrid loss function to better limit these features and assist in mining crucial data regarding global features. The results show that the R @ 1 and AP criteria on the University-1652 dataset have achieved 90.01% and 91.78% advanced performance in drone target matching, and 95.02% and 89.90% competitive capability in drone navigation. Danni Zhang, Xuan Rong, Fangli Guan |
CW | 5 |
| 2024 | Autonomous wireless positioning system using crowdsourced Wi-Fi fingerprinting and self-detected FTM stations
Fangli Guan, Kexin Tang, Sheng Bao, Liang Chen 0007, Ruizhi Chen, Yue Yu 0003 |
Expert Syst. Appl. | 1 |
| 2024 | Road-SAM: Adapting the Segment Anything Model to Road Extraction From Large Very-High-Resolution Optical Remote Sensing ImagesabstractWe propose road-segment anything model (SAM), a universal model for extracting roads from large, very-high-resolution (VHR), optical, remote sensing (RS) images. Unlike previous methods, Road-SAM builds upon the foundation of the SAM, a large-scale image-segmentation model, to explore a new paradigm for customizable road extraction (RE). Within the framework, we introduce three variants that allow for flexible insertion of adapters at different positions within the transformer block. Additionally, the model employs a task-specific input module of explicit visual prompting (EVP) during training that uses embedded features and high-frequency component (HFC) information as prompts. Road-SAM also utilizes a carefully designed frequency adapter fine-tuning mechanism, leveraging lightweight yet effective fine-tuning techniques to integrate domain-specific RS knowledge into the RE model, enhancing segmentation performance and making efficient use of computational resources. Comprehensive experiments on two sets of RE benchmark datasets demonstrate the effectiveness of the proposed method. Extensive ablation experiments further validate its superiority over multiple state-of-the-art (SOTA) RS RE algorithms, with updates applied to only 10% of the parameters. Wenqing Feng, Fangli Guan, Chenhao Sun |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | What Do We Actually Need During Self-localization in an Augmented Environment?
Fan Yang 0061, Zhixiang Fang, Fangli Guan |
W2GIS | 3 |