VLDB 2026 Research / reviewers in the wild / expert
Liang Yao 0001
dblp:71/8496-1
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-4588-3658ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AirNavigation: Let UAV Navigation Tell Its Own StoryabstractTesting autonomous navigation algorithms of Unmanned Aerial Vehicles (UAVs) in real-world scenarios often entails significant safety risks. In this paper, we aim to build a flexible yet user-friendly UAV autonomous navigation simulator. Ideally, it should closely emulate real-world environments, support diverse UAV models and algorithms, and provide a flexible evaluation framework. Existing frameworks fail to satisfy all three requirements simultaneously. To this end, we present AirNavigation, an integrated simulation platform designed to support the end-to-end workflow of UAV navigation research. Specifically, our system leverages Unreal Engine to simulate highly realistic environments and diverse UAV models. It further facilitates semi-automated scene generation and multi-modal synthetic training data production. To lower the barrier of adoption, we develop a suite of user-friendly interfaces to enable seamless integration of diverse navigation algorithms. Moreover, we introduce a novel evaluation system powered by large language models to deliver personalized and fine-grained performance analysis. Jianyu Jiang, Zequan Wang, Liang Yao 0001, Shengxiang Xu, Fan Liu 0003 |
AAAI | 3 |
| 2026 | RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowabstractRemote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to construct an Earth observation workflow to handle complex queries by reasoning about spatial context and user intent. As a reasoning workflow, it should autonomously explore and construct its own inference paths, rather than being confined to predefined ground‑truth sequences. Ideally, its architecture ought to be unified yet generalized, possessing capabilities to perform diverse reasoning tasks through one model without requiring additional fine-tuning. Existing remote sensing approaches rely on supervised fine-tuning paradigms and task‑specific heads, limiting both autonomous reasoning and unified generalization. To this end, we propose RemoteReasoner, a unified workflow for geospatial reasoning. The design of RemoteReasoner integrates a multi-modal large language model (MLLM) for interpreting user instructions and localizing targets, together with task transformation strategies that enable multi-granularity tasks, including object-, region-, and pixel-level. In contrast to existing methods, our framework is trained with reinforcement learning (RL) to endow the MLLM sufficient reasoning autonomy. At the inference stage, our transformation strategies enable diverse task output formats without requiring task-specific decoders or further fine-tuning. Experiments demonstrated that RemoteReasoner achieves state-of-the-art performance across multi-granularity reasoning tasks. Furthermore, it retains the MLLM's inherent generalization capability, demonstrating robust performance on unseen tasks and categories. Liang Yao 0001, Fan Liu 0003, Hongbo Lu, Chuanyi Zhang, Shengxiang Xu, Shimin Di |
AAAI | 1 |
| 2026 | Heterogeneous Knowledge Distillation Fostered Pre-training for remote sensing object detection
Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001 |
Pattern Recognit. | 4 |
| 2025 | Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image CaptionsabstractWhile densely annotated image captions significantly facilitate the learning of robust visionlanguage alignment, methodologies for systematically optimizing human annotation efforts remain underexplored.We introduce CHAIN-OF-TALKERS (COTALK), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time).The framework is built upon two key insights.First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the "residual"-the missing visual information that previous annotations have not covered.Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency.We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment.Experiments with eight participants show our CHAIN-OF-TALKERS (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%)over the parallel method.per minute? a review and meta-analysis of reading rate. Delong Chen, Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001, Yuhui Zheng |
EMNLP | 6 |
| 2025 | RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image ClassificationabstractSince high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of remote sensing images, resulting in significant accuracy loss after pruning. To this end, we propose an effective structural pruning approach for remote sensing image classification. Specifically, a pruning strategy that amplifies the differences in channel importance of the model is introduced. Then an adaptive mining loss function is designed for the fine-tuning process of the pruned model. Finally, we conducted experiments on two remote sensing classification datasets. The experimental results demonstrate that our method achieves minimal accuracy loss after compressing remote sensing classification models, achieving state-of-the-art (SoTA) performance. Guangwenjie Zou, Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Xin Li 0090, Shengxiang Xu, Jun Zhou 0001 |
ICASSP | 2 |
| 2025 | RemoteSAM: Towards Segment Anything for Earth ObservationabstractWe aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM. Liang Yao 0001, Fan Liu 0003, Delong Chen, Chuanyi Zhang, Ziyun Chen 0004, Shimin Di, Yuhui Zheng |
ACM Multimedia | 1 |
| 2025 | UEMM-Air: Enable UAVs to Undertake More Multi-modal TasksabstractThe development of multi-modal Unmanned Aerial Vehicles (UAVs) environment perception systems is hindered by three critical gaps in existing datasets: (1) insufficient modalities and pixel misalignment, (2) noisy labels, and (3) limited task types. To address these gaps, we propose an automatic data construction approach and construct a multi-modal UAV-based environment perception dataset, UEMM-Air. Its synthetic nature ensures scalability, reproducibility, and rare-event coverage, making it suitable for large-scale model pre-training. Benefiting from our automated data collection and annotation pipeline, UEMM-Air encompasses 120k data pairs across 6 aligned modalities and supports 4 perception tasks, significantly exceeding existing datasets (max 60k data, 3 modalities, 2 tasks). Compared to existing synthetic datasets like SynDrone, UEMM-Air provides more accurate annotations by avoiding noisy labels from direct coordinate computation. Notably, models pre-trained on UEMM-Air achieve a 5.8% accuracy improvement compared to those utilizing other synthetic datasets, while requiring less than half the data. This benchmark establishes performance evaluation of UAV multi-modal environmental perception models, and hopefully encourages more research efforts towards enabling UAVs to undertake more multi-modal tasks. The dataset and its generation engine are openly accessible under a permissive license at https://github.com/1e12Leon/UEMM-Air. Liang Yao 0001, Fan Liu 0003, Shengxiang Xu, Chuanyi Zhang, Shimin Di, Jianyu Jiang, Zequan Wang, Jun Zhou 0001 |
ACM Multimedia | 1 |
| 2025 | Unifying Foundation Model and Segment Anything Model for Remote Sensing Weakly Supervised Semantic SegmentationabstractDue to its reliance on fewer precise annotations, weakly supervised semantic segmentation (WSSS) techniques are in high demand in the field of remote sensing (RS) image processing. Despite mainstream WSSS approaches achieve remarkable dense prediction accuracies, they still face challenges such as insufficient pre-trained and ambiguous segment predictions. To this end, we propose to improve the accuracy of weakly supervised semantic segmentation by unifying the vision language foundation model and the Segment Anything Model (SAM). Specifically, we leverage a remote sensing vision language foundational model, RemoteCLIP, to provide sufficient pre-trained knowledge. Subsequently, we employ a decoder to transform the high-level feature representations extracted by RemoteCLIP into the segmentation predictions. Then, we introduce a multi-prompt fusion (MPF) approach via the Segment Anything Model (SAM) to obtain high-quality segment results with well-defined boundaries. To the best of our knowledge, this is the first study to apply a unified framework of foundation model and Segment Anything Model for RS WSSS. Experimental results demonstrate that our method achieves remarkable performance across three remote sensing datasets. Jinfeng Cui, Liang Yao 0001, Guoyan Xu, Fan Liu 0003 |
SMC | 4 |
| 2025 | Domain-Invariant Progressive Knowledge Distillation for UAV-Based Object DetectionabstractKnowledge distillation (KD) is an effective method for compressing models in object detection tasks. Due to limited computational capability, unmanned aerial vehicle-based object detection (UAV-OD) widely adopt the KD technique to obtain lightweight detectors. Existing methods often overlook the significant differences in feature space caused by the large gap in scale between the teacher and student models. This limitation hampers the efficiency of knowledge transfer during the distillation process. Furthermore, the complex backgrounds in aerial images make it challenging for the student model to efficiently learn the object features. In this letter, we propose a novel KD framework for UAV-OD. Specifically, a progressive distillation approach is designed to alleviate the feature gap between teacher and student models. Then, a new feature alignment method is provided to extract object-related features for enhancing the student model’s knowledge reception efficiency. Finally, extensive experiments are conducted to validate the effectiveness of our proposed approach. The results demonstrate that our proposed method achieves state-of-the-art performance on two datasets. Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Zhiquan Ou |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Boost UAV-Based Object Detection via Scale-Invariant Feature Disentanglement and Adversarial LearningabstractDetecting objects from Unmanned Aerial Vehicles (UAV) is often hindered by a large number of small objects, resulting in low detection accuracy. To address this issue, mainstream approaches typically utilize multi-stage inferences. Despite their remarkable detecting accuracies, real-time efficiency is sacrificed, making them less practical to handle real applications. To this end, we propose to improve the single-stage inference accuracy through learning scale-invariant features. Specifically, a Scale-Invariant Feature Disentangling module is designed to disentangle scale-related and scale-invariant features. Then an Adversarial Feature Learning scheme is employed to enhance disentanglement. Finally, scale-invariant features are leveraged for robust UAV-based object detection. Furthermore, we construct a multi-modal UAV object detection dataset, State-Air, which incorporates annotated UAV state parameters. We apply our approach to three lightweight detection frameworks on two benchmark datasets. Extensive experiments demonstrate that our approach can effectively improve model accuracy and achieve state-of-the-art (SoTA) performance on three datasets. Our code and dataset are publicly available at https://github.com/1e12Leon/SIFDAL. Fan Liu 0003, Liang Yao 0001, Chuanyi Zhang, Xiruo Jiang, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | AerialFace: A Light Weight Framework for Unmanned Aerial Vehicle Face RecognitionabstractUnmanned Aerial Vehicle (UAV) are widely applied in multiple fields due to their simple structure and high flexibility. Applying facial recognition technology to UAV can improve their intelligence and diversity of application scenarios. However, UAV face recognition is often hindered by low resolution face, resulting in low accuracy. To alleviate these issues, we propose an efficient face recognition framework named AerialFace. Firstly, we utilize the Residual SRGAN (ResSR-GAN) model to enhance image quality and generate high-resolution face images. Then, we propose Semantic-improved MobileFaceNet (SeMFNet) to relieve the impact of complex backgrounds. Finally, we leverage two pruning algorithms for face detection and recognition models, respectively. It can reduce their parameters to meet the deployment requirements of the algorithm on UAVs. Furthermore, we apply our AerialFace on a UAV face dataset and employ it on an edge computing device. Extensive experiments demonstrate that our approach can effectively improve UAV face recognition accuracy and have real-time performance in embedded UAV devices. Zhiquan Ou, Liang Yao 0001, Fan Liu 0003 |
FG | 2 |