Guangyao Shi

dblp:217/3436 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
abstract
Vision Language models (VLMs) have achieved remarkable success in video understanding tasks. Yet, a key question remains: Do they comprehend visual information or merely learn superficial mappings between visual and textual patterns? Understanding visual cues, particularly those related to physics and common sense, is crucial for AI systems interacting with the physical world. However, existing VLM evaluations primarily rely on positive-control tests using real-world videos that resemble training distributions. While VLMs perform well on such benchmarks, it is unclear whether they grasp underlying visual and contextual signals or simply exploit visual-language correlations. To fill this gap, we propose incorporating negative-control tests, i.e., videos depicting physically impossible or logically inconsistent scenarios, and evaluating whether models can recognize these violations. True visual understanding should evince comparable performance across both positive and negative tests. Since such content is rare in the real world, we introduce VideoHallu, a synthetic video dataset featuring physics- and commonsense-violating scenes generated using state-of-the-art tools such as Veo2, Sora, and Kling. The dataset includes expert-annotated question-answer pairs spanning four categories of physical and commonsense violations, designed to be straightforward for human reasoning. We evaluate several leading VLMs, including Qwen-2.5-VL, Video-R1, and VideoChat-R1. Despite their strong performance on real-world benchmarks (e.g., MVBench, MMVU), these models hallucinate or fail to detect physical or logical violations, revealing fundamental weaknesses in visual understanding. Finally, we explore reinforcement learning-based post-training on our negative dataset: fine-tuning improves performance on VideoHallu without degrading results on standard benchmarks, indicating enhanced visual reasoning in VLMs. Our data is available at https://github.com/zli12321/VideoHallu.git.
Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du 0002, Tianyi Zhou 0001, Dinesh Manocha, Jordan L. Boyd-Graber
NeurIPS3
2025 Patch Tensor-Based Geometric Structure Representation for Hyperspectral Imagery Classification
abstract
Hyperspectral remote sensing images (HSIs) typically have nanometer-level spectral resolution, and the rich spatial and spectral information they contain makes it possible to perform detailed land cover analysis. In this letter, a patch tensor-based geometric structure representation (PTGSR) method was proposed for HSIs classification. At first, based on the high-order structure of tensor samples, a novel representation learning model based on tensor neighborhood structure is constructed to capture the relationships between different tensor samples. Then, by utilizing the spatial coordinates of central pixel in the tensor samples, the distribution probabilities between different samples are computed to improve the effectiveness of the representation coefficients. Finally, the classification errors of each class across different tensor orders are integrated to enhance the overall classification performance. Experimental results on three HSI datasets demonstrate that the PTGSR algorithm improves OA by 2.78%, 3.15%, 1.84% over TLRSR-IDL for Indian Pines, Fanglu, and Houston2013, respectively.
Guangyao Shi, Wenhao Xiang
IEEE Geosci. Remote. Sens. Lett.1
2025 MMD-MLP: LiDAR-Guided Hyperspectral Data Classification Using Local-Global Directional-MLP With Multiresolution Multiscale Representation
abstract
Hyperspectral images (HSIs) are rich in spectral information and are widely used in the field of land-cover classification. However, existing deep learning methods ignore the combination of multiscale and multiresolution information, while not making better use of light detection and ranging (LiDAR) elevation information to assist in the enhancement of HSI data. To solve the problem above, we proposed a multiresolution, multiscale local-global directional MLP (MMD-MLP) for HSI classification with a LiDAR-guided feature enhancement module in this article. The algorithm first introduced a local-global-directed MLP structure, which effectively combines local and global features. Second, a multiresolution and multiscale feature extraction strategy is invented for the accurate acquisition of the detailed information and features of different land covers under different scale sizes. Subsequently, a LiDAR-guided feature enhancement module, which introduces the elevation information from LiDAR for improving the feature representation of HSI, adopts a cross-attention mechanism to reduce the semantic gap to improve the features of HSI. The proposed algorithm was evaluated on multiple hyperspectral-LiDAR datasets, and the results demonstrate that it achieves state-of-the-art (SOTA) performance. The code will be available athttps://github.com/sanxian-svg/MMD-MLP.
Fulin Luo, Yiyan Hua, Chuan Fu, Tan Guo, Guangyao Shi
IEEE Trans. Geosci. Remote. Sens.6
2024 AG-Cvg: Coverage Planning with a Mobile Recharging UGV and an Energy-Constrained UAV
abstract
In this paper, we present an approach for coverage path planning for a team of an energy-constrained Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV). Both the UAV and the UGV have predefined areas that they have to cover. The goal is to perform complete coverage by both robots while minimizing the coverage time. The UGV can also serve as a mobile recharging station. The UAV and UGV need to occasionally rendezvous for recharging. We propose a heuristic method to address this NP-Hard planning problem. Our approach involves initially determining coverage paths without factoring in energy constraints. Subsequently, we cluster segments of these paths and employ graph matching to assign UAV clusters to UGV clusters for efficient recharging management. We perform numerical analysis on real-world coverage applications and show that compared with a greedy approach our method reduces rendezvous overhead on average by 11.33%. We demonstrate proof-of-concept with a team of a VOXL m500 drone and a Clearpath Jackal ground vehicle, providing a complete system from the offline algorithm to the field execution.
Nare Karapetyan, Ahmad Bilal Asghar, Amisha Bhaskar, Guangyao Shi, Dinesh Manocha, Pratap Tokekar
ICRA4
2024 LAVA: Long-horizon Visual Action based Food Acquisition
abstract
Robotic Assisted Feeding (RAF) addresses the fundamental need for individuals with mobility impairments to regain autonomy in feeding themselves. The goal of RAF is to use a robot arm to acquire and transfer food to individuals from the table. Existing RAF methods primarily focus on solid foods, leaving a gap in manipulation strategies for semisolid and deformable foods. We present Long-horizon Visual Action-based (LAVA) food acquisition of liquid, semisolid, and deformable foods. Long-horizon refers to the goal of "clearing the bowl" by sequentially acquiring the food from the bowl. LAVA is hierarchical: (1) At the highest level, we determine primitives using ScoopNet. (2) At the mid-level, LAVA finds parameters for the low-level primitives. (3) At the lowest level, LAVA carries out action execution using behavior cloning. We validate LAVA on real-world acquisition trials involving granular, liquid, semisolid, and deformable foods along with fruit chunks and soup. Across 46 bowls, LAVA acquires much more efficiently than baselines with a success rate of 89±4%, and generalizes across realistic plate variations such as varying positions, varieties, and amount of food in the bowl. Datasets and supplementary materials can be found on our website.
Amisha Bhaskar, Rui Liu 0040, Vishnu Dutt Sharma, Guangyao Shi, Pratap Tokekar
IROS4
2024 Inverse Submodular Maximization with Application to Human-in-the-Loop Multi-Robot Multi-Objective Coverage Control
abstract
We consider a new type of inverse combinatorial optimization, Inverse Submodular Maximization (ISM), for human-in-the-loop multi-robot coordination. Forward combinatorial optimization - solving a combinatorial problem given the reward (cost)-related parameters - is widely used in multi-robot coordination. In the standard pipeline, the reward (cost)-related parameters are designed offline by domain experts. These parameters are utilized for coordinating robots online. What if non-expert human supervisors desire to change these parameters during task execution to adapt to some new requirements? We are interested in the case where human supervisors can suggest what actions to take, and the robots need to change these internal parameters accordingly. We study such problems from the perspective of inverse combinatorial optimization, i.e., the process of finding parameters given solutions to the problem. Specifically, we propose a new formulation for ISM, in which we aim to find a new set of parameters that minimally deviate from the current parameters while causing a greedy algorithm to output actions which are the same as those desired by the human supervisors. We show that such problems can be formulated as a Mixed Integer Quadratic Program (MIQP) which is intractable for existing solvers when the problem size is large. We propose a new Branch & Bound algorithm to solve such problems. In numerical simulations, we demonstrate how to use ISM in multi-robot multi-objective coverage control, and we show that the proposed algorithm provides significant advantages in running time and peak memory usage compared to directly using an existing solver.
Guangyao Shi, Gaurav S. Sukhatme
IROS1
2024 TCDM: Effective Large-Factor Image Super-Resolution via Texture Consistency Diffusion
abstract
Recently, remote sensing super-resolution (SR) tasks have been widely studied and achieved remarkable performance. However, due to the complex texture and serious image degeneration, the conventional methods (e.g. CNN-based, GAN-based) cannot reconstruct high-resolution (HR) remote sensing images with a large SR factor (≥ ×8). In this paper, we model the large-factor super-resolution (LFSR) task as a referenced diffusion process and explore how to embed pixel-wise constraint into the popular diffusion model. Following this motivation, we propose the first diffusion-based LFSR method named texture consistency diffusion model (TCDM) for remote sensing images. Specifically, we build a novel conditional truncated noise generator (CTNG) in TCDM to simultaneously generate the expectation of posterior probabilityp(xt-1|xt) and the truncated noise image. With the predicted truncated noise image, sampling an SR image using CTNG saves nearly 90% processing time compared to the naive diffusion model. Additionally, we design a new denoising process named texture consistency diffusion (TC-diffusion) to explicitly embed pixel-wise constraints into the LFSR diffusion model during the training stage. Universal experiments on five commonly used remote sensing datasets demonstrate that the proposed TCDM surpasses the latest SR methods by a large margin and reports new SOTA results on several evaluation metrics. Additionally, the proposed method demonstrates impressive visual quality on reconstructed remote sensing image texture and details.
Yan Zhang 0108, Hanqi Liu, Xinbo Gao 0001, Guangyao Shi, Jianan Jiang
IEEE Trans. Geosci. Remote. Sens.5
2023 Risk-aware Recharging Rendezvous for a Collaborative Team of UAVs and UGVs
abstract
We introduce and investigate the recharging rendezvous problem for a collaborative team of Unmanned Aerial Vehicles (UAVs) and Unmanned Ground Vehicles (UGVs), in which UAVs with limited battery capacity and UGVS persistently monitor an area. The UGVs also act as mobile recharging stations for the UAVs. In contrast to prior work on such problems, we consider the challenge of dealing with stochastic energy consumption in a risk-aware fashion. Specifically, we consider a bi-criteria optimization problem of minimizing the time taken by the UAVs on recharging detours while ensuring that the probability that no UAV runs out of charge is greater than a user-defined risk tolerance. This problem (termed Risk-aware Recharging Rendezvous Problem (RRRP)) is a combinatorial problem with a matching constraint — to ensure UAVs are assigned to the limited UGV recharging slots, and a knapsack constraint — to capture the risk tolerance. We propose a novel bicriteria approximation algorithm to solve RRRP and demonstrate its effectiveness in the context of a persistent monitoring mission compared to baseline methods.
Ahmad Bilal Asghar, Guangyao Shi, Nare Karapetyan, James Humann, Jean-Paul Reddinger, James Dotterweich, Pratap Tokekar
ICRA2
2023 Data-Driven Distributionally Robust Optimal Control with State-Dependent Noise
abstract
Distributionally Robust Optimal Control (DROC) is a technique that enables robust control in a stochastic setting when the true distribution is not known. Traditional DROC approaches require given ambiguity sets or a KL divergence bound to represent the distributional uncertainty. These may not be known a priori and may require hand-crafting. In this paper, we lift this assumption by introducing a data-driven technique for estimating the uncertainty and a bound for the KL divergence. We call this technique D3ROC. To evaluate the effectiveness of our approach, we consider a navigation problem for a car-like robot with unknown noise distributions. The results demonstrate that D3ROC provides robust and efficient control policies that outperform the iterative Linear Quadratic Gaussian (iLQG) control. The results also show the effectiveness of our proposed approach in handling different noise distributions.
Rui Liu 0040, Guangyao Shi, Pratap Tokekar
IROS2
2023 Decision-Oriented Learning with Differentiable Submodular Maximization for Vehicle Routing Problem
abstract
We study the problem of learning a function that maps context observations (input) to parameters of a submodular function (output). Our motivating case study is a specific type of vehicle routing problem, in which a team of Unmanned Ground Vehicles (UGVs) can serve as mobile charging stations to recharge a team of Unmanned Ground Vehicles (UAVs) that execute persistent monitoring tasks. We want to learn the mapping from observations of UAV task routes and wind field to the parameters of a submodular objective function, which describes the distribution of landing positions of the UAVs. Traditionally, such a learning problem is solved independently as a prediction phase without considering the downstream task optimization phase. However, the loss function used in prediction may be misaligned with our final goal, i.e., a good routing decision. Good performance in the isolated prediction phase does not necessarily lead to good decisions in the downstream routing task. In this paper, we propose a framework that incorporates task optimization as a differentiable layer in the prediction phase. Our framework allows end-to-end training of the prediction model without using engineered intermediate loss that is targeted only at the prediction performance. In the proposed framework, task optimization (submodular maximization) is made differentiable by introducing stochastic perturbations into deterministic algorithms (i.e., stochastic smoothing). We demonstrate the efficacy of the proposed framework using synthetic data. Experimental results of the mobile charging station routing problem show that the proposed framework can result in better routing decisions, e.g. the average number of UAVs recharged increases, compared to the prediction-optimization separate approach.
Guangyao Shi, Pratap Tokekar
IROS1
2023 Pansharpening Method Based on Deep Nonlocal Unfolding
abstract
Although deep neural networks (DNNs) have achieved great success in pansharpening, most of them lack transparency and interpretability. Currently, some DNNs methods utilize deep unfolding techniques to alleviate this problem. However, they do not consider the regularization term separately when solving the energy function that represents the image degradation process, making it difficult to extract complex prior information in the unfolding module. Therefore, this paper proposes a pansharpening method based on deep non-local unfolding. Specifically, we expand the iterative process of solving the energy function into the corresponding neural network modules, making each module have a certain physical meaning. Then, we decouple the prior operator containing the prior knowledge of the remote sensing image and approximate the solution using the network module. Meanwhile, we incorporate local and non-local self-similarity priors into the prior operator and design a two-branch prior module for learning the prior features and contribution weights adaptively. Finally, the fused image is corrected with the learned prior features to approximate the real image. Experimental results on datasets from two different types of satellites demonstrate the superiority of our approach.
Guangyao Shi, Liping Zhang 0012, Weisheng Li 0001, Dajiang Lei
IEEE Trans. Geosci. Remote. Sens.3
2023 Robust Multiple-Path Orienteering Problem: Securing Against Adversarial Attacks
abstract
The multiple-path orienteering problem asks for paths for a team of robots that maximize the total reward collected while satisfying budget constraints on the path length. This problem models many multirobot routing tasks, such as exploring unknown environments and information gathering for environmental monitoring. In this article, we focus on how to make the robot team robust to failures when operating in adversarial environments. We introduce the robust multiple-path orienteering problem (RMOP), where we seek worst case guarantees against an adversary that is capable of attacking at most$\alpha$robots. We consider two versions of this problem: RMOP offline and RMOP online. In the offline version, there is no communication or replanning when robots execute their plans, and our main contribution is a general approximation scheme with a bounded approximation guarantee that depends on$\alpha$and the approximation factor for single-robot orienteering. In particular, we show that the algorithm yields a: 1) constant-factor approximation when the cost function is modular; 2)$\log$factor approximation when the cost function is submodular; and 3) constant-factor approximation when the cost function is submodular, but the robots are allowed to exceed their path budgets by a bounded amount. In the online version, the RMOP is modeled as a two-player sequential game and solved adaptively in a receding horizon fashion based on Monte Carlo tree search. In addition to theoretical analysis, we perform simulation studies for ocean monitoring and tunnel information-gathering applications to demonstrate the efficacy of our approach.
Guangyao Shi, Lifeng Zhou 0001, Pratap Tokekar
IEEE Trans. Robotics1
2022 Interactive Multi-Robot Aerial Cinematography Through Hemispherical Manifold Coverage
abstract
This paper presents a distributed interactive framework to provide high-level position instructions for multi-robot aerial cinematography based on coverage over a hemisphere. The control strategy based on optimization of the coverage functional and geometric relationships over a hemisphere is presented. It enables multiple Unmanned Aerial Vehicles (UAVs) to coordinate their motion while tracking a dynamic (real or virtual) target, and can accommodate high-level human inputs to influence UAV concentration. In this framework, each UAV uses local information combined with exogenous inputs to determine its motion. The two inputs to the system, i.e., the predicted trajectory of the target and user-defined aesthetic preferences, are agnostic to the size of the multi-robot system (MRS). The proposed framework is validated using the PX4 SITL Autopilot simulator in Gazebo, and the scalability of the framework is verified via simulations.
Guangyao Shi, Pratap Tokekar, Yancy Diaz-Mercado
IROS2
2021 Communication-Aware Multi-robot Coordination with Submodular Maximization
abstract
Submodular maximization has been widely used in many multi-robot task planning problems including information gathering, exploration, and target tracking. However, the interplay between submodular maximization and communication is rarely explored in the multi-robot setting. In many cases, maximizing the submodular objective may drive the robots in a way so as to disconnect the communication network. Driven by such observations, in this paper, we consider the problem of maximizing submodular function with connectivity constraints. Specifically, we propose a problem called Communication-aware Submodular Maximization (CSM), in which communication maintenance and submodular maximization are jointly considered in the decision-making process. One heuristic algorithm that consists of two stages, i.e. topology generation and deviation minimization is proposed. We validate the formulation and algorithm through numerical simulation. We find that our algorithm on average suffers only slightly performance decrease compared to the pure greedy strategy.
Guangyao Shi, Ishat E. Rabban, Lifeng Zhou 0001, Pratap Tokekar
ICRA1
2021 Local Structure Graph Discriminant Embedding for Hyperspectral Image Classification
abstract
Graph learning is an effective technique to reduce the dimensionality of hyperspectral image (HSI) and improve classification result. However, the previous graph methods don't consider the local structure of each pixel in HSI. HSI has a complex non-linear structure, and the local structure can be regarded as a linear distribution. Therefore, we propose a local structure graph discriminant embedding (LSGDE) method to better reveal the intrinsic properties of HSI. This method constructs an intraclass and an interclass structure graphs to compact the intraclass samples and separate the interclass samples. Meanwhile, an interclass weighted scatter based on probability distribution is designed to enhance the discriminative ability of different classes. Then, a projection matrix can be obtained to map high-dimensional data into a low-dimensional space. Experiments on the HoustonU data set show that LSGDE can achieve better performance than the related DR methods.
Zehua Zou, Fulin Luo, Guangyao Shi
IGARSS4
2020 Multi-manifold locality graph preserving analysis for hyperspectral image classification
Guangyao Shi, Hong Huang 0002, Zhengying Li
Neurocomputing1
2020 Two-stream feature aggregation deep neural network for scene classification of remote sensing images
Kejie Xu, Hong Huang 0002, Peifang Deng, Guangyao Shi
Inf. Sci.4
2020 Unsupervised Dimensionality Reduction for Hyperspectral Imagery via Local Geometric Structure Feature Learning
abstract
Hyperspectral images (HSIs) possess a large number of spectral bands, which easily lead to the curse of dimensionality. To improve the classification performance, a huge challenge is how to reduce the number of spectral bands and preserve the valuable intrinsic information in the HSI. In this letter, we propose a novel unsupervised dimensionality reduction method called local neighborhood structure preserving embedding (LNSPE) for HSI classification. At first, LNSPE reconstructs each sample with its spectral neighbors and obtains the optimal weights for constructing the adjacency graph by modifying its loss function. Then, to discover the scatter information of the training samples, LNSPE minimizes the scatter between the pixels and the corresponding neighbors and maximizes the total scatter of the HSI data. Finally, it incorporates the scatter information and the dual graph structure to enhance the aggregation of the HSI. As a result, LNSPE can effectively reveal the intrinsic structure and improve the classification performance of the HSI data. The experimental results on two real hyperspectral data sets exhibit the efficiency and superiority of LNSPE to some state-of-the-art methods.
Guangyao Shi, Hong Huang 0002
IEEE Geosci. Remote. Sens. Lett.1
2020 Multilayer Feature Fusion Network for Scene Classification in Remote Sensing
abstract
The scene classification of high spatial resolution (HSR) images is a challenging task in the remote sensing community. How to construct a discriminative representation of the HSR scene is a key step to improve classification performance. In this letter, we propose a novel feature extraction method termed multilayer feature fusion network (MF2Net) for scene classification. At first, the transferred VGGNet-16 model is employed as a feature extractor to acquire multilayer convolutional features. Then, several layers including pooling, transformation, and fusion layers are designed to process hierarchical features in four branches, and the prediction probability can be obtained for classification. Finally, the proposed model is optimized by fine-tuning techniques, where a novel data augmentation approach is explored to improve generalization ability. As a result, MF2Net effectively applies useful information from multilayers to improve the accuracy of scene classification. The experimental results on AID and NWPU-RESISC45 data sets exhibit that the MF2Net method obtains quite competitive classification results compared with many state-of-the-art methods.
Kejie Xu, Hong Huang 0002, Yuan Li 0061, Guangyao Shi
IEEE Geosci. Remote. Sens. Lett.4
2020 DLPNet: A deep manifold network for feature extraction of hyperspectral imagery
Zhengying Li, Hong Huang 0002, Guangyao Shi
Neural Networks4
2020 Dimensionality Reduction of Hyperspectral Imagery Based on Spatial-Spectral Manifold Learning
abstract
The graph embedding (GE) methods have been widely applied for dimensionality reduction of hyperspectral imagery (HSI). However, a major challenge of GE is how to choose the proper neighbors for graph construction and explore the spatial information of HSI data. In this paper, we proposed an unsupervised dimensionality reduction algorithm called spatial-spectral manifold reconstruction preserving embedding (SSMRPE) for HSI classification. At first, a weighted mean filter (WMF) is employed to preprocess the image, which aims to reduce the influence of background noise. According to the spatial consistency property of HSI, SSMRPE utilizes a new spatial-spectral combined distance (SSCD) to fuse the spatial structure and spectral information for selecting effective spatial-spectral neighbors of HSI pixels. Then, it explores the spatial relationship between each point and its neighbors to adjust the reconstruction weights to improve the efficiency of manifold reconstruction. As a result, the proposed method can extract the discriminant features and subsequently improve the classification performance of HSI. The experimental results on the PaviaU and Salinas hyperspectral data sets indicate that SSMRPE can achieve better classification results in comparison with some state-of-the-art methods.
Hong Huang 0002, Guangyao Shi, Haibo He, Fulin Luo
IEEE Trans. Cybern.2
2020 Local Linear Spatial-Spectral Probabilistic Distribution for Hyperspectral Image Classification
abstract
A key challenge in hyperspectral image (HSI) classification is how to effectively utilize the spectral and spatial information of limited labeled training samples in the data set. In this article, a new spatial-spectral combined classification method, termed local linear spatial-spectral probabilistic distribution (LSPD), has been proposed on the basis of local geometric structure and spatial consistency of HSI. LSPD extracts discriminating spatial-spectral information from limited labeled training samples and their spatial-spectral neighbors. Then, it constructs a multiclass probability map by exploiting the local linear representation and spatial information of HSI. Finally, the spatial-spectral weighted reconstruction has been performed on the probability map, and the class of test sample can be predicted by the maximum value of LSPD. LSPD not only exploits spectral information to discover more intrinsic properties of the labeled training data but also utilizes the spatial relationship between samples to effectively improve discriminating power for classification. Experimental results on the Indian Pines, PaviaU, and HoustonU hyperspectral data sets demonstrate that the proposed LSPD method possesses better classification performance by comparing with some state-of-the-art classifiers.
Hong Huang 0002, Haibo He, Guangyao Shi
IEEE Trans. Geosci. Remote. Sens.4