VLDB 2026 Research / reviewers in the wild / expert
Xiaoguang Zhao
dblp:50/5365
· DBLP profile ↗
44ranked-venue papers
1as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Computer networks · 5 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Motion planning with uncertainty in human-populated environments via model-based reinforcement learning for social robot navigation
Xingyuan Gao, Shiying Sun, Chuanbao Zhou, Xiaoguang Zhao, Min Tan 0001 |
Neurocomputing | 4 |
| 2026 | ProPy : Building interactive and efficient prompt pyramids upon CLIP for partially relevant video retrieval
Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao |
Neural Networks | 4 |
| 2026 | Hierarchical Perception Spatio-Temporal Network for Distributed Multi-Robot Collision AvoidanceabstractMulti-robot collision avoidance in complex environments poses a significant challenge, as robots must not only avoid collisions with obstacles but also with other robots. However, most existing methods only use neighbor robot states or raw sensor data as input, which is difficult to deal with scenarios containing both unknown obstacles and neighbor robots. To address this problem, we propose a novel hierarchical perception spatio-temporal network (HPSTN) to capture the spatio-temporal features of the environment from both the sensor level and the agent level. At the agent level, we use the reciprocal velocity obstacle (RVO) to represent the neighbor robot information. At the sensor level, we propose a geometric representation method called lidar obstacle (LO) to describe the lidar-detected occluded sectors, and use it to represent the collision-relevant information of obstacles, so as to achieve a compact and informative observation representation unified with RVO. The robot’s ego state and the above two types of observation information are encoded and fused through spatial and temporal encoders of the HPSTN, enabling the robot to perceive the surroundings more comprehensively and adaptively focus on more important information in complex environments. Extensive simulation experiments with various scenarios demonstrate that our approach outperforms several state-of-the-art methods in terms of safety and efficiency. Furthermore, we also conduct physical experiments with multiple differential-drive robots to validate the effectiveness of our approach in real-world scenarios. Chuanbao Zhou, Shiying Sun, Xiaoguang Zhao, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | RefCap: Zero-shot Video Corpus Moment Retrieval Based on Refined Dense Video CaptioningabstractVideo corpus moment retrieval (VCMR) is a challenging task aimed at localizing specific segments from untrimmed videos within a vast video collection. It has long been addressed using end-to-end supervised or weakly-supervised methods, which often lack explainability and rely on laborious annotations. To address these issues, we exploit generated captions from Vision Large Language Models (VLLMs) and propose the first zero-shot training-free VCMR system, RefCap. The system consists of two decoupled stages: the Construction Stage to propose dense events and construct corresponding captions, and the Retrieval Stage to retrieve events based on text similarities between queries and captions. In the Construction Stage, we design a Sliding-Window Denoiser and a Quality-Measured Event Generator to produce high-quality dense captions, and derive the Indexing Keyword Sets to fully utilize the generated captions. In the Retrieval Stage, we propose a multi-granularity retrieval strategy integrating both sentence-level and word-level textual similarities between queries and event captions. Extensive experiments on the Charades and ActivityNet datasets demonstrate the effectiveness and competitiveness of our method. Code is available at https://github.com/BUAAPY/RefCap. Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao |
ICASSP | 4 |
| 2025 | FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive LearningabstractVideo Corpus Moment Retrieval (VCMR) is a challenging task that aims to localize query-specified moments from a collection of untrimmed videos. The recent state-of-the-art method, JSG, tries to tackle this task using only video-level annotations in a weakly-supervised setting. However, the late fusion strategy of JSG suffers from insufficient alignment, and the proposal-level video representations aggregated with Gaussian Distributions result in semantic inconsistency across frames. To address these issues, we propose a novel weakly-supervised VCMR method, FAWL, which incorporates frame-wise auxiliary alignment and weighted contrastive learning. Two frame-wise auxiliary alignment tasks, namely Query-guided Saliency Alignment (QSA) and Event-aware Boundary Alignment (EBA), are designed first. During training, QSA projects video and text representations through a shared layer, and computes frame-level saliency losses for sufficient multimodal alignment. EBA guides frames to predict distances to proposal boundaries making each frame aware of the corresponding events and thus ensuring semantic consistency. To help the model distinguish proposals with similar visual semantics, we further propose the Weighted Contrastive Learning (WCL) which integrates inter-video similarities into the conventional InfoNCE loss. FAWL enjoys both improved alignment and inference efficiency, achieving new state-of-the-art performance on two challenging datasets. Code is available at https://github.com/BUAAPY/FAWL. Yujia Zhang 0001, Xiaoguang Zhao |
ICASSP | 3 |
| 2025 | Compositional Text-Modality Completion Model for Partially Relevant Video RetrievalabstractPartially Relevant Video Retrieval (PRVR) is a challenging task aimed at retrieving videos based on partially relevant text queries. Previous PRVR models typically adopt the Multiple Instance Learning (MIL) framework, which are limited to annotated segments and are ineffective when dealing with unlabeled segments. To address this limitation, we approach the PRVR task from a novel missing modality completion perspective to provide supplementary textual supervision signals for better alignment. To this end, we propose a plug-and-play Text-Modality Completion (CTC) Model to generate pseudo-descriptions based on existing annotations and the intrinsic attribute of the PRVR task. Our key insight is that the semantics of unlabeled segments can be composed of intra-video and inter-video semantics. Thus, we first design a Semantic Decomposing Module to decompose features into fine-grained semantic units, with a Textual Semantic Sampling strategy to further enrich text representations. Subsequently, a Semantic Composing Module is introduced to compose intra-video and inter-video semantics based on the decomposed semantic units. The decomposed textual semantics, weighted by visual semantic similarities, serve as intra-video semantics, mapping visual representations of the same objects to their corresponding concepts. For inter-video semantics, we design a Generative Semantic Memory that accumulates shared semantics across videos by generating textual semantics given visual semantics. Finally, both intra-video and inter-video semantics are integrated to form pseudo-descriptions, which are leveraged as enriched textual supervision signals for further contrastive alignment. We evaluate the CTC model with different base PRVR models on different public datasets. The consistent improvements demonstrate the effectiveness and generalizability of our method. Yujia Zhang 0001, Xiaoguang Zhao |
ICME | 3 |
| 2025 | Crowd-Aware Robot Navigation for Scenario Generalization via Deep Reinforcement Learning and Knowledge DistillationabstractNavigating a mobile robot through crowded environments presents a significant challenge. Despite extensive research, existing methods often fail to ensure consistent effectiveness across different scenarios. To enhance the ability of the navigation policy to be applied to different scenarios, we propose a crowd navigation policy architecture based on state prediction and value estimation, which includes a state predictor based on pedestrian trajectory prediction and a value estimator using a graph convolutional network, which can navigate the robot to the target position safely and efficiently. Additionally, we propose a knowledge distillation-based training framework to distill the state prediction model, thereby improving its adaptability to human motions in different scenarios. Our approach is evaluated in a simulation environment through multi-scenario experiments using both simulated data and real pedestrian datasets. The experimental results demonstrate the superiority of our approach in terms of success rate and navigation efficiency, indicating that our approach can effectively improve the generalization performance of the navigation policy across different scenarios. Chuanbao Zhou, Shiying Sun, Xingyuan Gao, Xiaoguang Zhao, Min Tan 0001 |
IJCNN | 6 |
| 2025 | AquaScan: A Sonar-based Underwater Sensing System for Human Activity MonitoringabstractHuman activity monitoring in the water is essential for pool management and drowning prevention. Existing camera-based solutions pose significant concerns about privacy and extra installation costs. Although sonars have been widely used for underwater sensing in open aquatic environments such as oceans and lakes, monitoring human activities with sonars in a pool setup is challenging. In this work, we propose AquaScan, the first scanning sonar-based underwater sensing system for human activity monitoring. To overcome the low frame rate, we propose a novel scanning strategy and apply an image reconstruction method to accelerate the scanning speed without compromising the performance of motion detection. We develop a novel signal processing pipeline based on a physical model to remove noises and localize human subjects. We extract features like motion, time, and spatial information from sonar images and develop a state-transfer-based activity recognition system to recognize five common water activities. We deployed AquaScan on three public swimming pools for a total period of 94 hours. The evaluation results show that AquaScan can successfully recognize the five activities in the water at about 91.5%. Haozheng Hou, Sitong Cheng, Xiaoguang Zhao, Peiheng Wu, Lixing He, Yunqi Guo, Guoliang Xing, Zhenyu Yan 0002 |
MobiCom | 4 |
| 2025 | A Hybrid Learning and Optimization Framework for Reactive Whole-Body Motion Planning of Mobile ManipulatorsabstractAs an important branch of embodied artificial intelligence, mobile manipulators are increasingly applied in intelligent services, but their redundant degrees of freedom also limit efficient motion planning in cluttered environments. To address this issue, this paper proposes a hybrid learning and optimization framework for reactive whole-body motion planning of mobile manipulators. We develop the Bayesian distributional soft actor-critic (Bayes-DSAC) algorithm to improve the quality of value estimation and the convergence performance of the learning. Additionally, we use a quadratic programming method to calculate and constrain joint velocities, thereby improving the safety of the whole-body motion planning. We conduct experiments and make comparison with standard benchmark. The experimental results verify that our proposed framework significantly improves the efficiency of reactive whole-body motion planning, reduces the planning time, and improves the success rate of motion planning. Additionally, the proposed reinforcement learning method ensures a rapid learning process in the whole-body planning task. The novel framework allows mobile manipulators to adapt to complex environments more safely and efficiently. Shiying Sun, Chuanbao Zhou, Xiaoguang Zhao, Min Tan 0001 |
SMC | 5 |
| 2025 | Prompt-guided bidirectional deep fusion network for referring image segmentation
Junxian Wu 0001, Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao |
Neurocomputing | 4 |
| 2025 | HierGAT: hierarchical spatial-temporal network with graph and transformer for video HOI detection
Junxian Wu 0001, Yujia Zhang 0001, Michael Kampffmeyer, Shiying Sun, Xiaoguang Zhao |
Multim. Syst. | 8 |
| 2024 | Coarse-to-Fine Recurrently Aligned Transformer with Balance Tokens for Video Moment Retrieval and Highlight DetectionabstractVideo moment retrieval (MR) and highlight detection (HD) are two user-oriented video understanding tasks aimed at extracting query-dependent or highlighted moments to provide valuable content for users. While many recent works have proposed solutions for the joint task of MR and HD leveraging transformer architecture, we argue that existing approaches have not adequately aligned the video and text modalities using basic transformer encoders, and have overlooked the misalignment between irrelevant video clips and text queries. To address these issues, we introduce COREBA: a Coarse-to-Fine Recurrently Aligned Transformer with Balance Tokens. Firstly, we design a plug-and-play Coarse-to-Fine Cross-modal interaction (CFC) module, replacing the original transformer encoder to align the two modalities in a progressive manner. Secondly, we present a novel Recurrent Alignment Mechanism (RAM) to deeply align the modalities in a recurrent fashion. Thirdly, to mitigate the misalignment problem, we append text queries with learnable Balance Tokens to restrict the text information fused with irrelevant clips. Extensive experiments validate the effectiveness and superiority of our proposed method. Yujia Zhang 0001, Shiying Sun, Feihu Zhou, Xiaoguang Zhao |
IJCNN | 6 |
| 2024 | PSAIR: A Neuro-Symbolic Approach to Zero-Shot Visual GroundingabstractSupervised methods for Visual Grounding often require costly annotations of paired sentences and images with ground truth boxes. Recent zero-shot approaches to visual grounding such as ReCLIP and ChatRef aim to avoid the need for costly annotation of paired sentences and images with ground truth boxes. However, these approaches leverage an inflexible detect-then-reasoning paradigm, which leads to a notable semantic information loss. Additionally, these approaches are highly dependent on the definition of predefined keywords or the potentially inconsistent reasoning capabilities of Large Language Models (LLMs). To address these limitations, we propose a neuro-symbolic visual grounding method, PSAIR that incorporates two novel mechanisms, namely Parallel Scoring and Active Information Retrieval. PSAIR equips an LLM with external encapsulated reasoning functions, which the LLM can invoke, ensuring a flexible and stable reasoning process. The proposed Parallel Scoring mechanism is then used to reframe the sequential reasoning process found in prior approaches to facilitate robustness to noise in the detection process. Subsequently, the Active Information Retrieval mechanism is designed to address the loss of semantic information by having the ability to retrieve essential visual information, resulting in a new detect-reason-retrieve paradigm. These innovations result in superior performance and robustness across three public datasets compared to recent state-of-the-art zero-shot visual grounding methods. Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao |
IJCNN | 4 |
| 2024 | Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous DrivingabstractRecently, smart roadside infrastructure (SRI) has demonstrated the potential of achieving fully autonomous driving systems. To explore the potential of infrastructure-assisted autonomous driving, this paper presents the design and deployment of Soar, the first end-to-end SRI system specifically designed to support autonomous driving systems. Soar consists of both software and hardware components carefully designed to overcome various system and physical challenges. Soar can leverage the existing operational infrastructure like street lampposts for a lower barrier of adoption. Soar adopts a new communication architecture that comprises a bi-directional multi-hop I2I network and a downlink I2V broadcast service, which are designed based on off-the-shelf 802.11ac interfaces in an integrated manner. Soar also features a hierarchical DL task management framework to achieve desirable load balancing among nodes and enable them to collaborate efficiently to run multiple data-intensive autonomous driving applications. We deployed a total of 18 Soar nodes on existing lampposts on campus, which have been operational for over two years. Our real-world evaluation shows that Soar can support a diverse set of autonomous driving applications and achieve desirable real-time performance and high communication reliability. Our findings and experiences in this work offer key insights into the development and deployment of next-generation smart roadside infrastructure and autonomous driving systems. Shuyao Shi, Neiwen Ling, Zhehao Jiang, Xuan Huang 0001, Xiaoguang Zhao, Bufang Yang, Chen Bian, Jingfei Xia, Zhenyu Yan 0002, Raymond W. Yeung, Guoliang Xing |
MobiCom | 6 |
| 2024 | Multi-Stage Image-Language Cross-Generative Fusion Network for Video-Based Referring Expression ComprehensionabstractVideo-based referring expression comprehension is a challenging task that requires locating the referred object in each video frame of a given video. While many existing approaches treat this task as an object-tracking problem, their performance is heavily reliant on the quality of the tracking templates. Furthermore, when there is not enough annotation data to assist in template selection, the tracking may fail. Other approaches are based on object detection, but they often use only one adjacent frame of the key frame for feature learning, which limits their ability to establish the relationship between different frames. In addition, improving the fusion of features from multiple frames and referring expressions to effectively locate the referents remains an open problem. To address these issues, we propose a novel approach called the Multi-Stage Image-Language Cross-Generative Fusion Network (MILCGF-Net), which is based on one-stage object detection. Our approach includes a Frame Dense Feature Aggregation module for dense feature learning of adjacent time sequences. Additionally, we propose an Image-Language Cross-Generative Fusion module as the main body of multi-stage learning to generate cross-modal features by calculating the similarity between video and expression, and then refining and fusing the generated features. To further enhance the cross-modal feature generation capability of our model, we introduce a consistency loss that constrains the image-language similarity and language-image similarity matrices during feature generation. We evaluate our proposed approach on three public datasets and demonstrate its effectiveness through comprehensive experimental results. Yujia Zhang 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | CoEdge: A Cooperative Edge System for Distributed Real-Time Deep Learning TasksabstractRecent years have witnessed the emergence of a new class of cooperative edge systems in which a large number of edge nodes can collaborate through local peer-to-peer connectivity. In this paper, we propose CoEdge, a novel cooperative edge system that can support concurrent data/compute-intensive deep learning (DL) models for distributed real-time applications such as city-scale traffic monitoring and autonomous driving. First, CoEdge includes a hierarchical DL task scheduling framework that dispatches DL tasks to edge nodes based on their computational profiles, communication overhead, and real-time requirements. Second, CoEdge can dramatically increase the execution efficiency of DL models by batching sensor data and aggregating the inferences of the same model. Finally, we propose a new edge containerization approach that enables an edge node to execute concurrent DL tasks by partitioning the CPU and GPU workloads into different containers. We extensively evaluate CoEdge on a self-deployed smart lamppost testbed on a university campus. Our results show that CoEdge can achieve up to reduction on deadline missing rate compared to baselines. Zhehao Jiang, Neiwen Ling, Xuan Huang 0001, Shuyao Shi, Chenhao Wu 0006, Xiaoguang Zhao, Zhenyu Yan 0002, Guoliang Xing |
IPSN | 6 |
| 2023 | Visual enhanced hierarchical network for sentence-based video thumbnail generation
Junxian Wu 0001, Yujia Zhang 0001, Xiaoguang Zhao |
Appl. Intell. | 3 |
| 2022 | Hand Acupoint Detection from Images Based on Improved HRNetabstractAs an important component of Traditional Chinese Medicine (TCM), the acupoint therapy has achieved significant success in clinical practice. However, at present, the effect of acupoint therapy heavily depends on the skills of doctors and the acupuncture medical resources are seriously insufficient. The introduction of artificial intelligence technology in acupoint therapy can reduce the workload of doctors and ensure the consistency of acupoint operations, which is of great significance. The key process of acupoint therapy is acupoint detection. In this paper, we apply the deep learning method in automatic acupoint detection using images and propose an improved High-Resolution Network (HRNet) method for hand acupoint detection. What's more, we build a hand acupoint detection dataset and propose an evaluation metric. Experiments on the proposed dataset verify the effectiveness of the proposed method. Shiying Sun, Hongduo Xu, Lingyao Sun, Yuanbo Fu, Yujia Zhang 0001, Xiaoguang Zhao |
IJCNN | 6 |
| 2022 | Multi-Task Learning for Pavement Disease Segmentation Using Wavelet TransformabstractPavement safety is a significant part of transportation safety. In order to realize pavement inspection, many existing approaches have been proposed to detect and segment long-thin cracks while ignoring other pavement diseases. However, other diseases can also cause safety hazards if they are not maintained timely. It is essential to realize the fine-grained segmentation for different pavement diseases to monitor the exact condition of the pavement. Existing algorithms for crack segmentation cannot deal with the segmentation for different diseases with different shapes in the complex environment. To address this, we propose a novel multi-task network for fine-grained pavement disease segmentation. Specifically, the low-frequency information is first introduced to suppress the background noise. The semantic boundary branch is then designed to extract multi-scale feature maps, which can guide the generation of segmentation results for pavement diseases with different shapes. Furthermore, attention mechanisms are utilized to alleviate the effects of extreme imbalances in the number of pixels of different diseases and backgrounds. The network jointly learns the semantic boundary task and the segmentation task in an end-to-end manner to produce the final predictions, and the extensive experiments on a newly collected airport pavement dataset verify the ef-fectiveness of our approach. All the codes are available at https://github.com/wjx1198/MTSSN-WT. Junxian Wu 0001, Yujia Zhang 0001, Xiaoguang Zhao |
IJCNN | 3 |
| 2022 | Generalized zero-shot emotion recognition from body gestures
Jinting Wu, Yujia Zhang 0001, Shiying Sun, Qianzhong Li, Xiaoguang Zhao |
Appl. Intell. | 5 |
| 2022 | Cross-modality synergy network for referring expression comprehension and segmentation
Qianzhong Li, Yujia Zhang 0001, Shiying Sun, Jinting Wu, Xiaoguang Zhao, Min Tan 0001 |
Neurocomputing | 5 |
| 2022 | Beyond Crack: Fine-Grained Pavement Defect Segmentation Using Three-Stream Neural NetworksabstractPavement defect segmentation is a fundamental task in the field of transport infrastructure inspection. Existing methods mainly focus on detection/segmentation for long and thin cracks. However, there are many other types of defects with various sizes and shapes that are also essential to segment, which brings more challenges toward detailed road inspection. To address the above problems and provide a more comprehensive understanding of the overall road conditions, we propose a three-stream neural network that combines spatial, contextual and boundary information for fine-grained defect segmentation. Specifically, the spatial stream captures rich low-level spatial features. The contextual stream utilizes an attention mechanism and models high-level contextual relationships over local features. To further refine the segmentation results, the boundary stream encodes detailed boundaries using a global gated convolution and generates additional boundary maps. By combining the above different information, our model can effectively produce pixel-wise predictions for fine-grained road inspection. The network is trained using a dual-task loss in an end-to-end manner, and experiments were performed on three newly collected datasets, i.e., a fine-grained defect dataset and two crack datasets, which shows that the proposed method achieves favorable segmentation results on complex multi-class defects, and is also able to segment single-class cracks. Specifically, on the fine-grained dataset, it achieved state-of-the-art performance over other competing baselines (mPA of 0.54, mIoU of 0.38, Mic_F of 0.78 and Mac_F of 0.65), where each image is resized to 512$\times 512$and the processing speed is 21 FPS on average. Yujia Zhang 0001, Junxian Wu 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Stress Detection Using Wearable Devices based on Transfer LearningabstractExcessive stress will have a negative impact on people’s physical and mental health, especially for some special occupations. Because stressful stimuli can trigger a variety of physiological responses, analyzing physiological signals collected by wearable devices has become an important way to evaluate the stress state in recent years. However, the number of available subjects of a target group may be small, and collecting a large amount of data when the target group changes is costly and time-consuming. To solve this problem, we propose a stress detection framework for a small target group which uses adversarial transfer learning method to learn shared knowledge about stress between different groups. In order to verify the performance of the framework, we establish a dataset consisting of 264 ordinary college students and 32 police school students, aiming to evaluate the acute stress state of police school students under video stimuli for psychological training in the future. Comprehensive experiments show that our algorithm has achieved a significant improvement in the target group compared with the baseline methods. Jinting Wu, Yujia Zhang 0001, Xiaoguang Zhao |
BIBM | 3 |
| 2021 | TB-Net: A Three-Stream Boundary-Aware Network for Fine-Grained Pavement Disease SegmentationabstractRegular pavement inspection plays a significant role in road maintenance for safety assurance. Existing methods mainly address the tasks of crack detection and segmentation that are only tailored for long-thin crack disease. However, there are many other types of diseases with a wider variety of sizes and patterns that are also essential to segment in practice, bringing more challenges towards fine-grained pavement inspection. In this paper, our goal is not only to automatically segment cracks, but also to segment other complex pavement diseases as well as typical landmarks (markings, runway lights, etc.) and commonly seen water/oil stains in a single model. To this end, we propose a three-stream boundary-aware network (TB-Net). It consists of three streams fusing the low-level spatial and the high-level contextual representations as well as the detailed boundary information. Specifically, the spatial stream captures rich spatial features. The context stream, where an attention mechanism is utilized, models the contextual relationships over local features. The boundary stream learns detailed boundaries using a global-gated convolution to further refine the segmentation outputs. The network is trained using a dual-task loss in an end-to-end manner, and experiments on a newly collected fine-grained pavement disease dataset show the effectiveness of our TB-Net. Yujia Zhang 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001 |
WACV | 3 |
| 2021 | Rethinking semantic-visual alignment in zero-shot object detection via a softplus margin focal loss
Qianzhong Li, Yujia Zhang 0001, Shiying Sun, Xiaoguang Zhao, Min Tan 0001 |
Neurocomputing | 4 |
| 2020 | A Prototype-Based Generalized Zero-Shot Learning Framework for Hand Gesture RecognitionabstractHand gesture recognition plays a significant role in human-computer interaction for understanding various human gestures and their intent. However, most prior works can only recognize gestures of limited labeled classes and fail to adapt to new categories. The task of Generalized Zero-Shot Learning (GZSL) for hand gesture recognition aims to address the above issue by leveraging semantic representations and detecting both seen and unseen class samples. In this paper, we propose an end-to-end prototype-based GZSL framework for hand gesture recognition which consists of two branches. The first branch is a prototype-based detector that learns gesture representations and determines whether an input sample belongs to a seen or unseen category. The second branch is a zero-shot label predictor which takes the features of unseen classes as input and outputs predictions through a learned mapping mechanism between the feature and the semantic space. We further establish a hand gesture dataset that specifically targets this GZSL task, and comprehensive experiments on this dataset demonstrate the effectiveness of our proposed approach on recognizing both seen and unseen gestures. Jinting Wu, Yujia Zhang 0001, Xiaoguang Zhao |
ICPR | 3 |
| 2019 | Learning to Navigate in Human Environments via Deep Reinforcement Learning
Xingyuan Gao, Shiying Sun, Xiaoguang Zhao, Min Tan 0001 |
ICONIP (1) | 3 |
| 2019 | A Novel Development of Robots with Cooperative Strategy for Long-term and Close-proximity Autonomous Transmission-line InspectionabstractWe develop two cooperative robots for power transmission lines (PTLs) inspection - a light climbing robot (CBR) which can stably move on the overhead ground wire (OGW) for sensor data collection and an unmanned aerial vehicle (UAV) with a grabbing mechanism, which can automatically put the CBR on the OGW and take it off. In order to guarantee the safety, the mechanical structures of the connectors are designed in the shape of a trumpet. Further, a self-locked structure of the CBR is developed to automatically seize and release the OGW. For autonomous navigation, the UAV is equipped with a movable sliding rail and a 2D Laser Range Finder (LRF). The LRF can not only detect the position and orientation of the OGW but also detect the top beam of the CBR and the grabbing position in it. Furthermore, the action of the grabbing mechanism is automatically triggered by a microswitch. Finally, by the developed UAV and CBR platforms, we test the whole loading and unloading strategy in an artificially constructed PTLs environment outdoors and achieve an encouraging result1. Combining the flexible motion of the UAV and the high inspection accuracy of the CBR, the CBR can negotiate any obstacle by flying and abandon the traditional heavy obstacle crossing mechanism to effectively realize close-proximity inspection. Due to the light weight and low power consumption, the CBRs can be deployed once in many power corridors to conduct a long-term inspection. Jiang Bian 0004, Xiaolong Hui, Xiaoguang Zhao, Min Tan 0001 |
ICRA | 3 |
| 2018 | Unfamiliar Dynamic Hand Gestures Recognition Based on Zero-Shot Learning
Jinting Wu, Xiaoguang Zhao, Min Tan 0001 |
ICONIP (5) | 3 |
| 2018 | A Novel Monocular-Based Navigation Approach for UAV Autonomous Transmission-Line InspectionabstractThis paper proposes a unique and robust UAV autonomous navigation approach along one side of overhead transmission lines for inspection. To this end, we establish a perspective model and develop a novel Pan/Tilt monocular-based navigation scheme. Simultaneously, the following three key issues are addressed. First, to locate the effective landmark - transmission tower timely and reliably, we customize a neural network for tower detection and combine it with a fast and smooth tracking. Second, to provide UAV with a robust and precise heading, we detect the transmission lines and compute and optimize their vanishing point. Third, to keep a safe distance from transmission lines, we optimize a homography matrix to restore the parallel nature of transmission lines and perceive the distance variation by a point set registration model. Finally, by the designed UAV platform, we test the whole system in a real-world transmission-line inspection scenario under different weather condition and achieve an encouraging result. Our approach provides great flexibility for refined inspection and effectively improves inspection safety. Jiang Bian 0004, Xiaolong Hui, Xiaoguang Zhao, Min Tan 0001 |
IROS | 3 |
| 2017 | Deep Belief Networks for EEG-Based Concealed Information Test
Xiaoguang Zhao, Zeng-Guang Hou, Hongguang Liu 0003 |
ISNN (2) | 2 |
| 2016 | Fast localization for emergency monitoring and rescue in disaster scenarios based on WSNabstractNatural disasters, such as earthquake, typhoon and flood, have caused great loss of lives and property each year, which makes emergency monitoring and rescue an imperative problem to be addressed. In this paper, we designed a novel monitoring and rescue system based on wireless sensor network (WSN) for disaster scenarios, which combines environment monitoring, information transmission, and emergency localization. In our system, fast localization through WSN is a crucial technique for searching and rescuing. We proposed an improved weighted centroid algorithm with adaptive error correction to increase localization accuracy of sensor networks with sparse anchors. This method based on received signal strength indication (RSSI) is low cost, expandable, easy to implement, and suitable for disaster rescues. Mingxiao Lu, Xiaoguang Zhao, Yikun Huang |
ICARCV | 2 |
| 2016 | Online RGB-D tracking via detection-learning-segmentationabstractIn this paper, we address the problem of online RGB-D tracking where the target object undergoes significant appearance changes. To sufficiently exploit the color and depth cues, we propose a novel RGB-D tracking framework (DLS) that simultaneously builds the target 2D appearance model and 3D distribution model. The framework decomposes the tracking task into detection, learning and segmentation. The detection and segmentation components locate the target collaboratively by using the two target models. An adaptive depth histogram is proposed in the segmentation component to efficiently locate the target in depth frames. The learning component estimates the detection and segmentation errors, updates the target models from the most confident frames by identifying two kinds of distractors: potential failure and occlusion. Extensive experimental results on a large-scale benchmark dataset show that the proposed method performs favourably against state-of-the-art RGB-D trackers in terms of efficiency, accuracy, and robustness. Xiaoguang Zhao, Zeng-Guang Hou |
ICPR | 2 |
| 2014 | Survey of single-target visual tracking methods based on online learningabstractVisual tracking is a popular and challenging topic in computer vision and robotics. Owing to changes in the appearance of the target and complicated variations that may occur in various scenes, online learning scheme is necessary for advanced visual tracking framework to adopt. This paper briefly introduces the challenges and applications of visual tracking and focuses on discussing the state‐of‐the‐art online‐learning‐based tracking methods by category. We provide detail descriptions of representative methods in each category, and examine their pros and cons. Moreover, several most representative algorithms are implemented to provide quantitative reference. At last, we outline several trends for future visual tracking research. Xiaoguang Zhao, Zeng-Guang Hou |
IET Comput. Vis. | 2 |
| 2013 | Deployment and Routing Method for Fast Localization Based on RSSI in Hierarchical Wireless Sensor NetworkabstractThis paper proposes a method of hierarchical WSN deployment and routing for fast localization. Position-known anchor nodes form a static upper-layer network organization and several mobile nodes constitute the lower-layer network which need dynamic message exchange with the upper network. In the upper static network, near-optimal routing table is allocated to each anchor node based on their known-position and the strength of signal for communication during network initialization. Mobile nodes in the lower-layer can access to the upper network dynamically and at the same time report the RSSI value of nearby anchors to the control-center. At the center, the computer will calculate the real-time location of the target based on the reported RSSI and node-IDs. Algorithm presented in this paper is low cast, easy to implement and can be used for fast location estimation and motion tracking indoors and outdoors. Xiaoguang Zhao, Min Tan 0001 |
MASS | 2 |
| 2012 | Carrier-based sensor deployment by a mobile robot for wireless sensor networksabstractWe consider a realistic wireless sensor deployment strategy by which mobile robot deploys sensor nodes when it moves along the linear backbone network with some branches. Because of its finite load capacity, the robot has to repeatedly move back to the position where all sensors are temporarily stored and reload sensor nodes, which leads to the robot has to travel the path many times and consumes more energy. We present a Shortest Traveling Path for Robot (STPR) algorithm by which the robot can reduce the traveling path and achieve the required coverage. All nodes are stored at the temporary starting point and mobile robot continuously loads sensors and moves to the destination, dropping some sensor to meet the basic coverage and connectivity requirements. The mobile robot arrives at the far intersection and deploys sensors along the branches in accordance with the algorithm rules, and then returns the recent branch when having enough sensors. Otherwise the robot moves back and deploys sensors on the back path. The robot reloads sensors and repeats the deployment process until all the branches and backbone is finished. The paper proves that the algorithm is effective compared with the common methods. Simulation results show that the algorithm effectively reduces the moving distance at the randomly generated network. Xiaoguang Zhao |
ICARCV | 2 |
| 2012 | Unscented Particle Filter with Systematic Resampling Localization Algorithm Based on RSS for Mobile Wireless Sensor NetworksabstractWe consider the problem of mobile sensor node localization and propose an unscented particle filter algorithm in wireless sensor networks (WSNs) consisting of mobile nodes and static anchor nodes. Because the received signal strength (RSS) varies obviously, we employ particle filter to decrease the bad effect. We first form the system state model, mobile model, and RSS model, and then apply unscented particle filter and utilize the systematic resample to decrease the degeneration of the particle. We eliminate the uncertain of the RSS in wireless channel and get accurate location of mobile nodes. The predicted position of the mobile node is constrained by its velocity and the measurement value of RSS. We do a lot of simulation to validate the algorithm by assigning different parameters. Simulation results show that the algorithm enhances the localization accuracy of mobile node compared with the standard particle filter algorithm. Xiaoguang Zhao |
MSN | 2 |
| 2010 | A motion controller for a pan-tilt camera on an autonomous helicopterabstractIn this paper, a motion controller for a pan-tilt camera mounted on an autonomous helicopter is presented. The motion planner is designed according to the kinematics of the pan-tilt camera. However, the solution of the inverse kinematics of the pan-tilt camera is not unique. To deal with this situation, a decision maker is designed. The decision maker makes use of a cost function to decide which solution should be chose. And the error signal is defined as the difference between pan-tilt joint angles and the desired joint angles. The dynamics of the error system is derived, and proportional controllers are designed according to the error dynamics to control the pan and tilt joints respectively. This system is validated in a virtual reality environment, and the simulation results show that this system can track the target successfully. Haibo Deng, Xiaoguang Zhao, Zeng-Guang Hou |
ICARCV | 2 |
| 2009 | A Vision-Based Ground Target Tracking System for a Small-Scale Autonomous HelicopterabstractIn this paper, we present a visual system to drive an autonomous small-scale helicopter to track a ground target. The vision system is monocular, and the on-board camera is fixed. Since the time-varying attitude of helicopter can bring trouble to vision algorithm design, we developed an attitude compensation algorithm. The real flight experiments show that the vision system can perform in real-time, and track a ground target successfully. Haibo Deng, Xiaoguang Zhao, Z. G. How |
ICIG | 2 |
| 2007 | Application of Neural Network to the Alignment of Strapdown Inertial Navigation System
Meng Bai, Xiaoguang Zhao, Zeng-Guang Hou |
ICIC (1) | 2 |
| 2006 | Motion Deblurring for a Power Transmission Line Inspection Robot
Si-Yao Fu, Yun-Chu Zhang, Xiaoguang Zhao, Zi-ze Liang, Zeng-Guang Hou, An-Min Zou, Min Tan 0001, Wenbo Ye, Lian Bo |
ICIC (2) | 3 |
| 2006 | Visual Navigation for a Power Transmission Line Inspection Robot
Yun-Chu Zhang, Si-Yao Fu, Xiaoguang Zhao, Zi-ze Liang, Min Tan 0001, Yong-Qian Zhang |
ICIC (2) | 3 |
| 2006 | A Remote Aerial Robot for Topographic SurveyabstractIn this paper a seminal system for the topographic survey with an unmanned aerial robot is presented. The proposed system has demonstrated the feasibility of acquiring initial 3D ground models using active laser range sensors on a low-flying helicopter platform. The robot system consists of a remote control helicopter with a laser sensor and a GPS (global positioning system) for collecting ground point data and a post-processing data sub-system for drawing a topographic map. This robot system has many potential applications, such as terrain modeling, structure inspection or climate and weather measurement, etc. The experiment results verify the proposal robot system for topographic survey Xiaoguang Zhao, Min Tan 0001 |
IROS | 1 |
| 2005 | A camera calibration technique based on plane square
Xiaoguang Zhao |
IGARSS | 2 |