Hao Zhang 0011

dblp:55/2270-11 · DBLP profile ↗
← Back
50ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0001-8043-9184ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 9 first-author · 12 since 2021Systems, architecture and hardware · 27 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Coordinated Multi-Robot Navigation with Formation Adaptation
abstract
Coordinated multi-robot navigation is an essential ability for a team of robots operating in diverse environments. Robot teams often need to maintain specific formations, such as wedge formations, to enhance visibility, positioning, and efficiency during fast movement. However, complex environments such as narrow corridors challenge rigid team formations, which makes effective formation control difficult in real-world environments. To address this challenge, we introduce a novel Adaptive Formation with Oscillation Reduction (AFOR) approach to improve coordinated multi-robot navigation. We develop AFOR under the theoretical framework of hierarchical learning and integrate a spring-damper model with hierarchical learning to enable both team coordination and individual robot control. At the upper level, a graph neural network facilitates formation adaptation and information sharing among the robots. At the lower level, reinforcement learning enables each robot to navigate and avoid obstacles while maintaining the formations. We conducted extensive experiments using Gazebo in the Robot Operating System (ROS), a high-fidelity Unity3D simulator with ROS, and real robot teams. Results demonstrate that AFOR enables smooth navigation with formation adaptation in complex scenarios and outperforms previous methods. More details of this work are provided on the project website: https://hcrlab.gitlab.io/project/afor.
Peng Gao 0009, Williard Joshua Jose, Christopher M. Reardon, Maggie B. Wigness, John G. Rogers III, Hao Zhang 0011
ICRA7
2025 Bandwidth-Adaptive Spatiotemporal Correspondence Identification for Collaborative Perception
abstract
Correspondence identification (CoID) is an essential capability in multi-robot collaborative perception, which enables a group of robots to consistently refer to the same objects within their respective fields of view. In real-world applications, such as connected autonomous driving, vehicles face challenges in directly sharing raw observations due to limited communication bandwidth. In order to address this challenge, we propose a novel approach for bandwidth-adaptive spatiotemporal CoID in collaborative perception. This approach allows robots to progressively select partial spatiotemporal observations and share with others, while adapting to communication constraints that dynamically change over time. We evaluate our approach across various scenarios in connected autonomous driving simulations. Experimental results validate that our approach enables CoID and adapts to dynamic communication bandwidth changes. In addition, our approach achieves 8%-56% overall improvements in terms of covisible object retrieval for CoID and data sharing efficiency, which outperforms previous techniques and achieves the state-of-the-art performance. More information is available at: https://gaopeng5.github.io/acoid.
Peng Gao 0009, Williard Joshua Jose, Hao Zhang 0011
ICRA3
2025 Self-Reflective Perceptual Adaptation for Robust Ground Navigation in Unstructured Off-Road Environments
abstract
Autonomous ground robots navigating unstructured off-road environments face perceptual challenges, such as sensor obscuration or failure, which can lead to inaccurate perception or navigation failures. While robot adaptation has recently gained increasing attention, self-reflective robot adaptation, where robots understand and adjust to their own sensor limitations, remains under-explored. This paper proposes a novel approach for self-reflective perceptual adaptation in order to enhance robust off-road navigation. Our approach enables a robot to identify its own perceptual difficulties and dynamically adapt in challenging environments. The key novelty is learning a modality-invariant perceptual representation that encodes shared sensor data into a compact feature space. Within this representation space, the robot's dynamics model is also learned, which enables accurate prediction of future navigation paths. Extensive experiments in off-road environments with sensor obstructions and failures demonstrate that our method significantly improves adaptive capabilities and outperforms baseline and state-of-the-art approaches. More details of this work are provided on the project website: https://hcrlab.gitlab.io/project/srpa.
Sriram Siva, Oscar Youngquist, Maggie B. Wigness, John G. Rogers III, Hao Zhang 0011
ICRA5
2023 Collaborative Scheduling with Adaptation to Failure for Heterogeneous Robot Teams
abstract
Collaborative scheduling is an essential ability for a team of heterogeneous robots to collaboratively complete complex tasks, e.g., in a multi-robot assembly application. To enable collaborative scheduling, two key problems should be addressed, including allocating tasks to heterogeneous robots and adapting to robot failures in order to guarantee the completion of all tasks. In this paper, we introduce a novel approach that integrates deep bipartite graph matching and imitation learning for heterogeneous robots to complete complex tasks as a team. Specifically, we use a graph attention network to represent attributes and relationships of the tasks. Then, we formulate collaborative scheduling with failure adaptation as a new deep learning-based bipartite graph matching problem, which learns a policy by imitation to determine task scheduling based on the reward of potential task schedules. During normal execution, our approach generates robot-task pairs as potential allocations. When a robot fails, our approach identifies not only individual robots but also subteams to replace the failed robot. We conduct extensive experiments to evaluate our approach in the scenarios of collaborative scheduling with robot failures. Experimental results show that our approach achieves promising, generalizable and scalable results on collaborative scheduling with robot failure adaptation.
Peng Gao 0009, Sriram Siva, Anthony Micciche, Hao Zhang 0011
ICRA4
2023 Deep Masked Graph Matching for Correspondence Identification in Collaborative Perception
abstract
Correspondence identification (CoID) is an essential component for collaborative perception in multi-robot systems, such as connected autonomous vehicles. The goal of CoID is to identify the correspondence of objects observed by multiple robots in their own field of view in order for robots to consistently refer to the same objects. CoID is challenging due to perceptual aliasing, object non-covisibility, and noisy sensing. In this paper, we introduce a novel deep masked graph matching approach to enable CoID and address the challenges. Our approach formulates CoID as a graph matching problem and we design a masked neural network to integrate the multimodal visual, spatial, and GPS information to perform CoID. In addition, we design a new technique to explicitly address object non-covisibility caused by occlusion and the vehicle's limited field of view. We evaluate our approach in a variety of street environments using a high-fidelity simulation that integrates the CARLA and SUMO simulators. The experimental results show that our approach outperforms the previous approaches and achieves state-of-the- art CoID performance in connected autonomous driving applications. Our work is available at: https://github.com/gaopeng5/DMGM.git.
Peng Gao 0009, Qingzhao Zhu, Hongsheng Lu, Chuang Gan 0001, Hao Zhang 0011
ICRA5
2023 Failure Explanation in Privacy-Sensitive Contexts: An Integrated Systems Approach
abstract
In this paper, we explore how robots can properly explain failures during navigation tasks with privacy concerns. We present an integrated robotics approach to generate visual failure explanations, by combining a language-capable cognitive architecture (for recognizing intent behind commands), an object- and location-based context recognition system (for identifying the locations of people and classifying the context in which those people are situated) and an infeasibility proof-based motion planner (for explaining planning failures on the basis of contextually mediated privacy concerns). The behavior of this integrated system is validated using a series of experiments in a simulated medical environment.
Sihui Li, Sriram Siva, Terran Mott, Tom Williams 0001, Hao Zhang 0011, Neil Dantam
RO-MAN5
2022 Asynchronous Collaborative Localization by Integrating Spatiotemporal Graph Learning with Model-Based Estimation
abstract
Collaborative localization is an essential capability for a team of robots such as connected vehicles to collaboratively estimate object locations from multiple perspectives with reliant cooperation. To enable collaborative localization, four key challenges must be addressed, including modeling complex relationships between observed objects, fusing observations from an arbitrary number of collaborating robots, quantifying localization uncertainty, and addressing latency of robot communications. In this paper, we introduce a novel approach that integrates uncertainty-aware spatiotemporal graph learning and model-based state estimation for a team of robots to collaboratively localize objects. Specifically, we introduce a new uncertainty-aware graph learning model that learns spatiotemporal graphs to represent historical motions of the objects observed by each robot over time and provides uncertainties in object localization. Moreover, we propose a novel method for integrated learning and model-based state estimation, which fuses asynchronous observations obtained from an arbitrary number of robots for collaborative localization. We evaluate our approach in two collaborative object localization scenarios in simulations and on real robots. Experimental results show that our approach outperforms previous methods and achieves state-of-the-art performance on asynchronous collaborative localization.
Peng Gao 0009, Brian Reily, Hongsheng Lu, Qingzhao Zhu, Hao Zhang 0011
ICRA6
2022 NAUTS: Negotiation for Adaptation to Unstructured Terrain Surfaces
abstract
When robots operate in real-world off-road environments with unstructured terrains, the ability to adapt their navigational policy is critical for effective and safe navigation. However, off-road terrains introduce several challenges to robot navigation, including dynamic obstacles and terrain uncertainty, leading to inefficient traversal or navigation failures. To address these challenges, we introduce a novel approach for adaptation by negotiation that enables a ground robot to adjust its navigational behaviors through a negotiation process. Our approach first learns prediction models for various navigational policies to function as a terrain-aware joint local controller and planner. Then, through a new negotiation process, our approach learns from various policies' interactions with the environment to agree on the optimal combination of policies in an online fashion to adapt robot navigation to unstructured off-road terrains on the fly. Additionally, we implement a new optimization algorithm that offers the optimal solution for robot negotiation in real-time during execution. Experimental results have validated that our method for adaptation by negotiation outperforms previous methods for robot navigation, especially over unseen and uncertain dynamic terrains.
Sriram Siva, Maggie B. Wigness, John G. Rogers III, Long Quang, Hao Zhang 0011
IROS5
2021 Multi-view Sensor Fusion by Integrating Model-based Estimation and Graph Learning for Collaborative Object Localization
abstract
Collaborative object localization aims to collaboratively estimate locations of objects observed from multiple views or perspectives, which is a critical ability for multi-agent systems such as connected vehicles. To enable collaborative localization, several model-based state estimation and learning-based localization methods have been developed. Given their encouraging performance, model-based state estimation often lacks the ability to model the complex relationships among multiple objects, while learning-based methods are typically not able to fuse the observations from an arbitrary number of views and cannot well model uncertainty. In this paper, we introduce a novel spatiotemporal graph filter approach that integrates graph learning and model-based estimation to perform multi-view sensor fusion for collaborative object localization. Our approach models complex object relationships using a new spatiotemporal graph representation and fuses multi-view observations in a Bayesian fashion to improve location estimation under uncertainty. We evaluate our approach in the applications of connected autonomous driving and multiple pedestrian localization. Experimental results show that our approach outperforms previous techniques and achieves the state-of-the-art performance on collaborative localization.
Peng Gao 0009, Hongsheng Lu, Hao Zhang 0011
ICRA4
2021 Adaptation to Team Composition Changes for Heterogeneous Multi-Robot Sensor Coverage
abstract
We consider the problem of multi-robot sensor coverage, which deals with deploying a multi-robot team in an environment and optimizing the sensing quality of the overall environment. As real-world environments involve a variety of sensory information, and individual robots are limited in their available number of sensors, successful multi-robot sensor coverage requires the deployment of robots in such a way that each individual team member’s sensing quality is maximized. Additionally, because individual robots have varying complements of sensors and both robots and sensors can fail, robots must be able to adapt and adjust how they value each sensing capability in order to obtain the most complete view of the environment, even through changes in team composition. We introduce a novel formulation for sensor coverage by multi-robot teams with heterogeneous sensing capabilities that maximizes each robot's sensing quality, balancing the varying sensing capabilities of individual robots based on the overall team composition. We propose a solution based on regularized optimization that uses sparsity-inducing terms to ensure a robot team focuses on all possible event types, and which we show is proven to converge to the optimal solution. Through extensive simulation, we show that our approach is able to effectively deploy a multi-robot team to maximize the sensing quality of an environment, responding to failures in the multi-robot team more robustly than non-adaptive approaches.
Brian Reily, Terran Mott, Hao Zhang 0011
ICRA3
2021 Team Assignment for Heterogeneous Multi-Robot Sensor Coverage through Graph Representation Learning
abstract
Sensor coverage is the critical multi-robot problem of maximizing the detection of events in an environment through the deployment of multiple robots. Large multi-robot systems are often composed of simple robots that are typically not equipped with a complete set of sensors, so teams with comprehensive sensing abilities are required to properly cover an area. Robots also exhibit multiple forms of relationships (e.g., communication connections or spatial distribution) that need to be considered when assigning robot teams for sensor coverage. To address this problem, in this paper we introduce a novel formulation of sensor coverage by multi-robot systems with heterogeneous relationships as a graph representation learning problem. We propose a principled approach based on the mathematical framework of regularized optimization to learn a unified representation of the multi-robot system from the graphs describing the heterogeneous relationships and to identify the learned representation’s underlying structure in order to assign the robots to teams. To evaluate the proposed approach, we conduct extensive experiments on simulated multi-robot systems and a physical multi-robot system as a case study, demonstrating that our approach is able to effectively assign teams for heterogeneous multi-robot sensor coverage.
Brian Reily, Hao Zhang 0011
ICRA2
2021 Edge-Assisted Collaborative Perception in Autonomous Driving: A Reflection on Communication Design
Ruozhou Yu, Dejun Yang, Hao Zhang 0011
SEC3
2021 An Integrated Approach to Context-Sensitive Moral Cognition in Robot Cognitive Architectures
abstract
Acceptance of social robots in human-robot collaborative environments depends on the robots’ sensitivity to human moral and social norms. Robot behavior that violates norms may decrease trust and lead human interactants to blame the robot and view it negatively. Hence, for long-term acceptance, social robots need to detect possible norm violations in their action plans and refuse to perform such plans. This paper integrates the Distributed, Integrated, Affect, Reflection, Cognition (DIARC) robot architecture (implemented in the Agent Development Environment (ADE)) with a novel place recognition module and a norm-aware task planner to achieve context-sensitive moral reasoning. This will allow the robot to reject inappropriate commands and comply with context-sensitive norms. In a validation scenario, our results show that the robot would not comply with a human command to violate a privacy norm in a private context.
Ryan Blake Jackson, Sihui Li, Santosh Balajee Banisetty, Sriram Siva, Hao Zhang 0011, Neil Dantam, Tom Williams 0001
IROS5
2020 Long-Term Loop Closure Detection through Visual-Spatial Information Preserving Multi-Order Graph Matching
abstract
Loop closure detection is a fundamental problem for simultaneous localization and mapping (SLAM) in robotics. Most of the previous methods only consider one type of information, based on either visual appearances or spatial relationships of landmarks. In this paper, we introduce a novel visual-spatial information preserving multi-order graph matching approach for long-term loop closure detection. Our approach constructs a graph representation of a place from an input image to integrate visual-spatial information, including visual appearances of the landmarks and the background environment, as well as the second and third-order spatial relationships between two and three landmarks, respectively. Furthermore, we introduce a new formulation that formulates loop closure detection as a multi-order graph matching problem to compute a similarity score directly from the graph representations of the query and template images, instead of performing conventional vector-based image matching. We evaluate the proposed multi-order graph matching approach based on two public long-term loop closure detection benchmark datasets, including the St. Lucia and CMU-VL datasets. Experimental results have shown that our approach is effective for long-term loop closure detection and it outperforms the previous state-of-the-art methods.
Peng Gao 0009, Hao Zhang 0011
AAAI2
2020 Long-term Place Recognition through Worst-case Graph Matching to Integrate Landmark Appearances and Spatial Relationships
abstract
Place recognition is an important component for simultaneously localization and mapping in a variety of robotics applications. Recently, several approaches using landmark information to represent a place showed promising performance to address long-term environment changes. However, previous approaches do not explicitly consider changes of the landmarks, i,e., old landmarks may disappear and new ones often appear over time. In addition, representations used in these approaches to represent landmarks are limited, based upon visual or spatial cues only. In this paper, we introduce a novel worst-case graph matching approach that integrates spatial relationships of landmarks with their appearances for long-term place recognition. Our method designs a graph representation to encode distance and angular spatial relationships as well as visual appearances of landmarks in order to represent a place. Then, we formulate place recognition as a graph matching problem under the worst-case scenario. Our approach matches places by computing the similarities of distance and angular spatial relationships of the landmarks that have the least similar appearances (i.e., worst-case). If the worst appearance similarity of landmarks is small, two places are identified to be not the same, even though their graph representations have high spatial relationship similarities. We evaluate our approach over two public benchmark datasets for long-term place recognition, including St. Lucia and CMU-VL. The experimental results have validated that our approach obtains the state-of-the-art place recognition performance, with a changing number of landmarks.
Peng Gao 0009, Hao Zhang 0011
ICRA2
2020 Correspondence Identification in Collaborative Robot Perception through Maximin Hypergraph Matching
abstract
Correspondence identification is an essential problem for collaborative multi-robot perception, with the objective of deciding the correspondence of objects that are observed in the field of view of each robot. In this paper, we introduce a novel maximin hypergraph matching approach that formulates correspondence identification as a hypergraph matching problem. The proposed approach incorporates both spatial relationships and appearance features of objects to improve representation capabilities. It also integrates the maximin theorem to optimize the worst-case scenario in order to address distractions caused by non-covisible objects. In addition, we design an optimization algorithm to address the formulated non-convex non-continuous optimization problem. We evaluate our approach and compare it with seven previous techniques in two application scenarios, including multi-robot coordination on real robots and connected autonomous driving in simulations. Experimental results have validated the effectiveness of our approach in identifying object correspondence from partially overlapped views in collaborative perception, and have shown that the proposed maximin hypergraph matching approach outperforms previous techniques and obtains state-of-the-art performance.
Peng Gao 0009, Ziling Zhang, Hongsheng Lu, Hao Zhang 0011
ICRA5
2020 Representing Multi-Robot Structure through Multimodal Graph Embedding for the Selection of Robot Teams
abstract
Multi-robot systems of increasing size and complexity are used to solve large-scale problems, such as area exploration and search and rescue. A key decision in human-robot teaming is dividing a multi-robot system into teams to address separate issues or to accomplish a task over a large area. In order to address the problem of selecting teams in a multi-robot system, we propose a new multimodal graph embedding method to construct a unified representation that fuses multiple information modalities to describe and divide a multi-robot system. The relationship modalities are encoded as directed graphs that can encode asymmetrical relationships, which are embedded into a unified representation for each robot. Then, the constructed multimodal representation is used to determine teams based upon unsupervised learning. We per-form experiments to evaluate our approach on expert-defined team formations, large-scale simulated multi-robot systems, and a system of physical robots. Experimental results show that our method successfully decides correct teams based on the multifaceted internal structures describing multi-robot systems, and outperforms baseline methods based upon only one mode of information, as well as other graph embedding-based division methods.
Brian Reily, Christopher M. Reardon, Hao Zhang 0011
ICRA3
2020 Simultaneous Learning from Human Pose and Object Cues for Real-Time Activity Recognition
abstract
Real-time human activity recognition plays an essential role in real-world human-centered robotics applications, such as assisted living and human-robot collaboration. Although previous methods based on skeletal data to encode human poses showed promising results on real-time activity recognition, they lacked the capability to consider the context provided by objects within the scene and in use by the humans, which can provide a further discriminant between human activity categories. In this paper, we propose a novel approach to real-time human activity recognition, through simultaneously learning from observations of both human poses and objects involved in the human activity. We formulate human activity recognition as a joint optimization problem under a unified mathematical framework, which uses a regression-like loss function to integrate human pose and object cues and defines structured sparsity-inducing norms to identify discriminative body joints and object attributes. To evaluate our method, we perform extensive experiments on two benchmark datasets and a physical robot in a home assistance setting. Experimental results have shown that our method outperforms previous methods and obtains real-time performance for human activity recognition with a processing speed of 104Hz.
Brian Reily, Qingzhao Zhu, Christopher M. Reardon, Hao Zhang 0011
ICRA4
2020 Voxel-Based Representation Learning for Place Recognition Based on 3D Point Clouds
abstract
Place recognition is a critical component towards addressing the key problem of Simultaneous Localization and Mapping (SLAM). Most existing methods use visual images; whereas, place recognition using 3D point clouds, especially based on the voxel representations, has not been well addressed yet. In this paper, we introduce the novel approach of voxel-based representation learning (VBRL) that uses 3D point clouds to recognize places with long-term environment variations. VBRL splits a 3D point cloud input into voxels and uses multi-modal features extracted from these voxels to perform place recognition. Additionally, VBRL uses structured sparsity-inducing norms to learn representative voxels and feature modalities that are important to match places under long-term changes. Both place recognition, and voxel and feature learning are integrated into a unified regularized optimization formulation. As the sparsity-inducing norms are non-smooth, it is hard to solve the formulated optimization problem. Thus, we design a new iterative optimization algorithm, which has a theoretical convergence guarantee. Experimental results have shown that VBRL performs place recognition well using 3D point cloud data and is capable of learning the importance of voxels and feature modalities.
Sriram Siva, Zachary Nahman, Hao Zhang 0011
IROS3
2019 Visual Place Recognition via Robust ℓ2-Norm Distance Based Holism and Landmark Integration
abstract
Visual place recognition is essential for large-scale simultaneous localization and mapping (SLAM). Long-term robot operations across different time of the days, months, and seasons introduce new challenges from significant environment appearance variations. In this paper, we propose a novel method to learn a location representation that can integrate the semantic landmarks of a place with its holistic representation. To promote the robustness of our new model against the drastic appearance variations due to long-term visual changes, we formulate our objective to use non-squared ℓ2-norm distances, which leads to a difficult optimization problem that minimizes the ratio of the ℓ2,1-norms of matrices. To solve our objective, we derive a new efficient iterative algorithm, whose convergence is rigorously guaranteed by theory. In addition, because our solution is strictly orthogonal, the learned location representations can have better place recognition capabilities. We evaluate the proposed method using two large-scale benchmark data sets, the CMU-VL and Nordland data sets. Experimental results have validated the effectiveness of our new method in long-term visual place recognition applications.
Kai Liu 0018, Hua Wang 0007, Fei Han 0002, Hao Zhang 0011
AAAI4
2019 Learning Robust Multi-label Sample Specific Distances for Identifying HIV-1 Drug Resistance
Lodewijk Brand, Kai Liu 0018, Saad El Beleidy, Hua Wang 0007, Hao Zhang 0011
RECOMB6
2019 Collaborative Localization for Occluded Objects in Connected Vehicular Platform
abstract
Localizing occluded object is a long-term challenge in Advanced Driving Assistant System (ADAS) and autonomous driving research. In this paper, we propose a novel graph-matching based approach that leverages the challenge by adopting the deep learning and multiple-view geometry analysis. Specifically, the 3D scene reconstruction is firstly built by associating the comprehensive graph representations of the multiple-view observations, incorporated with the spatial relationship of the co-visible objects so as their discriminant appearance features. Followed by, the localization for occluded object is achieved by inferring from the reconstructed 3D geometry. We conduct experiments to validate the system in connected vehicular platform in the advanced traffic simulation dataset. The experimental results convincingly indicate the effectiveness of the proposed system in real- time object detection, graph generation, matching and location inference for occluded objects.
Hongsheng Lu, Peng Gao 0009, Ziling Zhang, Hao Zhang 0011
VTC Fall5
2018 Learning Integrated Holism-Landmark Representations for Long-Term Loop Closure Detection
abstract
Loop closure detection is a critical component of large-scale simultaneous localization and mapping (SLAM) in loopy environments. This capability is challenging to achieve in long-term SLAM, when the environment appearance exhibits significant long-term variations across various time of the day, months, and even seasons. In this paper, we introduce a novel formulation to learn an integrated long-term representation based upon both holistic and landmark information, which integrates two previous insights under a unified framework: (1) holistic representations outperform keypoint-based representations, and (2) landmarks as an intermediate representation provide informative cues to detect challenging locations. Our new approach learns the representation by projecting input visual data into a low-dimensional space, which preserves both the global consistency (to minimize representation error) and the local consistency (to preserve landmarks’ pairwise relationship) of the input data. To solve the formulated optimization problem, a new algorithm is developed with theoretically guaranteed convergence. Extensive experiments have been conducted using two large-scale public benchmark data sets, in which the promising performances have demonstrated the effectiveness of the proposed approach.
Fei Han 0002, Hua Wang 0007, Hao Zhang 0011
AAAI3
2018 Learning Multi-Instance Enriched Image Representations via Non-Greedy Ratio Maximization of the l1-Norm Distances
abstract
Multi-instance learning (MIL) has demonstrated its usefulness in many real-world image applications in recent years. However, two critical challenges prevent one from effectively using MIL in practice. First, existing MIL methods routinely model the predictive targets using the instances of input images, but rarely utilize an input image as a whole. As a result, the useful information conveyed by the holistic representation of an input image could be potentially lost. Second, the varied numbers of the instances of the input images in a data set make it infeasible to use traditional learning models that can only deal with single-vector inputs. To tackle these two challenges, in this paper we propose a novel image representation learning method that can integrate the local patches (the instances) of an input image (the bag) and its holistic representation into one single-vector representation. Our new method first learns a projection to preserve both global and local consistencies of the instances of an input image. It then projects the holistic representation of the same image into the learned subspace for information enrichment. Taking into account the content and characterization variations in natural scenes and photos, we develop an objective that maximizes the ratio of the summations of a number of ℓ1-norm distances, which is difficult to solve in general. To solve our objective, we derive a new efficient non-greedy iterative algorithm and rigorously prove its convergence. Promising results in extensive experiments have demonstrated improved performances of our new method that validate its effectiveness.
Kai Liu 0018, Hua Wang 0007, Feiping Nie 0001, Hao Zhang 0011
CVPR4
2018 Omnidirectional Multisensory Perception Fusion for Long-Term Place Recognition
abstract
Over the recent years, long-term place recognition has attracted an increasing attention to detect loops for largescale Simultaneous Localization and Mapping (SLAM) in loopy environments during long-term autonomy. Almost all existing methods are designed to work with traditional cameras with a limited field of view. Recent advances in omnidirectional sensors offer a robot an opportunity to perceive the entire surrounding environment. However, no work has existed thus far to research how omnidirectional sensors can help long-term place recognition, especially when multiple types of omnidirectional sensory data are available. In this paper, we propose a novel approach to integrate observations obtained from multiple sensors from different viewing angles in the omnidirectional observation in order to perform multi-directional place recognition in longterm autonomy. Our approach also answers two new questions when omnidirectional multisensory data is available for place recognition, including whether it is possible to recognize a place with long-term appearance variations when robots approach it from various directions, and whether observations from various viewing angles are the same informative. To evaluate our approach and hypothesis, we have collected the first large-scale dataset that consists of omnidirectional multisensory (intensity and depth) data collected in urban and suburban environments across a year. Experimental results have shown that our approach is able to achieve multi-directional long-term place recognition, and identifies the most discriminative viewing angles from the omnidirectional observation.
Sriram Siva, Hao Zhang 0011
ICRA2
2017 Minimum uncertainty latent variable models for robot recognition of sequential human activities
abstract
Recognition of sequential human activities, such as “sitting down” and “standing up”, is a common but challenging problem in human-robot interaction, which requires modeling their underlying temporal patterns. Although previous sequence modeling methods, such as Hidden Conditional Random Fields (HCRFs), demonstrated satisfactory recognition accuracy, they do not explicitly model the uncertainty in underlying temporal patterns, which can provide valuable information to characterize sequential activities. To address this problem, we introduce a novel Minimum Uncertainty HCRF (MU, or μHCRF). Different from traditional HCRF-based techniques that only utilize the negative log-likelihood of the categories' conditional probability as the loss function, the proposed μ-HCRF also introduces a regularization term to model the underlying temporal pattern of the latent variables. As another theoretical contribution, we provide a derivation to show that the formulated problem has a closed-form solution, and prove that inference of the proposed μHCRF is tractable. Extensive empirical study is performed to evaluate our approach, using four public benchmark datasets. Experimental results have shown that our μHCRFs outperform previous techniques and achieve state-of-the-art performance on human activity recognition, especially on sequential activities.
Fei Han 0002, Christopher M. Reardon, Lynne E. Parker, Hao Zhang 0011
ICRA4
2017 Simultaneous Feature and Body-Part Learning for real-time robot awareness of human behaviors
abstract
Robot awareness of human actions is an essential research problem in robotics with many important real-world applications, including human-robot collaboration and teaming. Over the past few years, depth sensors have become a standard device widely used by intelligent robots for 3D perception, which can also offer human skeletal data in 3D space. Several methods based on skeletal data were designed to enable robot awareness of human actions with satisfactory accuracy. However, previous methods treated all body parts and features equally important, without the capability to identify discriminative body parts and features. In this paper, we propose a novel simultaneous Feature And Body-part Learning (FABL) approach that simultaneously identifies discriminative body parts and features, and efficiently integrates all available information together to enable real-time robot awareness of human behaviors. We formulate FABL as a regression-like optimization problem with structured sparsity-inducing norms to model interrelationships of body parts and features. We also develop an optimization algorithm to solve the formulated problem, which possesses a theoretical guarantee to find the optimal solution. To evaluate FABL, three experiments were performed using public benchmark datasets, including the MSR Action3D and CAD-60 datasets, as well as a Baxter robot in practical assistive living applications. Experimental results show that our FABL approach obtains a high recognition accuracy with a processing speed of the order-of-magnitude of 101Hz, which makes FABL a promising method to enable real-time robot awareness of human behaviors in practical robotics applications.
Fei Han 0002, Christopher M. Reardon, Hao Zhang 0011
ICRA5
2017 Sequence-based multimodal apprenticeship learning for robot perception and decision making
abstract
Apprenticeship learning has recently attracted a wide attention due to its capability of allowing robots to learn physical tasks directly from demonstrations provided by human experts. Most previous techniques assumed that the state space is known a priori or employed simple state representations that usually suffer from perceptual aliasing. Different from previous research, we propose a novel approach named Sequence-based Multimodal Apprenticeship Learning (SMAL), which is capable to simultaneously fusing temporal information and multimodal data, and to integrate robot perception with decision making. To evaluate the SMAL approach, experiments are performed using both simulations and real-world robots in the challenging search and rescue scenarios. The empirical study has validated that our SMAL approach can effectively learn plans for robots to make decisions using sequence of multimodal observations. Experimental results have also showed that SMAL outperforms the baseline methods using individual images.
Fei Han 0002, Hao Zhang 0011
ICRA4
2017 Space-time representation of people based on 3D skeletal data: A review
Fei Han 0002, Brian Reily, William A. Hoff, Hao Zhang 0011
Comput. Vis. Image Underst.4
2017 Real-time gymnast detection and performance analysis with a portable 3D camera
Brian Reily, Hao Zhang 0011, William A. Hoff
Comput. Vis. Image Underst.2
2016 Drosophila Gene Expression Pattern Annotations via Multi-Instance Biological Relevance Learning
abstract
Recent developments in biologyhave produced a large number of gene expression patterns, many of which have been annotated textually with anatomical and developmental terms. These terms spatially correspond to local regions of the images, which are attached collectively to groups of images. Because one does not know which term is assigned to which region of which image in the group, the developmental stage classification and anatomical term annotation turn out to be a multi-instance learning (MIL) problem, which considers input as bags of instances and labels are assigned to the bags. Most existing MIL methods routinely use the Bag-to-Bag (B2B) distances, which, however, are often computationally expensive and may not truly reflect the similarities between the anatomical and developmental terms. In this paper, we approach the MIL problem from a new perspective using the Class-to-Bag (C2B) distances, which directly assesses the relations between annotation terms and image panels. Taking into account the two challenging properties of multi-instance gene expression data, high heterogeneity and weak label association, we computes the C2B distance by introducing class specific distance metrics and locally adaptive significance coefficients.We apply our new approach to automatic gene expression pattern classification and annotation on the Drosophila melanogaster species. Extensive experiments have demonstrated the effectiveness of our new method.
Hua Wang 0007, Cheng Deng 0002, Hao Zhang 0011, Xinbo Gao 0001, Heng Huang 0001
AAAI3
2016 SRAC: Self-Reflective Risk-Aware Artificial Cognitive models for robot response to human activities
abstract
In human-robot teaming, interpretation of human actions, recognition of new situations, and appropriate decision making are crucial abilities for cooperative robots (“co-robots”) to interact intelligently with humans. Given an observation, it is important that human activities are interpreted the same way by co-robots as human peers so that robot actions can be appropriate to the activity at hand. A novel interpretability indicator is introduced to address this issue. When a robot encounters a new scenario, the pretrained activity recognition model, no matter how accurate in a known situation, may not produce the correct information necessary to act appropriately and safely in new situations. To effectively and safely interact with people, we introduce a new generalizability indicator that allows a co-robot to self-reflect and reason about when an observation falls outside the co-robot's learned model. Based on topic modeling and the two novel indicators, we propose a new Self-reflective Risk-aware Artificial Cognitive (SRAC) model, which allows a robot to make better decisions by incorporating robot action risks and identifying new situations. Experiments both using real-world datasets and on physical robots suggest that our SRAC model significantly outperforms the traditional methodology and enables better decision making in response to human behaviors.
Hao Zhang 0011, Christopher M. Reardon, Fei Han 0002, Lynne E. Parker
ICRA1
2016 Enforcing Template Representability and Temporal Consistency for Adaptive Sparse Tracking
Fei Han 0002, Hua Wang 0007, Hao Zhang 0011
IJCAI4
2016 Unified robot learning of action labels and motion trajectories from 3D human skeletal data
abstract
Currently, robot learning of human activities is mainly studied in two largely disconnected domains: high level semantics understanding in human activity recognition, and low level motion trajectory reproduction in robot imitation learning. The critical problem of human activity unified learning (HAUL) was not well studied in previous work. One important challenge is the lack of a representation that can be learned from both levels. We introduce a novel approach to address this HAUL problem at the representation level, by simultaneously learning action labels and motion trajectories from publicly available 3D human skeletal datasets, thus avoiding additional human labor for data collection. Our approach builds a subject and body position independent shared skeleton, and extracts features of skeletal activities based on this model. Then the extracted features are encoded by the parameter set of Gaussian Mixture Models to construct the unified representations. The proposed compact representation can be directly applied to identify activity labels when combined with Support Vector Machines, and can be also employed to generate trajectories of the learned activities when combined with Gaussian Mixture Regression on a robot. Finally, an inverse kinematic mapping is developed to transfer human skeletal trajectories to joint angle sequences in the robot's embodiment. Empirical studies using simulation and real humanoid robots demonstrate that our approach achieves promising performance on robot unified learning of human action labels and motion trajectories, effectively addressing the HAUL problem.
Hao Zhang 0011, Lynne E. Parker
RO-MAN2
2016 Web-video-mining-supported workflow modeling for laparoscopic surgeries
Rui Liu 0003, Xiaoli Zhang 0002, Hao Zhang 0011
Artif. Intell. Medicine3
2016 CoDe4D: Color-Depth Local Spatio-Temporal Features for Human Activity Recognition From RGB-D Videos
abstract
Human activity recognition has a variety of important real-world applications, such as video analysis, surveillance, and human-robot interaction. As a promising video representation method, local spatio-temporal (LST) features have received increasing attention from computer vision, machine learning, and robotics communities. However, approaches based on traditional LST features only use color information, which face several challenges, such as illumination changes and dynamic backgrounds. The recent availability of commercial color-depth cameras makes it much cheaper, faster, and easier to acquire depth information, which provides a potential to implement more discriminative and robust LST features. In this paper, we introduce the new 4-D color-depth (CoDe4D) LST feature that incorporates both intensity and depth information acquired from RGB-D cameras. Our feature detector constructs a saliency map through applying independent filters in xyzt dimension to represent texture, shape and pose variations, and selects its local maxima as interest points. Our multichannel orientation histogram descriptor applies a 4-D support region, which is adaptive to linear perspective view changes, on each interest point. Then, image gradients of color-depth patches within the support region are computed and quantized using a spherical coordinate-based method to form a final feature vector. We build a complete activity recognition system by combining our features with bag-of-features representations and support vector machines. To evaluate the performance of our CoDe4D LST features and the complete system, we conduct experiments using four benchmark color-depth human activity data sets, including UTK Action3-D, Berkeley MHAD, ACT42, and MSR daily activity 3-D data sets. Experimental results demonstrate the promising representative power of our CoDe4D features, which obtain the state-of-the-art performance on activity recognition from RGB-D visual data.
Hao Zhang 0011, Lynne E. Parker
IEEE Trans. Circuits Syst. Video Technol.1
2015 Bio-inspired predictive orientation decomposition of skeleton trajectories for real-time human activity prediction
abstract
Activity prediction is an essential task in practical human-centered robotics applications, such as security, assisted living, etc., which targets at inferring ongoing human activities based on incomplete observations. To address this challenging problem, we introduce a novel bio-inspired predictive orientation decomposition (BIPOD) approach to construct representations of people from 3D skeleton trajectories. Our approach is inspired by biological research in human anatomy. In order to capture spatio-temporal information of human motions, we spatially decompose 3D human skeleton trajectories and project them onto three anatomical planes (i.e., coronal, transverse and sagittal planes); then, we describe short-term time information of joint motions and encode high-order temporal dependencies. By estimating future skeleton trajectories that are not currently observed, we endow our BIPOD representation with the critical predictive capability. Empirical studies validate that our BIPOD approach obtains promising performance, in terms of accuracy and efficiency, using a physical TurtleBot2 robotic platform to recognize ongoing human activities. Experiments on benchmark datasets further demonstrate that our new BIPOD representation significantly outperforms previous approaches for real-time activity classification and prediction from 3D human skeleton trajectories.
Hao Zhang 0011, Lynne E. Parker
ICRA1
2015 Adaptive human-centered representation for activity recognition of multiple individuals from 3D point cloud sequences
abstract
Activity recognition of multi-individuals (ARMI) within a group, which is essential to practical human-centered robotics applications such as childhood education, is a particularly challenging and previously not well studied problem. We present a novel adaptive human-centered (AdHuC) representation based on local spatio-temporal features (LST) to address ARMI in a sequence of 3D point clouds. Our human-centered detector constructs affiliation regions to associate LST features with humans by mining depth data and using a cascade of rejectors to localize humans in 3D space. Then, features are detected within each affiliation region, which avoids extracting irrelevant features from dynamic background clutter and addresses moving cameras on mobile robots. Our feature descriptor is able to adapt its support region to linear perspective view variations and encode multi-channel information (i.e., color and depth) to construct the final representation. Empirical studies validate that the AdHuC representation obtains promising performance on ARMI using a Meka humanoid robot to play multi-people Simon Says games. Experiments on benchmark datasets further demonstrate that our adaptive human-centered representation outperforms previous approaches for activity recognition from color-depth data.
Hao Zhang 0011, Christopher M. Reardon, Lynne E. Parker
ICRA1
2015 Feature Space Decomposition for effective robot adaptation
abstract
Adaptation is an essential capability for intelligent robots to work in new environments. In the learning framework of Programming by Demonstration (PbD) and Reinforcement Learning (RL), a robot usually learns skills from a latent feature space obtained by dimension reduction techniques. Because the latent space is optimized for a specific environment during the training phase, it typically contains fewer variations. Accordingly, searching for a solution within the latent space can be less effective for robot adaptation to new environments with unseen changes. In this paper, we propose a novel Feature Space Decomposition (FSD) approach to effectively address the robot adaptation problem, which is directly applicable to the learning framework based on PbD and RL. Our FSD method decomposes the high-dimensional original features extracted from the demonstration data into principal and non-principal feature space. Then, the non-principal features are used to form a new low-dimensional search space for autonomous robot adaptation based on RL, which is initialized using a generalized trajectory represented by a Gaussian Mixture Model that is learned from the principal features. The scalability of our FSD approach guarantees that optimal solutions can be found in the new non-principal space, if they exist in the original feature space. Experimental results on real robots validate that our FSD approach enables the robots to effectively adapt to new environments, and is usually able to find optimal solutions more quickly than traditional approaches when significant environment changes occur.
Hao Zhang 0011, Lynne E. Parker
IROS2
2015 Response prompting for intelligent robot instruction of students with intellectual disabilities
abstract
Instruction of students with intellectual disability (ID) presents both unique challenges and a compelling opportunity for socially embedded robots to empower an important group in our population. We propose the creation of an autonomous, intelligent robot instructor (IRI) to teach socially valid life skills to students with ID. We present the construction of a complete IRI system for this purpose. Experimental results show the IRI is capable of teaching a non-trivial life skill to students with ID, and participants feel interaction with the IRI is beneficial.
Christopher M. Reardon, Hao Zhang 0011, Rachel Wright, Lynne E. Parker
RO-MAN2
2015 Fuzzy Temporal Segmentation and Probabilistic Recognition of Continuous Human Daily Activities
abstract
Understanding human activities is an essential capability for intelligent robots to help people in a variety of applications. Humans perform activities in a continuous fashion, and transitions between temporally adjacent activities are gradual. Our Fuzzy Segmentation and Recognition (FuzzySR) algorithm explicitly reasons about gradual transitions between continuous human activities. Our objective is to simultaneously segment a given video into a sequence of events and recognize the activity contained in each event. The algorithm uniformly segments the video into a sequence of nonoverlapping blocks, each lasting a short period of time. Then, a multivariable time series is formed by concatenating block-level human activity summaries that are computed using topic models over local spatiotemporal features extracted from each block. Through encoding an event as a fuzzy set with fuzzy boundaries to represent gradual transitions, our approach is capable of segmenting the continuous visual data into a sequence of fuzzy events. By incorporating all block summaries contained in an event, our algorithm determines the activity label for each event. To evaluate performance, we conduct experiments using six datasets. Our algorithm shows promising continuous activity segmentation results on these datasets and obtains the event-level activity recognition precision of 42.6%, 60.4%, 65.2%, and 78.9% on the Hollywood-2, CAD-60, ACT $4^2$, and UTK-CAP datasets, respectively.
Hao Zhang 0011, Wenjun Zhou 0001, Lynne E. Parker
IEEE Trans. Hum. Mach. Syst.1
2015 Correlation range query for effective recommendations
Wenjun Zhou 0001, Hao Zhang 0011
World Wide Web2
2014 Simplex-Based 3D Spatio-temporal Feature Description for Action Recognition
abstract
We present a novel feature description algorithm to describe 3D local spatio-temporal features for human action recognition. Our descriptor avoids the singularity and limited discrimination power issues of traditional 3D descriptors by quantizing and describing visual features in the simplex topological vector space. Specifically, given a feature's support region containing a set of 3D visual cues, we decompose the cues' orientation into three angles, transform the decomposed angles into the simplex space, and describe them in such a space. Then, quadrant decomposition is performed to improve discrimination, and a final feature vector is composed from the resulting histograms. We develop intuitive visualization tools for analyzing feature characteristics in the simplex topological vector space. Experimental results demonstrate that our novel simplex-based orientation decomposition (SOD) descriptor substantially outperforms traditional 3D descriptors for the KTH, UCF Sport, and Hollywood-2 benchmark action datasets. In addition, the results show that our SOD descriptor is a superior individual descriptor for action recognition.
Hao Zhang 0011, Wenjun Zhou 0001, Christopher M. Reardon, Lynne E. Parker
CVPR1
2014 Fuzzy segmentation and recognition of continuous human activities
abstract
Most previous research has focused on classifying single human activities contained in segmented videos. However, in real-world scenarios, human activities are inherently continuous and gradual transitions always exist between temporally adjacent activities. In this paper, we propose a Fuzzy Segmentation and Recognition (FuzzySR) algorithm to explicitly model this gradual transition. Our goal is to simultaneously segment a given video into events and recognize the activity contained in each event. Specifically, our algorithm uniformly partitions the video into a sequence of non-overlapping blocks, each of which lasts a short period of time. Then, a multi-variable time series is creatively formed through concatenating the block-level human activity summaries that are computed using topic models over each block's local spatio-temporal features. By representing an event as a fuzzy set that has fuzzy boundaries to model gradual transitions, our algorithm is able to segment the video into a sequence of fuzzy events. By incorporating all block summaries contained in an event, the proposed algorithm determines the most appropriate activity category for each event. We evaluate our algorithm's performance using two real-world benchmark datasets that are widely used in the machine vision community. We also demonstrate our algorithm's effectiveness in important robotics applications, such as intelligent service robotics. For all used datasets, our algorithm achieves promising continuous human activity segmentation and recognition results.
Hao Zhang 0011, Wenjun Zhou 0001, Lynne E. Parker
ICRA1
2013 Approximate l-Fold Cross-Validation with Least Squares SVM and Kernel Ridge Regression
abstract
Kernel methods have difficulties scaling to large modern data sets. The scalability issues are based on computational and memory requirements for working with a large matrix. These requirements have been addressed over the years by using low-rank kernel approximations or by improving the solvers' scalability. However, Least Squares Support Vector Machines (LS-SVM), a popular SVM variant, and Kernel Ridge Regression still have several scalability issues. In particular, the O(n^3) computational complexity for solving a single model, and the overall computational complexity associated with tuning hyper parameters are still major problems. We address these problems by introducing an O(nlog n) approximate l-fold cross-validation method that uses a multi-level circulant matrix to approximate the kernel. In addition, we prove our algorithm's computational complexity and present empirical runtimes on data sets with approximately one million data points. We also validate our approximate method's effectiveness at selecting hyper parameters on real world and standard benchmark data sets. Lastly, we provide experimental results on using a multi level circulant kernel approximation to solve LS-SVM problems with hyper parameters selected using our method.
Richard E. Edwards, Hao Zhang 0011, Lynne E. Parker, Joshua R. New
ICMLA (1)2
2013 Zebrafish Larva Locomotor Activity Analysis Using Machine Learning Techniques
abstract
Zebra fish larvae have become a popular model organism to investigate genetic and environmental factors affecting behavior. However, difficulties exist in the analysis of complex behaviors from a large array of larvae. In this paper, we present the new application of machine learning techniques in bioinformatics to automatically detect and investigate the locomotor activities of zebra fish larvae. To achieve this, twelve features were defined and seven unsupervised learning methods were implemented. Next, seven performance measures were applied to evaluate and compare these methods. In order to empirically evaluate the machine learning algorithms, a large dataset was collected that contained 6847 valid instances. Using this dataset, the characteristics of the features were analyzed and the most appropriate unsupervised learning algorithm, i.e., Unweighted Pair Group Method with Arithmetic mean (UPGMA), for locomotor activity analysis was identified. In addition, UPGMA's ability to reveal underlying patterns of zebra fish locomotor activities was demonstrated. In general, this study shows that machine learning techniques have the potential to construct effective, high-throughput systems to automate the process of identifying zebra fish behaviors influenced by genetic manipulation, pharmaceuticals, and environmental toxins.
Hao Zhang 0011, Scott C. Lenaghan, Michelle H. Connolly, Lynne E. Parker
ICMLA (1)1
2013 Correlation Range Query
Wenjun Zhou 0001, Hao Zhang 0011
WAIM2
2013 Real-Time Multiple Human Perception With Color-Depth Cameras on a Mobile Robot
abstract
The ability to perceive humans is an essential requirement for safe and efficient human-robot interaction. In real-world applications, the need for a robot to interact in real time with multiple humans in a dynamic, 3-D environment presents a significant challenge. The recent availability of commercial color-depth cameras allow for the creation of a system that makes use of the depth dimension, thus enabling a robot to observe its environment and perceive in the 3-D space. Here we present a system for 3-D multiple human perception in real time from a moving robot equipped with a color-depth camera and a consumer-grade computer. Our approach reduces computation time to achieve real-time performance through a unique combination of new ideas and established techniques. We remove the ground and ceiling planes from the 3-D point cloud input to separate candidate point clusters. We introduce the novel information concept, depth of interest, which we use to identify candidates for detection, and that avoids the computationally expensive scanning-window methods of other approaches. We utilize a cascade of detectors to distinguish humans from objects, in which we make intelligent reuse of intermediary features in successive detectors to improve computation. Because of the high computational cost of some methods, we represent our candidate tracking algorithm with a decision directed acyclic graph, which allows us to use the most computationally intense techniques only where necessary. We detail the successful implementation of our novel approach on a mobile robot and examine its performance in scenarios with real-world challenges, including occlusion, robot motion, nonupright humans, humans leaving and reentering the field of view (i.e., the reidentification challenge), human-object and human-human interaction. We conclude with the observation that the incorporation of the depth information, together with the use of modern techniques in new ways, we are able to create an accurate system for real-time 3-D perception of humans by a mobile robot.
Hao Zhang 0011, Christopher M. Reardon, Lynne E. Parker
IEEE Trans. Cybern.1
2012 Regularized Probabilistic Latent Semantic Analysis with Continuous Observations
abstract
Probabilistic latent semantic analysis (PLSA) has been widely used in the machine learning community. However, the original PLSAs are not capable of modeling real-valued observations and usually have severe problems with over fitting. To address both issues, we propose a novel, regularized Gaussian PLSA (RG-PLSA) model that combines Gaussian PLSAs and hierarchical Gaussian mixture models (HGMM). We evaluate our model on supervised human action recognition tasks, using two publicly available datasets. Average classification accuracies of 97.69% and 93.72% are achieved on the Weizmann and KTH Action Datasets, respectively, which demonstrate that the RG-PLSA model outperforms Gaussian PLSAs and HGMMs, and is comparable to the state of the art.
Hao Zhang 0011, Richard E. Edwards, Lynne E. Parker
ICMLA (1)1
2011 4-dimensional local spatio-temporal features for human activity recognition
abstract
Recognizing human activities from common color image sequences faces many challenges, such as complex backgrounds, camera motion, and illumination changes. In this paper, we propose a new 4-dimensional (4D) local spatio-temporal feature that combines both intensity and depth information. The feature detector applies separate filters along the 3D spatial dimensions and the 1D temporal dimension to detect a feature point. The feature descriptor then computes and concatenates the intensity and depth gradients within a 4D hyper cuboid, which is centered at the detected feature point, as a feature. For recognizing human activities, Latent Dirichlet Allocation with Gibbs sampling is used as the classifier. Experiments are performed on a newly created database that contains six human activities, each with 33 samples with complex variations. Experimental results demonstrate the promising performance of the proposed features for the task of human activity recognition.
Hao Zhang 0011, Lynne E. Parker
IROS1